Diagnostic Workbench
Join invoices, traces, application metadata, quality evaluations, and outcomes during a private engagement.
Production AI economics tools
Free, private tools for measuring production AI spend, comparing alternatives, and reviewing cost controls before you change a model, provider, or infrastructure stack.
01–04 / Public tools
Start with assumptions, a usage export, or the production code. Each tool is useful on its own, makes its limits explicit, and points to the next measurement needed for a defensible decision.
Compare two production AI scenarios using retries, quality, human review, operating labor, fixed commitments, and cost per successful outcome.
Open tool 02 / Free · Browser-onlyAnalyze an AI usage CSV locally to see spend concentration, workload attribution, period totals, and missing measurement fields.
Open tool 03 / Free · Open sourceGive Codex or Claude an evidence-based method for reviewing AI call paths, cost controls, retries, context, caching, and instrumentation.
Open tool 04 / Free · Browser-onlyCompare projected cost, latency, quality, routing, retries, prompt caching, and full-response caching before running a representative local benchmark.
Open tool05–08 / Specialist tools
Join invoices, traces, application metadata, quality evaluations, and outcomes during a private engagement.
Turn a public scenario into a repeatable local test against representative prompts, real endpoints, and a workload-specific quality rubric.
Simulate routing, fallback, quality thresholds, and cascade costs before changing a production path.
Normalize the baseline and verify savings without hiding quality, volume, price, or workload-mix changes.
When the public evidence is not enough
Bring the invoices, one representative workload, and the quality floor the system has to preserve.
Request a cost review →