How the workbench calculates—and what makes an estimate stronger.
The workbench is a transparent scenario model. It does not infer a price from a project label. Every output comes from visible labor, usage, vendor, operating, and uncertainty assumptions that a user can inspect and change.
Calculation model
Build range = professional-services labor + one-time costs + named uncertainty impacts.
Monthly run rate = model/API workloads + infrastructure/services + human operations and support.
Year-one range = build range + twelve months of the modeled run rate.
Model workload = requests per month × ((input tokens × applicable input/cached rate) + (output tokens × output rate)). Rates are normalized per one million tokens.
Blended input rate = list input rate × (1 − cached share) + cached rate × cached share. The cached share is your assumption about how much of each request the provider can serve from cache.
Low and high values are normalized if entered in reverse. Named risk impacts are calculated against the high build subtotal and displayed separately. They are not probabilities or compliance scores.
Why ranges and sensitivity appear
The U.S. Government Accountability Office’s cost-estimating guidance emphasizes a defined technical baseline, documented assumptions, sensitivity analysis, risk and uncertainty analysis, and updating estimates as actual information becomes available. The workbench borrows these general transparency practices; it does not claim to implement GAO’s full institutional cost-estimating method.
AI lifecycle and review lanes
The capability and review prompts are informed by the need to treat AI work as a lifecycle involving design, data, evaluation, deployment, monitoring, and governance—not merely model integration. Selecting a review lane records scope; it never certifies compliance.
Technology cost references
The embedded OpenAI reference catalog was checked on August 29, 2026 and reflects the July 30, 2026 reduction (Terra −20% to $2.00/$12.00, Luna −80% to $0.20/$1.20 per million tokens; Sol unchanged at its $5.00/$30.00 list rate). A promotional short-context rate for Sol announced on August 22, 2026 is deliberately not used, so the model does not understate cost if it lapses. Prices change. Provider terms may vary by processing mode, context length, caching, data residency, modality, tool use, region, volume, or contract. The workbench therefore keeps every workload assumption — requests, tokens and caching share — visible and editable, and prints the rate and its check date next to each result.
- OpenAI API pricing
- Google Cloud Pricing Calculator
- AWS Pricing Calculator
- FinOps Open Cost and Usage Specification
How the sensitivity ranking is built
The ranking is a one-at-a-time analysis. For every assumption that carries a low and a high, the workbench computes the width of that one range at year-one scale, holding every other assumption fixed:
Labor line: (high hours − low hours) × rate. One-time cost: high − low. Infrastructure or operating line: (high − low) × 12. Named risk driver: high build subtotal × impact percentage.
Those widths sum exactly to the width of the year-one range. That identity is the point of the panel: the bars are not a heuristic ranking, they are an exhaustive decomposition of the range the workbench just quoted, so the longest bar is provably the assumption most worth resolving first.
Model workloads are the exception, because they are entered as a single point estimate rather than a range. They are therefore stressed separately, by a percentage you choose, and are labelled as sitting outside the stated range rather than being silently folded into it.
Where the rates come from, and when to stop trusting them
The embedded OpenAI reference catalog was checked on August 29, 2026 and reflects the July 30, 2026 reduction (Terra −20% to $2.00/$12.00, Luna −80% to $0.20/$1.20 per million tokens; Sol unchanged at its $5.00/$30.00 list rate). A promotional short-context rate for Sol announced on August 22, 2026 is deliberately not used, so the model does not understate cost if it lapses.
The catalog is a dated planning snapshot typed in by hand and versioned with the site. It is not a live feed, and there is no runtime request to any provider. The workbench prints the check date beside every workload, shows how many days old the catalog is, and carries both into the CSV and JSON exports, so a scenario circulated three months later announces its own age instead of passing as current. The threshold language is deliberate: under thirty days the rates are usable for planning, past thirty days each rate needs verifying before it informs a commitment, and past ninety days they are placeholders rather than prices.
List price is also only one of the inputs that decide what a token actually costs. As published on the provider's own pricing page, batch processing is priced at half the standard input and output rate, data residency adds ten percent, the quoted standard rates apply to context lengths under 270K, cached input is priced at roughly a tenth of fresh input, and built-in tools are billed separately — web search was listed at $10.00 per thousand calls, with priority and flex processing tiers changing the arithmetic again. Verified against the archived OpenAI API pricing page of 16 June 2026; check the live page before procurement.
Because of all that, the practical answer to staleness is not a fresher table — it is an editable one. Every workload accepts a rate override: tick Use my own rate, type your negotiated, promotional, batch or regional figure, and the calculation, the exports and the share link all carry your number and mark it as yours rather than the catalog's.
What the model deliberately excludes
Naming the boundary is part of the estimate. The workbench does not model: fully loaded employment cost (benefits, payroll tax, recruiting, equipment, software seats, workspace, management overhead) when labor rates are entered as employee rates rather than vendor rates; calendar schedule, sequencing, or waiting on other people — hours are not weeks; price inflation or contract escalation, since all rates are held flat across the twelve months; taxes and currency; and the cost of the decision not being made. Risk drivers are percentage impacts you chose applied to the high build subtotal. They are not probabilities, confidence intervals, contingency reserves in the GAO sense, or a compliance posture.
Scenarios, sharing and where the data sits
Saved scenarios and the browser copy of the current scenario live in this browser's local storage and are never transmitted. A share link encodes the scenario into the URL fragment, which browsers do not send to the server; anyone holding the link can read every assumption in it, so treat one as you would treat the spreadsheet. Imported scenarios — from a link or from local storage — are re-validated against a bounded schema on load: array lengths, string lengths, numeric ranges, and the model key are all clamped, and anything unrecognized is discarded rather than repaired.
What specification confidence means
The score measures whether eight planning fields are populated: scenario name, users, job to be done, capabilities, planning maturity, labor detail, runtime detail, and review lanes. It does not measure project feasibility, estimate accuracy, safety, compliance, or likelihood of success.
Responsible use
Use the workbench to expose assumptions, compare approaches, and prepare discovery. Before committing money or making procurement, tax, legal, staffing, security, or regulatory decisions, verify current provider prices and obtain the relevant project-specific reviews.