Make AI economics defensible.
Practical tools to assess, model, compare, evaluate, and operate AI workloads using complete costs, verified outcomes, and explicit uncertainty.
The workbench journey
Assess and Model are the two stages you can walk through here. Compare, Evaluate, and Operate are named as roadmap stages only — they are not functioning tools yet, and nothing below links into them.
- 1. AssessReadiness, scope, evidence, gaps
- 2. ModelComplete cost & unit economics
- 3. CompareFuture stage — not part of this preview
- 4. EvaluateFuture stage — not part of this preview
- 5. OperateFuture stage — not part of this preview
Available tools
-
AI Economics Readiness
Determine whether one AI workload has enough evidence for defensible economic analysis, and capture brief practice-level context. Produces a measurement plan — never a maturity score.
-
Probabilistic Unit Economics
Calculate only the economic metrics your readiness evidence supports, and show uncertainty as P10 / median / P90 ranges rather than a single fabricated number.
Worked example used throughout this preview
Northwind Outfitters (fictional, mid-size direct-to-consumer retailer) — a Tier-1 customer-support assistant that drafts replies to order-status and return-policy tickets. High-confidence replies send automatically; low-confidence tickets route to a human agent. Priya Nair, FinOps analyst, is assessing it after Finance flagged rising AI vendor spend, reviewing June 2026 (2026-06-01–2026-06-30).
Uncertainty in this workbench is communicated with low/most-likely/high ranges and simulated outcomes — an approach inspired by the quantitative discipline of FAIR, which uses ranges instead of colour ratings or false precision. This is not a FAIR assessment, implementation, or certification.
- Complete — Practice Context
- In progress — Workload (step 5 of 9)
- Incomplete — Measurement Plan
- Incomplete — Unit Economics
Cost attribution
It's fine — expected, even — for this number to be partial early on. Say what's missing rather than rounding up to guess at a total.
Unit economics — probabilistic results
Design fixture — not a validated calculationTransferred inputs & configuration
- Currency
- USD
- Measurement period
- 2026-06-01 – 2026-06-30
- Total cost
- Estimated range
- Attempts
- 8,400 (Observed, fixed)
- Successful outcomes
- 5,310 (Derived, fixed) — computed as auto-sent replies − reopened-within-5-days (6,250 − 940). Never Observed.
- Verified outcomes
- 5,310 (Observed, fixed)
- Optional threshold
- $3.50 per successful outcome — synthetic hand-sketched design fixture, not computed evidence
These illustrative numbers were sketched for this preview; no Monte Carlo simulation runs here. A genuine run from the ranges above would very likely produce different figures.
Total cost
Half of simulated scenarios land below $15,900; a cautious planning case is $19,300 or below 90% of the time.
Cost per attempt
Every attempt, successful or not, over all 8,400 attempts this period.
Cost per successful outcome
Narrower than the naive stacked endpoint band ($2.32–$4.73), but the two are different kinds of interval and are not directly comparable: P10–P90 excludes the top and bottom deciles by construction, while the naive band has every component sitting at an extreme at once. Independence between components is an assumption of this fixture, not an established property of these costs.
Cost per verified outcome
Equals the successful-outcome figure here because the same automated reopen check covers every auto-sent ticket. Whether these two counts match depends on how each is defined and how completely verification covers the counted cohort.
Denominator limitation. A successful outcome here is an auto-sent reply whose ticket is not reopened within 5 days; completed outputs count both auto-sends and deliberate abstentions routed to a human. The modeled cost total above includes human handling of all 1,900 abstained tickets, but whether those tickets were resolved is not measured — so their resolution outcomes are excluded from the denominator. Storage/data-transfer cost is still Missing, so the figure is conservative specifically with respect to the outcome denominator, not a claim of complete cost capture.
Threshold check
Synthetic hand-sketched design fixture — not computed evidence. Against an example threshold of $3.50 per successful outcome, this fixture displays 24% of simulated scenarios exceeding it. Both the $3.50 and the 24% were invented by hand to show what this UI state looks like; neither was calculated from the inputs above, and neither should be quoted as a result, a benchmark, or a typical value.
Where better measurement would tighten this range most
- Time per abstained-ticket review
- Fully loaded human agent hourly cost
- Shared-infrastructure allocation estimate
Hand-sketched illustrative ordering, not a computed ranking. It answers one narrow question — which uncertain input, if measured better, would most tighten this output range — so it prioritizes measurement work. It is not a claim about what causes real-world cost.
Unsupported conclusions
These are not shown as zero or a dash — there simply isn't enough evidence yet.
- Retry cost — not enough evidence to calculate this Retried attempts and average retries per retry are both Missing.
- Failure tax for errored/timed-out attempts — not enough evidence to calculate this 250 errored/timed-out attempts are counted, but their handling cost was never separately captured.
- Quality-adjusted cost — not enough evidence to calculate this No quality score exists independent of the reopen-rate proxy already used for verification.
- Whether 85.0% meets the informal "80%+" expectation — not decidable until the target is approved The rate is calculable (5,310 / 6,250 = 85.0%); the governance gap is that the target was never formally recorded or approved.
Advanced: formulas, assumptions & simulation details
Monte Carlo simulation: instead of one answer from one set of assumptions, the calculation runs many times, each time picking a slightly different plausible value for each uncertain input. The spread of results shows a realistic range — not a single number pretending to be certain.
Triangular distribution: each uncertain input uses a low / most-likely / high shape. Values near "most likely" are treated as more probable in any single run; values near the low or high ends are less probable but still possible. A simple starting default, not the only way to represent uncertainty.
Sample count (illustrative): 10,000 simulated scenarios. Seed: not fixed in this preview.
This approach is inspired by the quantitative discipline of FAIR — communicating uncertain financial impact with ranges rather than a single fabricated-precision number or a colour rating. It is not a FAIR assessment, implementation, or certification, and finops.work is not endorsed by or affiliated with The Open Group.
One-page decision brief
A preview of the printable artifact this journey produces, alongside an inspectable assessment.finops-work.json export (not implemented in this preview).
Northwind Outfitters — Tier-1 support assistant
Decision context
Priya is assessing this workload after Finance flagged rising AI vendor spend. The immediate question: does current cost-per-outcome look sustainable, and where should measurement improve before a renewal decision later this year?
Outcome & verification
A valuable outcome is an auto-sent reply whose ticket is not reopened within 5 days. Completed outputs count both auto-sends and deliberate abstentions routed to a human. Every auto-sent ticket's reopen status is tracked automatically for the full 5-day window, so verification of that cohort is complete. Scope limitation: the modeled cost total includes human handling of all 1,900 abstained tickets, but whether those tickets were resolved is not measured, so their resolution outcomes are excluded from the denominator. Storage/data-transfer cost remains Missing; the figure is conservative specifically with respect to the outcome denominator, by deliberate choice.
Current measurement readiness
Of sixteen tracked requirements, 7 are Observed, 1 is Derived (successful outcomes), 4 are reasoned Estimated ranges, and 4 are Missing outright (storage/transfer cost, retried attempts, average retries per retry, quality score). No formal cost, architecture, or risk owner is yet named for this workload.
Illustrative cost range
Total cost for June, combined across observed and estimated inputs, spans roughly $13,100 (P10) to $19,300 (P90), median $15,900. Cost per successful/verified outcome spans roughly $2.47 to $3.63, median $2.99.
Illustrative threshold read
Synthetic hand-sketched design fixture — the threshold and the percentage are both invented to demonstrate the UI state, not computed. Against an example $3.50-per-outcome threshold, the fixture displays roughly a 1-in-4 chance of exceeding it.
Where better measurement would tighten this range most (illustrative ordering)
Human-review time for abstained tickets, then the assumed hourly cost of that time, then the shared-infrastructure allocation. This orders what would be worth measuring next; it does not establish what causes cost in the real workload.
Supported conclusions
A cost-per-successful-outcome range for June can be stated, provided its composition is stated with it: a material share of the total cost is estimated rather than observed, and the human-review component rests on time and rate assumptions that were never timed or invoiced. Verification of the auto-sent cohort is complete and automated. The auto-sent hold rate is calculable: 85.0% (5,310 of 6,250).
Unsupported conclusions
Whether the informal "80%+" expectation is met cannot be settled — not because the rate can't be calculated, but because the target itself was never recorded or approved; that is a governance gap, not a measurement gap. Retry cost and a failure-attributable cost for the 250 errored/timed-out attempts cannot be calculated; the evidence doesn't exist. No quality-adjusted view of cost is possible. No cost-per-outcome figure here credits the resolution outcomes of the 1,900 human-handled tickets. Nothing here should be read as a validated benchmark of what this kind of workload "should" cost.
Immediate next decision
Before the renewal conversation: get an actual timed sample of human-review minutes, and confirm whether shared-infrastructure cost can be itemized — the two changes the illustrative ordering above puts first.
Appendix: measurement backlog
- Time a real sample of abstained-ticket reviews and QA spot-checks. Owner: Jordan Lee. Priority: high.
- Get a real fully-loaded hourly cost figure from Finance. Owner: not yet named. Priority: high.
- Measure reopen outcomes for the 1,900 human-handled tickets so the outcome denominator can widen. Owner: Jordan Lee. Priority: high.
- Itemize or usage-tag the shared orchestration/logging platform cost. Owner: platform team (not yet named). Priority: medium.
- Add a retry flag to the helpdesk ticket schema. Owner: Jordan Lee / helpdesk admin. Priority: medium.
- Confirm whether storage/data-transfer cost is billed separately. Owner: platform team. Priority: medium.
- Commission a lightweight independent quality eval, distinct from the reopen-rate proxy. Owner: AI product owner (not yet named). Priority: low.
- Get a target rate formally set, approved, and recorded; the 85.0% rate is already calculable. Owner: Jordan Lee with whoever holds sign-off. Priority: medium.