Platform · ROI Hub

Every dollar, gradedby its evidence.

The ROI Hub is the Observatory's financial-justification surface. It measures what your production agents actually return — over the same run telemetry the platform already captures — and grades every reported dollar as Verified, Modeled or Assumed, so the number you take to the CFO carries its own audit trail.

verified · modeled · assumed — one tier per dollar

The measurement thesis
The agent-ROI crisis is not a value crisis.It is a measurement-credibility crisis.Credibility is an engineering property of the measurement system — not a rhetorical one.

95% of generative-AI pilots show no measurable P&L impact — not because agents don't create value, but because flat assumptions and self-reported time savings are figures a finance function correctly declines to capitalize. The fix is the one financial accounting found a century ago: don't make every number certain — grade it, audit it, and recognize it under the right tier.

VERA · Verifiable, Evidence-Graded Return on Agents
Three dimensions

Three measurements, one reconciled number.

The dimensions aren't alternatives — they're layers over one stream of agent-run telemetry, each visible, each feeding the confidence engine that decides how much value may be credited, and at what tier.

D1

Baseline

The intentionally simple floor: configured manual hours × role rate × run count. Cheap, deterministic, always present — the sanity check the other dimensions are reconciled against.

D2

Counterfactual estimator

Answers, for every run: how many human-hours would this specific output have taken? A usefulness gate admits the run, an estimator prices it, and calibration against your own labelled history keeps it conservative and unbiased in aggregate.

D3

Realized outcomes

A ledger of value that actually occurred: outcome contracts define what counts as a realized event, records link back to the originating run, and causal experiments (randomized holdouts) validate the effect. Only this dimension can produce Verified credit.

Confidence grading

Verified, Modeled, Assumed.

Borrowed from revenue recognition: a dollar backed by a randomized holdout is not the same asset as a dollar inferred from a model assumption — so each unit of value is counted under exactly one tier.

Verified

Backed by causal evidence

Value validated by a holdout experiment or a verified outcome record — and credited conservatively, at the lower bound of its confidence interval.

Modeled

Estimated, then calibrated

Per-run counterfactual estimates, corrected against your organization’s own labelled samples. Noisy per run, trustworthy in aggregate — and it never silently inflates.

Assumed

The declared floor

The flat baseline every ROI conversation starts from today. VERA keeps it visible — as the anchor the higher grades are reconciled against, not as the headline.

Read the research: VERA — Evidence-Graded ROI for Production Agents →

From number to decision

Graded ROI becomes portfolio decisions.

Portfolio decisions

Net ROI and payback per agent feed a standing review: SCALE what earns, WATCH what’s unclear, FIX what leaks, KILL what doesn’t pay.

Anti-gaming by design

An acceptance signal and a per-operation inclusion policy keep the measurement honest — an agent can’t farm credit for output nobody keeps.

Self-correcting

Realized outcomes feed back to re-calibrate the estimator, so the modeled numbers converge on the verified truth instead of drifting from it.

Next

Take a number theCFO will accept.

Bring one agent already in production — or one you're sizing. We'll show its return measured, graded and reconciled on the ROI Hub, evidence attached.