Baseline
The intentionally simple floor: configured manual hours × role rate × run count. Cheap, deterministic, always present — the sanity check the other dimensions are reconciled against.
Platform · ROI Hub
The ROI Hub is the Observatory's financial-justification surface. It measures what your production agents actually return — over the same run telemetry the platform already captures — and grades every reported dollar as Verified, Modeled or Assumed, so the number you take to the CFO carries its own audit trail.
verified · modeled · assumed — one tier per dollar
The agent-ROI crisis is not a value crisis.It is a measurement-credibility crisis.Credibility is an engineering property of the measurement system — not a rhetorical one.
95% of generative-AI pilots show no measurable P&L impact — not because agents don't create value, but because flat assumptions and self-reported time savings are figures a finance function correctly declines to capitalize. The fix is the one financial accounting found a century ago: don't make every number certain — grade it, audit it, and recognize it under the right tier.
VERA · Verifiable, Evidence-Graded Return on AgentsThe dimensions aren't alternatives — they're layers over one stream of agent-run telemetry, each visible, each feeding the confidence engine that decides how much value may be credited, and at what tier.
The intentionally simple floor: configured manual hours × role rate × run count. Cheap, deterministic, always present — the sanity check the other dimensions are reconciled against.
Answers, for every run: how many human-hours would this specific output have taken? A usefulness gate admits the run, an estimator prices it, and calibration against your own labelled history keeps it conservative and unbiased in aggregate.
A ledger of value that actually occurred: outcome contracts define what counts as a realized event, records link back to the originating run, and causal experiments (randomized holdouts) validate the effect. Only this dimension can produce Verified credit.
Borrowed from revenue recognition: a dollar backed by a randomized holdout is not the same asset as a dollar inferred from a model assumption — so each unit of value is counted under exactly one tier.
Value validated by a holdout experiment or a verified outcome record — and credited conservatively, at the lower bound of its confidence interval.
Per-run counterfactual estimates, corrected against your organization’s own labelled samples. Noisy per run, trustworthy in aggregate — and it never silently inflates.
The flat baseline every ROI conversation starts from today. VERA keeps it visible — as the anchor the higher grades are reconciled against, not as the headline.
Read the research: VERA — Evidence-Graded ROI for Production Agents →
Net ROI and payback per agent feed a standing review: SCALE what earns, WATCH what’s unclear, FIX what leaks, KILL what doesn’t pay.
An acceptance signal and a per-operation inclusion policy keep the measurement honest — an agent can’t farm credit for output nobody keeps.
Realized outcomes feed back to re-calibrate the estimator, so the modeled numbers converge on the verified truth instead of drifting from it.
Bring one agent already in production — or one you're sizing. We'll show its return measured, graded and reconciled on the ROI Hub, evidence attached.