Technical paper · VERA
Evidence-Graded ROI for Production Agents.
Bringing counterfactual measurement, causal validation, and confidence-graded accounting to the business case for enterprise AI agents.
- agent ROI
- counterfactual estimation
- causal validation
- confidence grading
- productivity measurement
- value outcome units
- AI governance
- FinOps for AI

Core thesis
The agent-ROI crisis is not a value crisis.It is a measurement-credibility crisis, and credibility is an engineering property of the measurement system, not a rhetorical one.
Abstract
Enterprises are deploying AI agents faster than they can prove the agents pay for themselves. Industry research reports that 95% of generative-AI pilots produce no measurable profit-and-loss impact, and analysts project that more than 40% of agentic-AI projects will be cancelled by 2027, primarily for unclear business value. We argue that this is not a value crisis but a measurement-credibility crisis: prevailing ROI practice rests on flat productivity assumptions and self-reported time savings that finance functions correctly decline to capitalize.
We present VERA (Verifiable, Evidence-Graded Return on Agents), a measurement framework that embeds the causal-inference gold standard (counterfactual estimation and randomized holdout validation) directly into the agent observability layer, and grades every reported dollar by the strength of its evidence.
VERA computes return across three reconciled dimensions: a deterministic per-agent baseline, a per-execution counterfactual estimator calibrated against an organization s own labelled history, and a realized-outcome ledger validated by causal experiments. Each figure is then classified as Verified, Modeled, or Assumed (an evidentiary grade analogous to revenue-recognition tiers in financial accounting), with verified value credited conservatively at the lower bound of its confidence interval.
We describe the architecture, the estimator and calibration mechanics, the confidence-grading and causal-validation protocols, and the governance surface that turns graded ROI into portfolio decisions. We illustrate the framework on a regulated-insurance deployment and discuss how evidence-graded accounting changes the conversation between AI teams and the CFO.
Part of the FlowX.AIpaper series.
Each paper names a framework and shows it running in production — governance, reliability, memory, and measurement, engineered rather than hoped for.