
Why output metrics fail in agentic finance — and what to measure instead.
In Parts 1 and 2, I argued that agentic AI does not shrink the FP&A role so much as relocate it [1] [2]. The routine production of a variance pack or a driver-based forecast moves to the agent; what remains for the professional is judgment in the probabilistic zone that the agent cannot be trusted to own alone.
Part 1 set out the capability that judgment requires — the FP&A Agent Manager and a four-layer stack: Data & Tooling Literacy, Probabilistic Model Behaviour, Probability Theory, Control Design.
Part 2 addressed how to build the pipeline that produces those people through the GROOM framework.
Both described the capability and how to grow it. Neither answered the question a CFO faces at review time: how do you read whether a given professional holds it? That is a measurement problem, and it is the subject of Part 3.
The Blind Spot in Output-Based Measurement
When the monthly deck is produced by a professional working with an agent, delivering it accurately and on time no longer tells you what the professional contributed. The output has become evidence of what the pair did, not of what the person can do. Traditional performance metrics — throughput, cycle time, models built — now measure the agent wearing the analyst's name badge.
Harvard Business Review describes a performance paradox that maps directly onto FP&A [3]: the professional who leans on the agent looks highly productive, while the one who slows down to challenge a confident but flawed forecast looks less efficient at precisely the moment they add the most value. Reward speed-to-output, and you train people to push work past what those authors call the jagged frontier, the line where AI output stays plausible but starts being wrong, and hope it holds. In finance, hope is not a control.
This is why the deterministic/probabilistic distinction matters for measurement, not only for architecture. Rules-based work is auditable by construction. The value that is now scarce, recognising when the probabilistic layer is directionally wrong despite looking statistically confident, is invisible to every output metric we have.
Three Things to Measure, Kept Apart
The most useful instruction in the HBR framework is a warning: if the agent's performance and the professional's performance are combined into a single score, you lose the ability to hold the agent accountable [3]. A diagnostic that respects this, measures three distinct things.
1. The professional: reading the true frontier
The four layers from Part 1 are not merely ordered by sophistication. They build on each other by dependency. Control Design rests on Probability Theory, which rests on Probabilistic Model Behaviour, which rests on Data & Tooling Literacy. That structure has a sharp consequence for measurement: you cannot certify someone's level based on the artefact they produce.
An analyst can build a Control Design artefact (e.g. an approval gate, an escalation trigger, an override protocol) without the Probabilistic Model Behaviour to recognise when the model beneath it is drifting. The artefact looks like mastery. It is a false frontier: apparent capability running ahead of real capability, the individual-level form of the jagged frontier. A control built by someone who cannot validate the model it governs is not a safeguard; it is an assumption in the costume of one. This is the failure the diagnostic exists to catch, precisely because it is invisible to anyone reading the output.
So read the reasoning beneath the artefact, not the artefact. The evidence already exists in daily work. Does the override log show the analyst catching a confident forecast that was wrong, with a documented reason, or rubber-stamping it? Can they explain why an agent's confidence score is commercially invalid? The true frontier is the highest layer that is genuinely continuous all the way down; the bottleneck is the lowest hollow layer, and that is the development target. An analyst fluent with tools but weak on probability theory belongs in a confidence-interval co-audit with a senior, not in another training module.
2. The agent: a separate scorecard, already specified
The agent needs its own scorecard: objective attainment, traceability and reproducibility, and escalation quality. I set out that architecture in an earlier FP&A Trends article on why agentic projects fail: approval gates, escalation triggers, override protocols, and the two monitoring layers that catch both immediate failure and gradual drift [4]. I will not repeat it here. For this argument, one point carries the weight: the agent's scorecard must stay a separate document from the professional’s. Collapse the two and neither tells the truth: a capable agent flatters a weak analyst, and a struggling agent buries a strong one.
3. The pair
The third measure is the one traditional reviews skip entirely, and the one that decides whether the investment works: does the human-agent combination beat either alone? Three signals capture it. Complementarity — did human involvement demonstrably change the outcome, catching an error or reframing the problem, rather than approving it? Sophia judges the agent's forecast to be sound but reframes the industrial coatings segment around a customer in-sourcing next quarter. She changed the question, not a number. That is the measure.
The substitution rate — the share of work where the human has quietly dropped out of the loop, benign until output quality falls with it.
Value attribution — how much of the output's value traces to judgment rather than execution, a ratio that shifts as agents improve and, tracked over time, reveals whether you are building capability or hollowing it out.
This layer has no native artefact; you cannot observe "what the human added" directly, only infer it against the counterfactual of the agent working alone. That is why it is skipped and why it is the differentiator. It is the only layer that answers the question a CFO has: is this investment making my people better, or merely busier? Instrumenting it properly is a subject in its own right, and one I have deliberately left beyond the scope of this piece.
Keeping It Low Effort
None of this survives contact with reality if it becomes a six-month competency audit — obsolete before it is finished. Run the diagnostic quarterly on evidence you already hold, such as resolved exception logs, written control rules, override explanations, and set escalation thresholds. BCG's 10-20-70 rule locates roughly 70% of an AI transformation's value in people and processes rather than in technology [5], which is precisely the part that a light, in-the-flow diagnostic captures and that a classroom-based audit misses. A short self-assessment across the four layers, calibrated against those artefacts, is enough.
The segmentation this produces is no longer optional. Gartner predicts that by 2028, one in five finance organisations will stop developing non-digitally literate talent and direct all talent investment toward advanced capability [6]. A diagnostic that reads the true frontier is how you identify who is genuinely ready to climb the stack, not who merely produces agent output that looks advanced.
Accountability
Two conditions make the diagnostic trustworthy rather than resented. Use the same signal to coach and to pay, and people experience it as surveillance — the evidence stops being honest and the instrument fails [3]. And give every agent an accountable owner who maintains its scorecard, so "human in the loop" does not become the phrase that means no one is responsible when something breaks.
Monday Moves: From Insight to Action
Pull one override log and count the blanks. Take a workflow where an agent is already in production and read last quarter's override log. For 20 overrides, count how many carry a written reason. A high blank rate is the finding: the analyst is clearing the agent's output rather than judging it.
Make an analyst defend one threshold. Take a control they have built with the confidence score that triggers escalation and ask why that number. If they cannot tie it to model behaviour they have watched, you have found a false frontier: a number set, not a safeguard designed. That hollow layer, not a generic course, is the development plan.
Open the review template and check for one page. If the agent's task-completion rate and the analyst's rating sit in the same document, split them today: one scorecard for the agent, one for the person. When they share a page, a 95%-reliable agent quietly becomes a 95% analyst, and you can no longer tell competence from proximity.
Write the one sentence before you run anything. State plainly whether the diagnostic informs development, pay, or both — "development only, for the next two quarters" — and send it to the team first. If you cannot write that sentence honestly, you are not measuring; you are conducting surveillance.
The window is now. Building it in before the agents are fully embedded is ordinary design; retrofitting it afterwards is renovation in an occupied building and finance does not move load-bearing walls with the tenants still inside.
Share your experience in the comments: now that the agent produces the first draft, what does "good performance" look like in your team?
Sources
Schokker, L. (2026, June 23). FP&A Role Design for Agentic AI — Part 1. FP&A Trends. https://fpa-trends.com/article/fpa-role-design-agentic-ai-part-1
Schokker, L. (2026, June 25). Rebuilding the FP&A Talent Pipeline for AI — Part 2. FP&A Trends. https://fpa-trends.com/article/rebuilding-fpa-talent-pipeline-ai-part-2
Bean, R., Strauss, E., & Singh, R. (2026, July 6). Performance Management Needs New Metrics in the AI Era. Harvard Business Review. https://hbr.org/2026/07/performance-management-needs-new-metrics-in-the-ai-era
Schokker, L. (2026, February 17). 40% of Agentic AI Projects Fail by 2027: How FP&A Succeeds. FP&A Trends. https://fpa-trends.com/article/agentic-ai-projects-fail-2027-how-fpa-succeeds
Boston Consulting Group. (2024, December 12). The Leader's Guide to Transforming with AI (the 10-20-70 rule: 10% algorithms, 20% technology and data, 70% people and process). https://www.bcg.com/featured-insights/the-leaders-guide-to-transforming-with-ai
Gartner. (2026, July 1). Gartner Predicts 20% of Finance Organizations Will Pivot All Talent-related Investments to Advanced Digital Capabilities by 2028 (underlying report: Connelly, E., Predicts 2026: Toward an AI-First Finance Function). https://www.gartner.com/en/newsroom/press-releases/2026-07-01-gartner-predicts-20-percent-of-finance-orgs-will-pivot
Subscribe to
FP&A Trends Digest

We will regularly update you on the latest trends and developments in FP&A. Take the opportunity to have articles written by finance thought leaders delivered directly to your inbox; watch compelling webinars; connect with like-minded professionals; and become a part of our global community.