Headless FP&A shows how AI, governed data architecture, and human oversight can turn business intent into...

LLMs can generate the work without producing the understanding behind it. That changes what sign-off means. LLMs can generate outputs without the underlying understanding. This shifts the meaning of sign-off.
Trust Is What Lets The Work Travel
No leader can personally verify every number, reconstruct every forecast, and reperform every analysis. There isn’t time, and that’s not their job. FP&A works because leaders can act on the analysis they haven’t done themselves. Trust is what makes that possible, and what makes it valuable to the whole organisation. And it enables faster decisions and less duplication because good analysis doesn’t have to be rebuilt at every level.
Trust isn’t just a soft skill or a nebulous concept. Trust is what lets real work have a bigger impact. When trust is weak, another team (or team member) is asked to review and validate the same numbers; a forecast is escalated because no one trusts it; parts are quietly rewritten to compensate; or the output stops being useful for decision-making and becomes a compliance checkbox. Low trust doesn’t remove the work. It multiplies it.
The core job of FP&A isn’t the close, the forecast, or the variance analysis. Those are tasks. The core job is judgment and decision support, helping the business see what’s really happening and make better decisions. That requires trust, and that’s the part I’ve been thinking about as Large language models (LLMs) are set to become an increasingly bigger part of FP&A work. The impact this will have isn’t obvious and I haven’t seen it much in the current AI-in-finance conversation.
What an Output Used to Prove
Take a common FP&A task, such as a variance commentary. Writing a strong one meant pulling the actuals, comparing them to the forecast, asking why the number moved, investigating different data sources, talking to team members with insight, and challenging the explanation. The commentary that landed was a visible signal that someone had done that work.
LLMs can generate commentary more quickly and often in a cleaner format than a typical human first draft. However, what they do not necessarily produce is the investigation and judgment that would normally accompany a variance comment. The resulting paragraph may now appear identical regardless of whether someone undertook the more rigorous investigative work. As a result, the connection between the output and the underlying understanding has become weaker.
Automation had already weakened this connection in parts of finance work. Analysts could run automated reports and recurring packs without deeply understanding or reviewing them. But LLMs push that separation into areas like explanations, recommendations, and narratives. These used to be strong evidence of interpretation and judgment being applied.
Trust Isn’t One Thing
I’ve noticed that many discussions about AI in finance treat trust as a single thing: accuracy. I believe it is, in fact, more complicated than that, and that there are three different aspects to it.
Output trust asks whether it is correct.
Process trust asks whether it was produced the right way.
Judgment trust asks whether the individual who produced the work understands it well enough to be able to defend it.
They differ in practice. A leader may agree that a figure is correct even if they doubt the explanation. Judgment trust refers to the belief that the person will be able to explain the drivers, identify the main assumptions, tell the difference between evidence and inference, and also state how confident they are and what is responsible for that confidence.
Output and process trust get most of the attention at the moment because they are more tangible. Hallucinations, formula errors, traceability, controls. Judgment trust is in the background because it was never challenged in this way by tools before. A leader can act on the analysis because they believe the person understands the business well enough to be insightful and honest enough to say when they’re not sure.
The Work Was Building More Than the Output
Some parts of FP&A work can look inefficient. But some of that is where judgment got built. Not because manual work guarantees understanding. Analysts can copy commentary forward or accept a surface explanation without considering root causes. Working through variance commentary or a business case builds an understanding of the business, how the financials flow, and effective methods of communicating.
The old workflows had the developmental parts embedded by default. Writing variance commentary properly required the analyst to investigate the drivers, and created the repeated exposure that made next month’s variances easier to recognise and understand. Building a business case from scratch surfaced assumptions and complexities, deepening the analyst’s understanding.
Doing the work built understanding. Understanding made reviews more meaningful. And the reviews increased confidence in approvals. This cycle occurred every time an analyst actually worked on solving the problem. Although the workflow was producing variance commentary, it was also producing better analysts.
LLMs can take over the production, and if they take the reasoning with it, the developmental parts go too. Once production is automated, that development has to be done separately, or it doesn’t happen. Judgment can still come from decision reviews, postmortems, or seeing your own recommendations play out. But the old workflow supplied it as a byproduct of the process, and the new one won’t.
Why Review Alone Doesn’t Close the Gap
The intuitive fix is that the time saved in production can be devoted to more careful review, which is harder than it sounds. Independent review is sometimes better precisely because the reviewer isn’t anchored to the original reasoning. But the problem is reconstructability: review becomes ineffective when the reviewer cannot access the source evidence, the underlying assumptions, the intermediate steps, or the points of uncertainty in the analysis. A reviewer can’t usually do much with a smooth final paragraph that leaves no trail behind.
Humans produce polished nonsense all the time. That isn’t unique to LLMs. The difference is how the errors show up. Human work tends to carry visible signs of trouble: messy logic, obvious gaps, a bridge between two assumptions that doesn’t quite hold. Fluency and coherence used to be at least weak signals that the work underneath was sound. LLMs produce both cheaply, so weak reasoning can now arrive looking finished. What still catches it is noticing that a driver doesn’t fit operating reality, or that a conclusion is more certain than it should be. That comes from knowing the business, not from spending more time on the page, which is why a human in the loop only helps when that person has the understanding or the trail to challenge the work meaningfully.
Two Risks, Not One
Two problems get blurred together, and they unfold on different timelines. The first is a current accountability risk: the person signing off can’t genuinely reconstruct or defend today’s analysis. That’s visible now, if anyone checks.
The second is a future capability risk that compounds quietly. If junior analysts get less practice doing investigative work where judgment forms, the team can keep producing acceptable output today while slowly losing the ability to review it tomorrow. And the bill comes later, when the people who were supposed to catch the subtle errors can’t, because they never learned what wrong feels like.
Designing for Trust Instead of Inheriting It
Trust used to be part of the work. In an LLM world, it has to be built into how the work gets done.
The first is probably already happening in some form on most teams: match review depth to the consequences, not the format. Forecast narratives and business cases need review to reach the sources and assumptions. Board materials and major investment calls need the accountable owner to reconstruct the logic and say what evidence would change the conclusion. A first-draft summary isn’t low-consequence just because it’s a first draft.
The second is to require reconstructability for material work. If a conclusion matters, someone should be able to trace how it was reached, not just confirm that it reads well.
When I’ve looked at what it actually takes to get an LLM to produce decent variance commentary, most of the effort goes into the inputs rather than the prompt. Actuals at the right grain. Which forecast version you’re comparing against, and what assumptions are behind it. And any other key context outside the financial system that can be captured, like a contract overview showing a renewal everyone knows slipped. If those aren’t captured and structured, the LLM doesn’t tell you they’re missing. It produces a plausible driver anyway. So even though the model itself is a black box, you can be certain about what went into it, and check the reasoning against that when you need to. And if the workflow also lists what it couldn’t find, the reviewer knows where the commentary is standing on thin ground.
The third is to preserve the activities that build judgment. Investigating drivers, testing explanations, challenging assumptions, and seeing decisions play out. Those build the skill and knowledge. If the tool does all of them, decay is likely to follow.
This does not imply that you should keep the tools at arm’s length. If used properly, LLMs have the potential to enhance judgment rather than replace it, for example, by producing alternative explanations for a variance or by considering what would have to be true in order for the recommendation to be wrong. The problem is not that the LLM takes part in the work. It is that it may remove the parts that build understanding and leave only the numbers and words behind.
LLMs don’t automatically erode trust in FP&A. They erode it when teams let a credible-looking output replace their own understanding, thereby weakening their accountability. Since trust is the prerequisite for FP&A doing its job, that’s not something to leave to chance. It’s more of a workflow design question than a technology question, and it’s the part of the AI conversation we need more of.
Subscribe to
FP&A Trends Digest

We will regularly update you on the latest trends and developments in FP&A. Take the opportunity to have articles written by finance thought leaders delivered directly to your inbox; watch compelling webinars; connect with like-minded professionals; and become a part of our global community.