Can we trust generative AI in Finance?

Here we will analyze the unnatural but relevant link—if placed under control—between a Finance function anchored in its need for truth and the arrival of generative AI* at its service, whose strength is precisely to invent narratives but also figures. An irreconcilable paradox?

Finance is also a story of explainability

Beyond the figure that allows a decision to be made, the explainability of that figure is also key. When announcing to a BU manager overseeing several billion that they are not meeting their objectives and that, consequently, corrective action will be taken, one must be able to argue, document, and be sure of the sources. This is the notion of explainability.

And we regularly come across management controllers who spend one hour producing an analysis and another hour checking the calculation method and sources. How many meetings have ended after 5 minutes because no one was aligned on the figures?

It is in the DNA of Finance, and it also allows for direct action without hesitating over the numbers.

Key takeaway:
A figure without a source or documented method is worthless to a BU manager.

Generative AI is first and foremost a story of creativity

Generative AI is a family of algorithms based on LLMs* that has the fascinating ability to generate natural language text that appears completely realistic. And for good reason: the major LLMs on the market have analyzed the entire global web and are capable, from a fragment of a sentence or a question, of predicting the most credible sequences of words relative to the context. These are not exactly the most relevant; they are the words that are statistically most probable according to its digital vision of the world. Therefore, it is not the most accurate result via logical reasoning that is produced, but the generation of the most probable.

Another important nuance: to give an even greater illusion of naturalness, the word sequences are not always those with the highest probability but are drawn at random from among the most probable. Why? To give an impression of naturalness. If you ask it the same question twice, it will phrase the answer in two different ways with small nuances. This moves even further away from a deterministic logic which, via documented deduction, leads to a stable result.

It is on this principle that ChatGPT, Claude, Mistral, or Gemini respond to you so fluidly, regardless of the context.

Key takeaway:
An LLM generates the most probable word, not the truest.

Generative AI knows neither how to count nor how to explain itself: and that is its main risk

An LLM alone has no computational logic. If it answers “2+2=” correctly, it is not because it calculated it, but because it has seen this answer millions of times in its training data. As soon as we move away from simple or highly documented calculations, an LLM alone can “hallucinate”: invent a result that appears real, without ever detecting the error, and sometimes build upon it with complete confidence.

This limitation is aggravated by a second problem: unexplainability. Even when an LLM is trained to break down its reasoning into more readable steps, these steps remain probabilities, never a logical and reproducible deduction that could be audited step-by-step. For your function, which thrives on traceability and argumentation before a BU manager, this is the number one point of vigilance.

The good news: this risk is identified, understood, and we already know how to put it under control; that is the whole subject of the rest of this article.

Key takeaway:
An LLM alone does not calculate and cannot justify its reasoning step-by-step.

How to reconcile AI hallucination and Finance?

We can see that on paper, the raw combination of Finance and generative AI seems antinomical and explosive. Yet there is enormous potential if two essential rules are respected:

  • Knowing its limits and strengths, only apply generative AI where it makes sense…
  • Put it under control and validate it…

Nothing very new, in short. Let’s detail this approach.

 

Key takeaway:
Only apply AI where it makes sense, and always validate it.

Generative AI is a bad calculator but an excellent analyst

In Finance, you don’t just work with numbers. The decisions you are asked to make are certainly based on figures, but also on contextual information which can be narrative, graphic, video… Market context, competitor behavior, the socio-economic climate, evolving standards, your own internal rules from one division to another… these are all formats that this type of algorithm is relevant for linking, analyzing, and synthesizing. Specifically trained for Finance, AI agents* are excellent analysts that can absorb and contextualize a mass of information that you could not process otherwise.

Key takeaway:
Generative AI excels at analyzing context, not at calculating.

Generative AI is the voice and the pilot of your calculator

While LLMs do not know how to calculate, they are the new translators and analysts between you and your complex tools.

Let’s take a concrete example: explaining an 8% variance on a cost line compared to the budget during a closing. Traditionally, a management controller manually crosses several sources—extractions from their EPM tool, comments sent back by subsidiaries, history from previous months—to write a reasoned variance comment. This work easily takes 30 to 45 minutes per significant line, repeated over dozens of lines each closing.

With a correctly configured AI agent, the mechanics change: the calculation tool (your EPM) produces the exact and verifiable variance: the deterministic data does not change hands. The AI agent then reads the subsidiary comments, cross-references them with the history, identifies recurring causes, and writes an initial structured draft of the comment. The management controller then spends their time no longer searching and writing, but verifying, challenging, and enriching, reducing the work to 10-15 minutes of quality control rather than 30-45 minutes of production.

The gain is therefore not in producing a different figure: the figure remains the same, produced by the same reliable calculation tool. But it frees up the management controller’s time from the most time-consuming and least analytical part of the work. Tools perform deterministic* calculations that can be 100% certain, and LLMs handle writing the context and transmitting the correct hypotheses to them. Each solution focuses on its strengths for a combination that multiplies the potential of your teams.

 

Key takeaway:
The figure remains the same; it is the context writing time that disappears.

Small specialized AI agents are easier to control

To reduce hallucinations in your future agents, a simple technique consists of breaking down the tasks assigned to them into sub-tasks, each assigned to a dedicated agent. This technique makes it easier to train these algorithms and reduces their hallucinations, as they are less exposed to contexts they do not know.

Thus, one could have a single agent to whom all questions are asked—documentary research, P&L analysis, consolidation of subsidiary comments, identification of causes, or production of reporting… But it would be extremely complex to train correctly and very sensitive to any change in scope or data.

We therefore favor a collection of agents that collaborate with each other. This often includes a “Pilot & Synthesis” agent that receives questions and redirects them to the right sub-agents (analyst, formatting, web search…) and synthesizes the results. Each is more specialized and easier to control and evolve.

Key takeaway:
Several specialized agents are better than a single generalist agent.

AI agents are very good at self-monitoring

In this galaxy of agents that will equip your department, many will be dedicated to monitoring and challenging the results of their peers. We know that LLMs hallucinate, but we can train other LLMs to detect these hallucinations, poor report layout, figures without real references… This is now a standard. Every LLM process has its integrated LLM safeguard.

 

Key takeaway:
Every AI agent should have its own AI safeguard.

Good governance and data quality improve the relevance of AI agents

This is the part no one wants to read. No, generative AI is not magic. If your data is chaotic, siloed, inconsistent from one source to another, and with business definitions of varying scope, then AI or no AI, you will produce hallucinations = figures you believe are real but have no reality. Adding an autonomous AI agent to these areas greatly increases the potential for hallucination.

Exploiting the potential of generative AI therefore implies up-to-standard governance and data quality. AI can help improve quality, but it cannot do it without you.

Key takeaway:
Without quality data, AI amplifies hallucinations rather than avoiding them.

Generative AI also raises the question of where your data goes

There is a risk we haven’t addressed yet that often causes concern faster than hallucination: what happens to data once it is entered into a generative AI tool? There is a fundamental difference between consumer use, where data can be kept and reused by the publisher, and secure enterprise use, where it is contractually protected and not reused for model training.

Pasting a P&L extract or a supplier negotiation into an unvalidated tool is equivalent to letting sensitive data leave your control perimeter—this is what is called “Shadow AI.” The rule is simple: the same confidentiality reflexes that apply to an email apply to an AI prompt. And this is not a regulatory vacuum: the European AI Act and your internal controls (SOX or equivalent) apply to the results produced by an AI exactly as they do to any other result.

Key takeaway:
An AI prompt deserves the same confidentiality reflexes as an email.

Always, always keep the human in the loop

Even with these precautions, risks of hallucination always remain. Yes, our analytical power will be multiplied, yes our time to action will be reduced, but let us never neglect the time for monitoring and validating results. It is in our nature, in our duties, and it should be natural. We remain responsible for the figures and analyses produced. Part of the evolution of our role will be to understand when and why they hallucinate and to enrich the safeguards and controls we apply to them.

Key takeaway:
AI multiplies your analytical power, but you remain responsible for the result.

In summary: There is therefore no paradox in mixing Finance and generative AI.

It is an extremely powerful new duo that does not just perform existing tasks faster, but opens up new analyses and projections that were previously humanly impossible. Its integration must be progressive, applying it where it makes sense, while controlling its potential deviations and, above all, never abdicating our role as those responsible for the result.

FAQ

No, not alone. An LLM can hallucinate on complex calculations. Figures must always come from a deterministic calculation tool (EPM, ERP…), with generative AI handling context and synthesis.

 

Two main risks: hallucination (an invented result that appears real) and data security (Shadow AI, use of tools not validated by the organization).

By breaking down tasks between specialized agents that are easier to train and control, adding dedicated quality control agents, and systematically maintaining human validation.

No. It automates research, cross-referencing, and first-level writing tasks, but the responsibility for the result and its validation remains human.

Generative AI: a family of algorithms based on LLMs, capable of generating text, images, or other content in natural language from an instruction.
LLM (Large Language Model): a language model trained on immense volumes of text, which predicts the most probable sequence of words in a given context.
AI Agent: a program built around one or more LLMs, specialized in a business task (analysis, research, synthesis, etc.) and capable of interacting with other tools.
Deterministic calculation: a calculation whose result is always identical for the same input data, as opposed to a probabilistic generation of the LLM type

Let’s discuss your challenges and our solutions

Let’s talk about your project and discover how MeltOne can turn your challenges into concrete, high-performing solutions.

Img Meltone 9