Investigation under ambiguity
Which evidence to request first, and why — when the catalog is larger than the time available.
An independent study of a single question: can AI systems investigate an ambiguous financial-performance problem the way an experienced FP&A professional would — reasoning from incomplete evidence to a defensible conclusion, under a fixed information boundary?
Most evaluations of AI in finance test computation: given clean inputs, does the arithmetic come out right? That is not the hard part of the job. The hard part is deciding what to look at, knowing when the obvious explanation is wrong, and being willing to say "the evidence supports this much and no further."
Judgment, as this benchmark defines it, is the ability to investigate an ambiguous problem, weigh conflicting signals, and commit to a defensible answer. Each scenario is built so that a diligent but shallow investigation reaches a plausible conclusion that happens to be wrong.
Which evidence to request first, and why — when the catalog is larger than the time available.
Whether a stated management theory is abandoned when the data contradicts it, or defended because it was stated.
Committing to a conclusion at the point the evidence supports it — and leaving genuinely unresolved items unresolved.
A synthetic scenario is only useful if practitioners recognise it. Before any scenario enters the benchmark, we ask experienced FP&A, controllership, and revenue-accounting professionals to review it for realism — the magnitude, the mechanism, the timing, the terminology, and whether the investigation resembles how the work is actually done.
Reviewers are asked to be adversarial. "A real finance team wouldn't do it this way, because…" is the most valuable response we receive, and it has changed scenarios before they were published.
Reviews are used to improve scenario realism. We do not publish reviewer names or affiliations without explicit permission, and participation is not disclosed to any third party.
Review is by invitation while Scenario 1 is in progress. If you work in FP&A or revenue accounting and would like to be considered for a future round, write to contact@atlasfinancialintelligence.com.
The scenario opens with a business situation and management's initial explanation for it. The explanation is plausible. It is not necessarily correct.
The investigator sees what evidence exists but not what it contains, and must choose what to request. Choosing well is part of what is measured.
Requested exhibits are revealed. Some are decisive, some are ordinary operating detail, and some invite a double-count that a careful analyst avoids.
The scenario is constructed so that part of the picture cannot be resolved with the information available. A strong answer says so rather than forcing a cause.
Scenario contents are deliberately not published while a scenario is active. Publishing the specifics — the figures, the mechanism, the intended conclusion — would place them in the training data of the systems being evaluated and destroy the scenario's value. Methodology is published; answers are not.
Atlas Financial Intelligence is an independent research and advisory firm working where the finance function meets AI. The benchmark is self-funded. We take no vendor sponsorship for it, and no AI company has reviewed, funded, or approved its design.
Questions about methodology, participation, or anything on this page: contact@atlasfinancialintelligence.com.