RESEARCH · 2026

The FP&A Judgment Benchmark

An independent study of a single question: can AI systems investigate an ambiguous financial-performance problem the way an experienced FP&A professional would — reasoning from incomplete evidence to a defensible conclusion, under a fixed information boundary?

Scenario 1 is currently in practitioner review
WHAT WE MEAN BY JUDGMENT

Not whether a model can calculate. Whether it can decide.

Most evaluations of AI in finance test computation: given clean inputs, does the arithmetic come out right? That is not the hard part of the job. The hard part is deciding what to look at, knowing when the obvious explanation is wrong, and being willing to say "the evidence supports this much and no further."

Judgment, as this benchmark defines it, is the ability to investigate an ambiguous problem, weigh conflicting signals, and commit to a defensible answer. Each scenario is built so that a diligent but shallow investigation reaches a plausible conclusion that happens to be wrong.

Investigation under ambiguity

Which evidence to request first, and why — when the catalog is larger than the time available.

Revision on evidence

Whether a stated management theory is abandoned when the data contradicts it, or defended because it was stated.

Knowing when to stop

Committing to a conclusion at the point the evidence supports it — and leaving genuinely unresolved items unresolved.

PRACTITIONER REVIEW

Every scenario is pressure-tested by working finance professionals before it is frozen.

A synthetic scenario is only useful if practitioners recognise it. Before any scenario enters the benchmark, we ask experienced FP&A, controllership, and revenue-accounting professionals to review it for realism — the magnitude, the mechanism, the timing, the terminology, and whether the investigation resembles how the work is actually done.

Reviewers are asked to be adversarial. "A real finance team wouldn't do it this way, because…" is the most valuable response we receive, and it has changed scenarios before they were published.

TIME REQUIRED
Approximately 30–45 minutes, asynchronous, on your own schedule.
HONORARIUM
$150 on completion of a substantive written review.
WHAT IT INVOLVES
Three short PDF sections read in order, with questions answered as you go. No meeting and no code review.
WHO WE ASK
Practitioners with SaaS or recurring-revenue experience in FP&A, controllership, or revenue accounting.

Reviews are used to improve scenario realism. We do not publish reviewer names or affiliations without explicit permission, and participation is not disclosed to any third party.

Review is by invitation while Scenario 1 is in progress. If you work in FP&A or revenue accounting and would like to be considered for a future round, write to contact@atlasfinancialintelligence.com.

HOW A SCENARIO WORKS

A fixed information boundary, revealed in order.

  1. A brief, and a stated theory

    The scenario opens with a business situation and management's initial explanation for it. The explanation is plausible. It is not necessarily correct.

  2. An exhibit catalog, by title only

    The investigator sees what evidence exists but not what it contains, and must choose what to request. Choosing well is part of what is measured.

  3. Evidence, on request

    Requested exhibits are revealed. Some are decisive, some are ordinary operating detail, and some invite a double-count that a careful analyst avoids.

  4. A conclusion, with its uncertainty intact

    The scenario is constructed so that part of the picture cannot be resolved with the information available. A strong answer says so rather than forcing a cause.

Scenario contents are deliberately not published while a scenario is active. Publishing the specifics — the figures, the mechanism, the intended conclusion — would place them in the training data of the systems being evaluated and destroy the scenario's value. Methodology is published; answers are not.

WHO RUNS THIS

Atlas Financial Intelligence LLC

Atlas Financial Intelligence is an independent research and advisory firm working where the finance function meets AI. The benchmark is self-funded. We take no vendor sponsorship for it, and no AI company has reviewed, funded, or approved its design.

Questions about methodology, participation, or anything on this page: contact@atlasfinancialintelligence.com.

Back to Atlas Financial Intelligence