Skip to content
AnrilX

Evidence

The method, published before the numbers

Every vendor in this category asserts accuracy and almost none of them says how it was measured. That is why the figures get discounted, and rightly. This page is the method we intend to be held to, including the parts of it that are weak.

Two scores

Being right and being checkable are different things

A single blended accuracy number hides the trade between them, and the trade is the entire product. A correct figure whose source line does not resolve is not an answer you can take into a close.

Answer Score

What fraction of a controller-quality answer the system produces. Graded against a gold answer written by someone who could have produced it by hand, not against a string match.

Source Score

What fraction of the figures carry a Source Line that actually resolves: run the stated entity, field, aggregation, filter and period yourself and get the stated number back. A figure that is right with an unusable source line fails this and passes the other, which is exactly why they are separate.

Why not an existing benchmark

Four reasons the public sets do not transfer

These gaps are already documented for text-to-SQL benchmarks generally. Every one of them is worse against an SAP schema, which is why a good score on a public set would tell you nothing useful about your own landscape.

Question complexity

Public sets favour questions with one clear answer. The questions worth asking an ERP cross grains, modules and time windows, and the hard part is assembling the calculation rather than retrieving a value.

Schema complexity

Benchmark schemas are tens of tables with readable names. An SAP landscape is thousands, with names that are neither English nor stable across releases.

Query complexity

Reversals, line-item grain, currency and the difference between a document date and a posting date are not edge cases here. They are most of the work.

Semantic alignment

The largest gap. A query can be valid SQL, return rows, and still not be your definition of net revenue. Public benchmarks do not measure agreement with a business definition because they have none.

The harness

Five steps, and the fourth is the expensive one

Nothing here is novel. It is written down because almost nobody writes it down, and an accuracy figure without it is a number you have no way to argue with.

  1. 1

    A question set, published

    Written to look like the questions a controller actually asks: across grains, across modules, with a definition behind them.

  2. 2

    A gold answer per question

    What somebody who could have produced it by hand would have produced. More than one where two calculations are both defensible.

  3. 3

    Run, and score twice

    Answer Score: how much of that gold answer came back. Source Score: how many figures carry a Source Line that actually resolves.

  4. 4

    Graded by a person where a model cannot

    Where automated grading cannot separate a right answer from a plausible one, a human grades it, and we publish what fraction needed that.

  5. 5

    Published with its limitations

    Including the weakest dimension. A scorecard with no low number on it has not been measured.

Results

Not yet

There is no scorecard on this page, and there will not be one until the run is finished and the method has not been quietly adjusted to improve it. Publishing an empty section is the honest version of a page like this; the alternative is to wait and then present the method and the numbers together, which is how a method gets shaped by its results.

If you are evaluating now, ask and you will get the current state as it stands, including what is not working.

Read the caveats first

What this measurement will not tell you

They are the part worth reading. A benchmark with no stated limits has not been thought about carefully enough to trust.

  • There are no results on this page yet. The method is published first on purpose: a number without a method is discounted on sight, and we would rather be judged on the method while the run is in progress than surprise anyone with a figure later.

  • It will be our own benchmark, run by us. That is worth exactly what it is worth: the questions, the gold answers and the grading are ours, and until somebody independent runs it the right posture is scepticism. Publishing the question set is what makes that scepticism actionable.

  • One question can have more than one correct answer. Where two defensible calculations exist we will hold several gold answers rather than mark the alternative wrong, and we will say how often that happened.

  • A model grading a model is not sufficient. Where automated grading cannot separate a right answer from a plausible one, the case is graded by a person, and we will publish what fraction needed that.

  • We will publish the weakest dimension next to the strongest. A scorecard with no low number on it has not been measured, it has been assembled.

Questions about the method

Why not just use an existing text-to-SQL benchmark?

Because they measure the easy half. The published gaps in those benchmarks (question complexity, schema complexity, query complexity and semantic alignment) all get worse on SAP rather than better, and the last one has no equivalent at all: there is no public benchmark where the correct answer depends on your organisation's definition of a metric. A strong score on one of those sets would not predict anything about your close.

When will there be numbers?

When the run is done and the method has not changed to flatter it. If you are evaluating now and want the current state, ask; you will get what exists rather than a page-shaped version of it.

Can we run it against our own system?

That is the version worth having, and it is the better use of a pilot than any figure we publish. Your questions, your definitions, your data, graded by your own people against what they would have produced by hand.

What is the reproducibility test?

A different and much simpler claim: that re-opening a saved view returns the same figure. A stored view holds a compiled query rather than a prompt, so there is no model in the read path and nothing to re-reason. Running the same view repeatedly over a period and publishing the identical-result rate is a measurement we can make without any grading at all, and it is the one a CFO cares about most: the number should not move between the board pre-read and the board meeting.

Measure it on your own questions

The version that settles it is a scoped pilot: your questions against your own SAP, graded by the people who would otherwise have produced the answer by hand. The first session is thirty minutes on a system we provide.