Skip to content
AnrilX
All capabilities

Forward-looking

Trained on your history. Tested on what it never saw.

History is pulled from SAP, a model is trained on it, and it is backtested on whole periods held out of training. Then it tells you what it thinks happens next, and what is pushing it.

  • Drivers ranked, so a forecast can be argued with
  • A model that cannot be trusted shows no figures at all
  • Error is stated as error
  • A forecast is never presented as an actual
See it live

Net sales forecast · six months · Sep 2026 – Feb 2027

Net sales forecast · six months · Sep 2026 – Feb 2027

Trained on 2024-01 to 2026-05 and scored on three months deliberately held out of training, so the error below is measured rather than asserted. The interval on every month is shown, the months resting on extrapolation are named, and the one month the model got wrong by more than its own stated band is on the panel rather than smoothed out of it.

First forecast month

1.24bnSAR

September 2026

Six-month total

7.46bnSAR

Sep 2026 – Feb 2027

Against last actual

−4.2%

vs August 2026

Typical error

±2.5%

measured on 3 held-out months

Actual and forecast, by month

SAR

JanAprJulOct*Jan*
Actual Forecast

Six-month forecast by sales organisation

SAR

2.1bn 1600 2.0bn 1800 2.0bn 1400 1.4bn 1200

Backtest: predicted against actual

the three months held out of training

SAR

JunJulAug
Actual Predicted

Where the error sits

mean absolute error on the held-out months, per cent

1400 · Dammam 4.1 1200 · Export 3.2 1800 · Jeddah 1.9 1600 · Riyadh 1.4

Forecast by month

MonthForecastLowHighBasisOutside observed range
Sep 20261,244,100,000 SAR1,213,000,000 SAR1,275,200,000 SARseasonal + trendno
Oct 20261,261,400,000 SAR1,229,900,000 SAR1,292,900,000 SARseasonal + trendno
Nov 20261,288,200,000 SAR1,256,000,000 SAR1,320,400,000 SARseasonal + trendno
Dec 20261,196,300,000 SAR1,166,400,000 SAR1,226,200,000 SARseasonal + trendorg 1400
Jan 20271,242,800,000 SAR1,211,700,000 SAR1,273,900,000 SARseasonal + trendorg 1400
Feb 20271,229,200,000 SAR1,198,500,000 SAR1,259,900,000 SARseasonal + trendorg 1400
Six months7,462,000,000 SAR7,275,500,000 SAR7,648,500,000 SAR±2.5% measured3 of 6

Backtest: months held out of training

The band is the mean error, not a guarantee. August fell outside it and is reported as such.

MonthPredictedActualErrorInside ±2.5%
Jun 20261,302,000,000 SAR1,330,000,000 SAR−2.1%yes
Jul 20261,341,000,000 SAR1,376,000,000 SAR−2.5%yes
Aug 20261,335,000,000 SAR1,298,000,000 SAR+2.9%no
3 months3,978,000,000 SAR4,004,000,000 SAR2.5% MAPE2 of 3

Where this forecast is least trustworthy

  • The six months total 7.46bn SAR, 4.2% below August. That is the shape of the last two Septembers repeating, not a downturn the model has detected; it has no input that would let it detect one.
  • Three of the six months for organisation 1400 fall below anything in its training history. Those rest on extrapolation rather than measurement, are marked in the table, and should be compared with history rather than with the other forecast months.
  • ±2.5% is the mean absolute error over three months held out of training. It is not a guarantee: August was 2.9% out, which is outside the band, and it is on the panel because a forecast that only reports the months it got right is not reporting an error at all.
  • 1400 Dammam carries most of the error as well as all of the extrapolation. The group total is more reliable than the 1400 line inside it, and splitting this view by organisation is where somebody gets misled.

Source: zsd_sales_order / ZA_SalesOrder · net sales = sum of TotalNetAmount · grouped by SalesOrganization and calendar month · trained on 2024-01 to 2026-05 · backtested on 3 held-out months (2026-06 … 2026-08), 12 figures scored · interval from the residual distribution of the backtest

Outside the observed range: 3 of 6 forecast months for organisation 1400 fall below anything in the training history. Those months rest on extrapolation, not measurement. Where a series is too short to hold months back from, no figure is returned at all: an untested forecast is not a forecast with a wide interval, it is a guess, and the platform does not print one.

How it works

01

History comes from SAP, month by month

The series a model trains on is extracted from your own system one period at a time and stored so a re-run cannot double-count it. An incomplete trailing month is excluded and re-read once it closes: a part-month looks like a collapse in demand, and a model that learns from one will forecast a collapse.

02

The backtest holds out whole periods

Accuracy is measured against periods the model never saw during training, held out as complete periods rather than as scattered rows. Holding out rows leaks the shape of the period into training and produces an error figure that is better than the truth.

03

A gap month is filled before the model reads it

Where a month has no activity, the calendar is completed with a zero before any window functions run. Otherwise “the previous month” silently means “the previous month that had data”, and every lag feature in the model is subtly wrong.

04

A retrained model has to earn its place

A new model is compared against the standing one on the same test window before it replaces it. Where the windows differ the comparison is not treated as evidence, because two accuracy figures measured over different periods are not comparable and pretending otherwise is how a worse model quietly takes over.

Three rules it holds to

These are constants in the product rather than intentions:

  1. A model that cannot be trusted shows no figures. Above a fixed error threshold, the forecast reports that it is unreliable and stops. It does not render a neat table that is indistinguishable in shape from a good one.
  2. Error is stated as error. The backtest figure travels with the forecast.
  3. A forecast is never presented as an actual. It is labelled everywhere it appears.

Why the holdout is by period

The tempting way to measure a forecast is to hold out a random sample of rows and see how well the model predicts them. It produces a better-looking number and it is wrong.

A time series leaks. If the model has seen most of a month, predicting the rest of that month is not forecasting; it is interpolation, and the accuracy figure you get describes a task nobody is going to ask it to do. Holding out whole periods measures the thing you actually care about: performance on a period the model has never seen at all.

We found this the hard way. A leaking backtest reported an error rate roughly a third better than the honest one. The honest figure is the one that ships.

The month that looks like a collapse

An extraction that runs mid-month picks up a partial month and stores it as if it were complete. The model then sees a period where volume fell by 80%, and learns that the series sometimes collapses.

Measured on real data, this produced a forecast with an error rate near 170%: a model that was confidently, comprehensively wrong. Excluding the incomplete trailing month and re-reading it once it closes took the same model to under 2%.

Neither number was visible from anywhere except the backtest, which is the argument for having one.

In more detail

How this one actually behaves

The three rules the engine holds itself to

A model that cannot be trusted shows no figures. Error is stated as error. A forecast is never presented as an actual. These are shipped constants rather than aspirations, and the first one is the one that costs something: it means a question you asked in good faith can come back with three of four organisations answered and the fourth declined, which is less satisfying than four numbers and considerably more useful than three good ones and a guess.

Why backtesting on held-out periods is the only claim worth making

A model evaluated on the history it was trained on will always look excellent, which is why most accuracy claims are worthless. Whole periods are held out of training and the model is scored on those, so the figure quoted above the forecast is what it achieved on months it had never seen, not how well it memorised the ones it had.

Outside the observed range is called out, not smoothed over

When a projection falls below anything in the training history, the months concerned are named and the reason is stated: the figures rest on extrapolation rather than on anything measured. The instruction that comes with it matters as much: compare those months against the history, never against the other forecast months, because that would only show whether the forecast is consistent with itself.

Where it stops

What this capability will not do

Where this one stops, said here rather than discovered in a pilot.

  • A model whose error is above the reliability threshold shows no figures at all; a confident-looking table from a bad model is worse than no table.

  • A forecast is labelled a forecast everywhere it appears. It is never blended into a figure that reads as an actual.

  • Forecasting needs enough history and enough density. Where a series is too short, too sparse or too stale, the platform says so instead of producing a number.

Questions people ask us

How do you know the forecast is any good?

It is backtested against whole periods held out of training, and the resulting error is reported with the forecast rather than kept internal. Above a fixed error threshold the platform shows no figures at all.

Can it tell us why a number is moving?

Yes. Drivers are ranked by contribution, which is what makes a forecast something you can argue with rather than a number you either accept or ignore.

Does it forecast anything, or a fixed list of KPIs?

The KPI is proposed from your question and then verified against your catalogue and checked for density before anything is trained. Nothing is hardcoded to a particular measure.

Bring the question your reports cannot answer

Thirty minutes against a live SAP system we provide: no access to yours, nothing to set up. If it cannot answer, you find that out in half an hour rather than three months into a pilot.