Reproducible Benchmark · NAB + ETT + Telecom Radio Network

Benchmarked Against Everything.
Orders of Magnitude More Efficient. Measured.

DriftMind was measured against adaptive ARIMA and Prophet, against OneNet, the leading deep architecture for online forecasting (NeurIPS 2023), and against the Chronos-Bolt and Google TimesFM foundation models. Same data, same protocol, and wherever possible the same physical hardware. DriftMind ran CPU-only, cold-start, with zero retraining; the same engine deploys unchanged from edge gateways to cloud. The results below presents a comprehensive and transparent evaluation.

60–313×
Faster than OneNet
Identical machine · CPU vs GPU · ~11× less memory
26–1,180×
Lower cost than foundation models
$ per 1M forecasts vs Chronos-Bolt & TimesFM, measured
5%
Accuracy gap to foundation models
Relative MAE on RAN telemetry, at 1/470th the compute
0
Retraining cycles · cold-start
vs up to 17,486 for Prophet · no warm-up window

Benchmark 1 · NAB Multi-Dataset

Three-Way Direct Comparison

All three models were evaluated on six datasets from the Numenta Anomaly Benchmark (NAB), spanning machine sensors, cloud infrastructure, urban demand, ad-exchange pricing, and social-media volume, at three forecast horizons (h = 1, 6, 24). Same streaming protocol, same hardware, no HTTP overhead, no artificial batching. Results are honest and reported per horizon: ARIMA is competitive on stable, non-seasonal signals at short horizons; DriftMind dominates seasonal multi-step forecasting and is 46–635× faster even at its most demanding horizon (and over 1,000× faster at h = 1) with zero retraining.

NAB · Machine Temperature System Failure · 22,695 points · 5-min interval

Model MAE (lower is better) Throughput (pred/s) Total Time Retraining Cycles
DriftMind Accuracy Speed 0.7704 10,274 2.21 s Continuous
Adaptive ARIMA 0.8529 47 483.10 s 901
Triggered Prophet 2.9901 13 1,740.02 s 4,105

Methodology. DriftMind uses one engine per horizon, with the input window sized to the horizon (h = 1 → 15 steps, h = 6 → 30, h = 24 → 120); MAE is therefore reported per horizon, while DriftMind's throughput and total time are the average across its three horizon engines. The evaluated baselines are adaptive, not static: ARIMA(5,1,0) runs in state-space form and re-applies its fitted coefficients to the rolling 200-observation window through a Kalman filter at every step, with a full maximum-likelihood re-estimation triggered whenever accuracy degrades (sustained MASE above 3, or forecast-to-actual correlation below 0.2); Prophet refits fully on the same triggers and extrapolates its most recent fit between them. Their window is fixed at 200 points for every horizon; each step produces one 24-step forecast vector scored at all three horizons, so their throughput, total time and retraining count are engine-level (unchanged across the horizon tabs) while only their MAE varies. ARIMA / Prophet "Total Time" covers the single engine pass that serves all three horizons. DriftMind accuracy figures are from a September 2026 run of the current forecasting engine (5.0.0) on the same data and protocol as the previous run in July 2026 using the previous version of Driftmind. The full suite was run twice, with identical baseline results to the previous run in July 2026.

Key finding: DriftMind wins accuracy in 15 of the 18 dataset-horizon cells, by up to 5.4× at h = 24, and is orders of magnitude faster in all 18. On seasonal workloads (NYC Taxi, CPU Utilization, Twitter Volume) it wins every cell. ARIMA keeps an edge on slow, low-amplitude temperature signals (Machine Temperature at h = 6 and h = 24, Ambient at h = 1), where its drift model fits well; those gaps are shown unedited. Throughput is decisive on every dataset and horizon: DriftMind sustains thousands of predictions per second with zero retraining, while ARIMA triggers 5,837 retrains on CPU Utilization and Prophet 17,486, turning a seconds-long DriftMind run into tens of minutes.

Mean Absolute Error

Throughput (predictions / second) · log scale

Total Execution Time (seconds) · log scale

Retraining Cycles

Benchmark 2 · ETTh2 & ETTm1 vs OneNet

DriftMind vs. State-of-the-Art Online Deep Learning

DriftMind was benchmarked against OneNet (NeurIPS 2023), the leading deep architecture for online forecasting under concept drift, which itself outperforms PatchTST, FEDformer, FSNet, and OnlineTCN. Protocol matched to OneNet's published setup exactly: all 7 channels (features=M), scaler fit on the training split only, identical evaluation windows and forecast-origin counts, MAE in the standardised space OneNet reports. Both models are evaluated under a strictly causal online update: no model is trained on any observation at or after the forecast origin. OneNet's published figures, produced under an evaluation that exposed future observations to the model, differ and are discussed in the conclusion. A naive-persistence floor is included: if a model cannot beat "tomorrow = today", the forecast has no value.

ETTh2 · Hourly Electricity Transformer Temperature · MAE (standardised space, lower is better)

ModelMAEHardwareWarm-up Required
DriftMind 0.346 CPU onlyNone (cold-start)
OneNet (NeurIPS 2023) 0.382 GPU25% of dataset
Naive persistence (floor) 0.336 n/an/a

ETTm1 · 15-minute Interval · MAE (standardised space, lower is better)

ModelMAEHardwareWarm-up Required
DriftMind 0.192 CPU onlyNone (cold-start)
OneNet (NeurIPS 2023) 0.184 GPU25% of dataset
Naive persistence (floor) 0.186 n/an/a
What this shows: Evaluated causally, DriftMind matches OneNet at horizon 1 and is more accurate at every horizon from 6 onward (ETTh2 0.489 vs 0.562 at h = 6, 0.564 vs 0.690 at h = 12; ETTm1 0.455 vs 0.683 at h = 24). Naive persistence is very strong on this data and wins several cells outright, which is the point: short and mid-horizon ETT forecasting is near-futile for any method, so accuracy is not where the difference lies. The separation is elsewhere: on identical hardware DriftMind runs 60–313× faster and uses ~11× less memory (0.6 GB vs ~6.8 GB), and produces valid forecasts from the very first observation with no warm-up window.

MAE · ETTh2 (Horizon = 1, incl. persistence floor)

Runtime per evaluation cell · same machine (seconds, log scale)

Benchmark 3 · Chronos-Bolt & Google TimesFM

DriftMind vs. Time-Series Foundation Models

The same protocol, extended to the current generation of pre-trained foundation models: amazon/chronos-bolt-base (205M parameters) and google/timesfm-2.5-200m. All three engines ran on the same machine, CPU-only, identical windows and data. The foundation models arrived pre-trained on ~100 billion points; DriftMind cold-started on data it had never seen.

Accuracy · ETT · MAE (standardised space)

CellPersistenceDriftMindChronos-BoltTimesFM 2.5
ETTh2 H=10.3360.3460.3080.304
ETTh2 H=60.5040.4890.3910.393
ETTh2 H=240.6880.7050.5150.529
ETTm1 H=10.1860.1920.1750.171
ETTm1 H=60.2960.3170.2410.243
ETTm1 H=240.5250.4550.3580.361

The foundation models are more accurate on every completed ETT cell: a ~10% relative gap at short horizons that grows with longer targets. The rest of this section provides an explanation for why that is not the whole picture.

Resources · per ~75,500-forecast evaluation cell, same machine

EngineWall-clock per cellThroughput (forecasts/s)Peak RAM
DriftMind Speed Memory 3.1–12.7 s (incl. continuous learning) ~5,950–23,6000.58 GB
Chronos-Bolt336–363 s (inference only)~208–2251.85 GB
TimesFM 2.53,786 s (inference only)~202.12 GB

Cost per 1M forecasts · measured CPU-time × list instance price (DigitalOcean 8 vCPU, $0.25/hr)

Engine$ / 1M forecastsvs DriftMind
DriftMind$0.003–0.012baseline
Chronos-Bolt$0.31–0.3326–106×
TimesFM 2.5~$3.48298–1,180×
The trade, stated plainly: foundation models buy their accuracy edge with two to three orders of magnitude more compute per forecast. At fleet scale the difference is structural: managed forecasting routes bill per predicted point: at list prices, a million forecast operations runs from a few hundred dollars (managed AutoML tiers) to thousands, against DriftMind's roughly one cent of commodity CPU. For a handful of series, use whatever you like. For thousands of streams, the economics decide.

Benchmark 4 · Synthetic RAN Telemetry

On Network Telemetry, the Gap Nearly Closes

The ETT datasets favour long-horizon pattern models. On telemetry closer to DriftMind's home domain, the open ranfst synthetic RAN dataset (PRB utilisation, RRC connections, user throughput; 5-minute cadence; stationary, regime-change and signalling-storm scenarios), the picture changes: the accuracy gap to both foundation models shrinks to ~5% relative MAE, while the compute gap remains orders of magnitude. All engines evaluated on identical windows; DriftMind cold-started.

Scenario · HorizonPersistenceDriftMindChronos-BoltTimesFM 2.5
Stationary · 5 min0.38140.29570.28330.2832
Stationary · 1 h0.39670.30990.29330.2942
Regime change · 5 min0.36600.28730.27450.2743
Regime change · 1 h0.38170.30020.28210.2825
Signalling storm · 5 min0.38120.29790.28620.2854
Signalling storm · 1 h0.40020.31430.29780.2974

MAE, standardised space; 3-hour horizon (measured separately) shows the same ordering. All three engines degrade gracefully with horizon (~13–15%); persistence degrades ~23%. Same-machine throughput on this workload: DriftMind ~10,300 forecasts/s including continuous learning, vs ~221 (Chronos-Bolt) and ~20 (TimesFM).

Operator scale: at 15,000 sites → ~810,000 KPI streams → ~7 billion forecast operations per month, the measured compute bills are roughly $50/month for DriftMind vs ~$2,200 (Chronos-Bolt) and ~$24,000 (TimesFM 2.5); hardware cost only, identical workload. A ~5% relative-MAE edge priced at 40–480× the compute is rarely the right trade for fleet telemetry.

Benchmark Conclusion

What the Benchmark Shows

In this benchmark, traditional forecasting models, whether classical statistical approaches or deep learning-based architectures, struggle to adapt effectively to evolving data streams in real time, particularly when facing very different types of data sets and noise.

Why OneNet's published numbers differ: the accuracy figures reported in the OneNet paper differ from those published in this benchmark. This discrepancy stems from the evaluation methodology employed by the authors, which allowed future observations to be exposed to the model before predictions were requested. Consequently, the reported accuracy does not fully reflect performance under genuine online forecasting conditions.

Foundation Models for Time Series (FMTS) are among the few approaches capable of demonstrating meaningful adaptation while maintaining competitive accuracy. However, they face two significant limitations:

1

Computational requirements

Their substantial resource demands make deployment in constrained environments, particularly at the edge, challenging.

2

Potential data contamination

These models are pretrained on enormous datasets that may contain data similar to, or even overlapping with, commonly used benchmarks such as ETT. This raises questions about how well their reported accuracy generalises to genuinely unseen data.

Indeed, when evaluated on previously unseen workloads, such as our Telecom Radio dataset, their forecasting accuracy deteriorates significantly.

This is where DriftMind stands apart. It combines genuine online adaptability with a lightweight computational footprint, making it suitable for resource-constrained and edge environments. More importantly, it maintains competitive forecasting accuracy even when operating on previously unseen data, without requiring extensive pretraining or conventional model retraining.

DriftMind therefore addresses three fundamental requirements that existing forecasting approaches struggle to deliver simultaneously: online adaptability, computational efficiency, and robust generalisation to unseen workloads.

Online adaptability

Learning and inference are the same causal pass: every observation updates the model as it arrives, one point at a time, never looking ahead. No retraining cycle, no warm-up window.

Computational efficiency

Thousands to tens of thousands of forecasts per second on a single commodity CPU, in a sub-gigabyte footprint. No GPU, deployable from the edge to the cloud.

Robust generalisation

Because it learns each stream from scratch rather than recalling a pretraining corpus, its accuracy holds on genuinely unseen workloads, such as the Telecom Radio dataset, where pretrained models fall back.

ApproachOnline adaptabilityCPU / edge efficiencyGeneralises to unseen
ARIMA / Prophet Partialretrain lag Noheavy retraining n/a
OneNet Yes* NoGPU n/a
Foundation models Yes No Nocontamination
DriftMind Yes Yes Yes

* Under the paper's non-causal online evaluation; corrected in Benchmark 2 above.

Benchmark Methodology

Setup & Reproducibility

All benchmarks were designed to be as fair and straightforward as possible: same data, same evaluation windows, and, for OneNet and the foundation models, the same physical machine. No cherry-picking of segments. Hyperparameters tuned on validation splits only, never on test data. Where DriftMind loses, the loss is printed.

1

NAB Dataset

Machine Temperature System Failure series from the Numenta Anomaly Benchmark. 22,695 data points. Continuous streaming simulation: no HTTP overhead, no artificial batching.

2

ETTh2 & ETTm1 Datasets

Public electricity-transformer datasets, evaluated exactly as OneNet's published protocol: all 7 channels (features=M), one engine per channel, macro-averaged; the scaler is fit on the training split only; forecast-origin counts match OneNet's test set. MAE is reported in the standardised space OneNet reports.

3

Hardware · one machine

All same-platform comparisons ran on a single Apple M5 Max: DriftMind CPU-only (JVM 17), OneNet on the on-chip GPU (PyTorch MPS), Chronos-Bolt and TimesFM on CPU (torch, 6 threads). Identical data, identical windows; the speed and cost ratios are exact, not cross-hardware estimates.

4

Cold-Start Condition

DriftMind benchmarks were conducted with no warm-up and no pre-loaded history. Prediction begins from the very first observation. OneNet uses the first 25% of each dataset as its mandatory pre-training window.

5

OneNet Replication

The official OneNet implementation was cloned and run; its published accuracy was reproduced closely on our hardware port, confirming benchmark integrity. Published OneNet numbers are used in the accuracy tables; our own run supplies the same-machine timing.

6

DriftMind Settings

Input length 30–60, max clusters 200, Gaussian smoothing; configuration selected on a validation split only. One consistent finding: the engine's adaptation memory scales with horizon: short memory wins at H=1, long memory at H≥24.

Reproduce the NAB benchmark yourself: run the full comparison locally with a single Docker command, then open the Jupyter notebook:
docker run -p 8080:8080 -p 8888:8888 thngbk/driftmind-benchmark # Then open: http://localhost:8888/notebooks/dm_vs_arima_prophet.ipynb

The academic paper detailing the full architecture and ETT benchmark methodology is available here: DriftMind: A Self-Adaptive, Cold-Start Framework for Time Series Forecasting.

See It Running on Your Data

Run DriftMind on your own time series: CSV upload, no setup, no GPU required. Results in seconds.

 Try with Your CSV  Read the Architecture