Benchmarked Against Everything.
Orders of Magnitude More Efficient. Measured.
DriftMind was measured against adaptive ARIMA and Prophet, against OneNet, the leading deep architecture for online forecasting (NeurIPS 2023), and against the Chronos-Bolt and Google TimesFM foundation models. Same data, same protocol, and wherever possible the same physical hardware. DriftMind ran CPU-only, cold-start, with zero retraining; the same engine deploys unchanged from edge gateways to cloud. The results below presents a comprehensive and transparent evaluation.
Benchmark 1 · NAB Multi-Dataset
Three-Way Direct Comparison
All three models were evaluated on six datasets from the Numenta Anomaly Benchmark (NAB), spanning machine sensors, cloud infrastructure, urban demand, ad-exchange pricing, and social-media volume, at three forecast horizons (h = 1, 6, 24). Same streaming protocol, same hardware, no HTTP overhead, no artificial batching. Results are honest and reported per horizon: ARIMA is competitive on stable, non-seasonal signals at short horizons; DriftMind dominates seasonal multi-step forecasting and is 46–635× faster even at its most demanding horizon (and over 1,000× faster at h = 1) with zero retraining.
NAB · Machine Temperature System Failure · 22,695 points · 5-min interval
| Model | MAE (lower is better) | Throughput (pred/s) | Total Time | Retraining Cycles |
|---|---|---|---|---|
| DriftMind Accuracy Speed | 0.7704 | 10,274 | 2.21 s | Continuous |
| Adaptive ARIMA | 0.8529 | 47 | 483.10 s | 901 |
| Triggered Prophet | 2.9901 | 13 | 1,740.02 s | 4,105 |
Methodology. DriftMind uses one engine per horizon, with the input window sized to the horizon (h = 1 → 15 steps, h = 6 → 30, h = 24 → 120); MAE is therefore reported per horizon, while DriftMind's throughput and total time are the average across its three horizon engines. The evaluated baselines are adaptive, not static: ARIMA(5,1,0) runs in state-space form and re-applies its fitted coefficients to the rolling 200-observation window through a Kalman filter at every step, with a full maximum-likelihood re-estimation triggered whenever accuracy degrades (sustained MASE above 3, or forecast-to-actual correlation below 0.2); Prophet refits fully on the same triggers and extrapolates its most recent fit between them. Their window is fixed at 200 points for every horizon; each step produces one 24-step forecast vector scored at all three horizons, so their throughput, total time and retraining count are engine-level (unchanged across the horizon tabs) while only their MAE varies. ARIMA / Prophet "Total Time" covers the single engine pass that serves all three horizons. DriftMind accuracy figures are from a September 2026 run of the current forecasting engine (5.0.0) on the same data and protocol as the previous run in July 2026 using the previous version of Driftmind. The full suite was run twice, with identical baseline results to the previous run in July 2026.
Mean Absolute Error
Throughput (predictions / second) · log scale
Total Execution Time (seconds) · log scale
Retraining Cycles
Benchmark 2 · ETTh2 & ETTm1 vs OneNet
DriftMind vs. State-of-the-Art Online Deep Learning
DriftMind was benchmarked against OneNet (NeurIPS 2023), the leading deep architecture for online forecasting under concept drift, which itself outperforms PatchTST, FEDformer, FSNet, and OnlineTCN. Protocol matched to OneNet's published setup exactly: all 7 channels (features=M), scaler fit on the training split only, identical evaluation windows and forecast-origin counts, MAE in the standardised space OneNet reports. Both models are evaluated under a strictly causal online update: no model is trained on any observation at or after the forecast origin. OneNet's published figures, produced under an evaluation that exposed future observations to the model, differ and are discussed in the conclusion. A naive-persistence floor is included: if a model cannot beat "tomorrow = today", the forecast has no value.
ETTh2 · Hourly Electricity Transformer Temperature · MAE (standardised space, lower is better)
| Model | MAE | Hardware | Warm-up Required |
|---|---|---|---|
| DriftMind | 0.346 | CPU only | None (cold-start) |
| OneNet (NeurIPS 2023) | 0.382 | GPU | 25% of dataset |
| Naive persistence (floor) | 0.336 | n/a | n/a |
ETTm1 · 15-minute Interval · MAE (standardised space, lower is better)
| Model | MAE | Hardware | Warm-up Required |
|---|---|---|---|
| DriftMind | 0.192 | CPU only | None (cold-start) |
| OneNet (NeurIPS 2023) | 0.184 | GPU | 25% of dataset |
| Naive persistence (floor) | 0.186 | n/a | n/a |
MAE · ETTh2 (Horizon = 1, incl. persistence floor)
Runtime per evaluation cell · same machine (seconds, log scale)
Benchmark 3 · Chronos-Bolt & Google TimesFM
DriftMind vs. Time-Series Foundation Models
The same protocol, extended to the current generation of pre-trained foundation models: amazon/chronos-bolt-base (205M parameters) and google/timesfm-2.5-200m. All three engines ran on the same machine, CPU-only, identical windows and data. The foundation models arrived pre-trained on ~100 billion points; DriftMind cold-started on data it had never seen.
Accuracy · ETT · MAE (standardised space)
| Cell | Persistence | DriftMind | Chronos-Bolt | TimesFM 2.5 |
|---|---|---|---|---|
| ETTh2 H=1 | 0.336 | 0.346 | 0.308 | 0.304 |
| ETTh2 H=6 | 0.504 | 0.489 | 0.391 | 0.393 |
| ETTh2 H=24 | 0.688 | 0.705 | 0.515 | 0.529 |
| ETTm1 H=1 | 0.186 | 0.192 | 0.175 | 0.171 |
| ETTm1 H=6 | 0.296 | 0.317 | 0.241 | 0.243 |
| ETTm1 H=24 | 0.525 | 0.455 | 0.358 | 0.361 |
The foundation models are more accurate on every completed ETT cell: a ~10% relative gap at short horizons that grows with longer targets. The rest of this section provides an explanation for why that is not the whole picture.
Resources · per ~75,500-forecast evaluation cell, same machine
| Engine | Wall-clock per cell | Throughput (forecasts/s) | Peak RAM |
|---|---|---|---|
| DriftMind Speed Memory | 3.1–12.7 s (incl. continuous learning) | ~5,950–23,600 | 0.58 GB |
| Chronos-Bolt | 336–363 s (inference only) | ~208–225 | 1.85 GB |
| TimesFM 2.5 | 3,786 s (inference only) | ~20 | 2.12 GB |
Cost per 1M forecasts · measured CPU-time × list instance price (DigitalOcean 8 vCPU, $0.25/hr)
| Engine | $ / 1M forecasts | vs DriftMind |
|---|---|---|
| DriftMind | $0.003–0.012 | baseline |
| Chronos-Bolt | $0.31–0.33 | 26–106× |
| TimesFM 2.5 | ~$3.48 | 298–1,180× |
Benchmark 4 · Synthetic RAN Telemetry
On Network Telemetry, the Gap Nearly Closes
The ETT datasets favour long-horizon pattern models. On telemetry closer to DriftMind's home domain, the open ranfst synthetic RAN dataset (PRB utilisation, RRC connections, user throughput; 5-minute cadence; stationary, regime-change and signalling-storm scenarios), the picture changes: the accuracy gap to both foundation models shrinks to ~5% relative MAE, while the compute gap remains orders of magnitude. All engines evaluated on identical windows; DriftMind cold-started.
| Scenario · Horizon | Persistence | DriftMind | Chronos-Bolt | TimesFM 2.5 |
|---|---|---|---|---|
| Stationary · 5 min | 0.3814 | 0.2957 | 0.2833 | 0.2832 |
| Stationary · 1 h | 0.3967 | 0.3099 | 0.2933 | 0.2942 |
| Regime change · 5 min | 0.3660 | 0.2873 | 0.2745 | 0.2743 |
| Regime change · 1 h | 0.3817 | 0.3002 | 0.2821 | 0.2825 |
| Signalling storm · 5 min | 0.3812 | 0.2979 | 0.2862 | 0.2854 |
| Signalling storm · 1 h | 0.4002 | 0.3143 | 0.2978 | 0.2974 |
MAE, standardised space; 3-hour horizon (measured separately) shows the same ordering. All three engines degrade gracefully with horizon (~13–15%); persistence degrades ~23%. Same-machine throughput on this workload: DriftMind ~10,300 forecasts/s including continuous learning, vs ~221 (Chronos-Bolt) and ~20 (TimesFM).
Benchmark Conclusion
What the Benchmark Shows
In this benchmark, traditional forecasting models, whether classical statistical approaches or deep learning-based architectures, struggle to adapt effectively to evolving data streams in real time, particularly when facing very different types of data sets and noise.
Foundation Models for Time Series (FMTS) are among the few approaches capable of demonstrating meaningful adaptation while maintaining competitive accuracy. However, they face two significant limitations:
Computational requirements
Their substantial resource demands make deployment in constrained environments, particularly at the edge, challenging.
Potential data contamination
These models are pretrained on enormous datasets that may contain data similar to, or even overlapping with, commonly used benchmarks such as ETT. This raises questions about how well their reported accuracy generalises to genuinely unseen data.
Indeed, when evaluated on previously unseen workloads, such as our Telecom Radio dataset, their forecasting accuracy deteriorates significantly.
This is where DriftMind stands apart. It combines genuine online adaptability with a lightweight computational footprint, making it suitable for resource-constrained and edge environments. More importantly, it maintains competitive forecasting accuracy even when operating on previously unseen data, without requiring extensive pretraining or conventional model retraining.
DriftMind therefore addresses three fundamental requirements that existing forecasting approaches struggle to deliver simultaneously: online adaptability, computational efficiency, and robust generalisation to unseen workloads.
Online adaptability
Learning and inference are the same causal pass: every observation updates the model as it arrives, one point at a time, never looking ahead. No retraining cycle, no warm-up window.
Computational efficiency
Thousands to tens of thousands of forecasts per second on a single commodity CPU, in a sub-gigabyte footprint. No GPU, deployable from the edge to the cloud.
Robust generalisation
Because it learns each stream from scratch rather than recalling a pretraining corpus, its accuracy holds on genuinely unseen workloads, such as the Telecom Radio dataset, where pretrained models fall back.
| Approach | Online adaptability | CPU / edge efficiency | Generalises to unseen |
|---|---|---|---|
| ARIMA / Prophet | Partialretrain lag | Noheavy retraining | n/a |
| OneNet | Yes* | NoGPU | n/a |
| Foundation models | Yes | No | Nocontamination |
| DriftMind | Yes | Yes | Yes |
* Under the paper's non-causal online evaluation; corrected in Benchmark 2 above.
Benchmark Methodology
Setup & Reproducibility
All benchmarks were designed to be as fair and straightforward as possible: same data, same evaluation windows, and, for OneNet and the foundation models, the same physical machine. No cherry-picking of segments. Hyperparameters tuned on validation splits only, never on test data. Where DriftMind loses, the loss is printed.
NAB Dataset
Machine Temperature System Failure series from the Numenta Anomaly Benchmark. 22,695 data points. Continuous streaming simulation: no HTTP overhead, no artificial batching.
ETTh2 & ETTm1 Datasets
Public electricity-transformer datasets, evaluated exactly as OneNet's published protocol: all 7 channels (features=M), one engine per channel, macro-averaged; the scaler is fit on the training split only; forecast-origin counts match OneNet's test set. MAE is reported in the standardised space OneNet reports.
Hardware · one machine
All same-platform comparisons ran on a single Apple M5 Max: DriftMind CPU-only (JVM 17), OneNet on the on-chip GPU (PyTorch MPS), Chronos-Bolt and TimesFM on CPU (torch, 6 threads). Identical data, identical windows; the speed and cost ratios are exact, not cross-hardware estimates.
Cold-Start Condition
DriftMind benchmarks were conducted with no warm-up and no pre-loaded history. Prediction begins from the very first observation. OneNet uses the first 25% of each dataset as its mandatory pre-training window.
OneNet Replication
The official OneNet implementation was cloned and run; its published accuracy was reproduced closely on our hardware port, confirming benchmark integrity. Published OneNet numbers are used in the accuracy tables; our own run supplies the same-machine timing.
DriftMind Settings
Input length 30–60, max clusters 200, Gaussian smoothing; configuration selected on a validation split only. One consistent finding: the engine's adaptation memory scales with horizon: short memory wins at H=1, long memory at H≥24.
The academic paper detailing the full architecture and ETT benchmark methodology is available here: DriftMind: A Self-Adaptive, Cold-Start Framework for Time Series Forecasting.
See It Running on Your Data
Run DriftMind on your own time series: CSV upload, no setup, no GPU required. Results in seconds.