In a Streaming World, Should You Stand Still?
A Comprehensive Benchmark of Anomaly Detection in Streams
Abstract.
Time series anomaly detection (TSAD) is increasingly deployed in streaming settings, where data arrive sequentially and may exhibit non-stationarity. As a result, several works from the recent literature propose streaming anomaly detection methods that rely on incremental updates to adapt over time. However, most of these approaches originate from the streaming outlier detection literature and largely ignore core characteristics of time series anomalies. Moreover, their empirical evaluation is typically conducted on synthetic or small-scale benchmarks with limited diversity, making it unclear whether streaming methods are truly advantageous in realistic TSAD scenarios. In this work, we carry out the first large-scale experimental study comparing streaming and static TSAD methods under a unified streaming evaluation benchmark. We consider a realistic setting in which an initial batch of data is available for model training, followed by online evaluation of both detection accuracy and computational efficiency. In addition, we propose a distribution-drift dataset of real time series, called TSB-drift, to isolate scenarios where streaming updates are theoretically justified. Our results show that, contrary to common assumptions, static TSAD methods significantly outperform streaming approaches in most streaming settings. Such finding highlights a critical gap between the design of existing streaming methods and the requirements of modern TSAD, and calls for a rethinking of how streaming capabilities should be integrated into TSAD.
Keywords:
Time Series, Stream, Anomaly Detection, BenchmarkKDD Availability Link:
The source code of this paper has been made publicly available at https://doi.org/10.5281/zenodo.20310651.
1. Introduction
Time series anomaly detection (TSAD) plays a critical role in a wide range of real-world applications, including industrial control (Lei et al., 2013), energy production (Alkuwari et al., 2022), healthcare (Goldberger et al., 2000) and Internet-of-Things (IoT) platforms (García-Teodoro et al., 2009). In many of these fields, data are generated continuously and must be processed sequentially, motivating the development of TSAD methods suitable for Streaming settings. As a consequence, recent years have seen increasing interest in streaming anomaly detection approaches (Vázquez et al., 2023; Cao et al., 2025; Salles et al., 2025) that incrementally update their models to handle evolving data distributions (Gama et al., 2014).
A common assumption is that streaming TSAD methods are inherently better suited for streaming data than static alternatives. In particular, update mechanisms (Cao et al., 2014; Kontaki et al., 2011), forgetting factors (Bhatia et al., 2022; Hartl et al., 2020), or incremental model updates (Sathe and Aggarwal, 2018; Manzoor et al., 2018), are often presented as essential tools for dealing with non-stationarity (Vázquez et al., 2023). However, it remains unclear whether existing streaming methods actually deliver superior performance when applied to realistic TSAD tasks.
This uncertainty lies in the conceptual gap between the literature on streaming anomaly detection and on TSAD. Much of the streaming literature focuses on point-wise outlier detection, where each observation is treated independently and anomalies are defined as individual points (Angiulli and Fassetti, 2007; Ramaswamy et al., 2000; Kontaki et al., 2011; Hartl et al., 2020; Manzoor et al., 2018; Sathe and Aggarwal, 2018). In contrast, TSAD fundamentally relies on the temporal structure of the data: anomalies often manifest as subsequences (Keogh et al., 2005; Boniol et al., 2024) or collective patterns that become apparent when contextualized within a broader temporal window. Ignoring these characteristics can lead to well-suited models for detecting isolated outliers but poorly designed to identify anomalies commonly encountered in real-world time series (Wu and Keogh, 2022; Paparrizos et al., 2022).
A second equally important issue concerns evaluation practices. Most streaming anomaly detection methods are evaluated on synthetic datasets (Cao et al., 2025; Vázquez et al., 2023; Ntroumpogiannis et al., 2023) or on a small number of narrowly scoped real-world time series (Salles et al., 2025; Vázquez et al., 2023). These benchmarks often lack diversity in terms of domains, temporal characteristics, anomaly types and degrees of non-stationarity. Meanwhile, recent work in TSAD literature has introduced large-scale, heterogeneous benchmarks (Liu and Paparrizos, 2024) that better reflect the variety of real-world applications. To date, however, these modern benchmarks have not been used to assess streaming anomaly detection methods. Therefore, the aforementioned limitations raise a fundamental question:
In this paper, we address this question through our proposed benchmark, called StrAD. We evaluate a large collection of static and streaming methods under a unified streaming evaluation framework. Our setting assumes an initial batch of data for model training, followed by sequential processing of incoming data. We operate in an unsupervised context, so the time series may contain anomalies within the training batch. Static methods are trained once and remain unchanged, while streaming methods apply their update mechanisms during the online phase. Importantly, we define streaming methods not merely as those capable of real-time execution, but as methods that explicitly rely on model updates to improve efficiency or accuracy in the presence of distribution drifts. Static methods used in streaming settings are called Online methods.
Our results reveal a striking finding: online TSAD methods consistently outperform streaming approaches in most streaming settings, often by a substantial margin. This performance gap persists even though streaming methods are explicitly designed to adapt over time. Results on TSB-drift, our novel distribution-drift subset of TSB-AD-M (Liu and Paparrizos, 2024), reveal that, although the performance gap narrows in the presence of concept drifts, streaming methods do not manage to reverse the trend. These findings challenge prevailing assumptions in the streaming TSAD literature and highlight important limitations of current streaming approaches. Rather than demonstrating an inherent advantage of streaming methods, our study suggests that effective time series modeling and anomaly representation remain more critical than incremental model updates alone. Overall, our contributions are as follows:
- •
We first properly define the concept of Static, Online and Streaming TSAD methods (Section 2.3).
- •
We carefully describe our experimental benchmark StrAD and introduce a novel dataset subset of TSB-AD-M, called TSB-drift that contains only real-world time series that manifest a concept drift (Section 3).
- •
We present our experimental results comparing static, online and streaming TSAD methodologies on both TSB-AD benchmark and TSB-drift (Section 4).
- •
We conclude by discussing the implications of these results for potential future research (Section 5).
- •
We release our code and results in our repository 11 1 https://github.com/magaliparrino/StrAD.git.
2. Background and Related Work
We now introduce formal definitions relevant for this paper.
A time series is a sequence of values , where is the length of , and is the point of . A subsequence of of length starting at position is defined as with .
A multivariate time series is a collection of univariate time series of equal length . Formally, we denote as , where for each dimension , the corresponding univariate time series is , with for . Moreover, a subsequence of dimension of length starting at position is defined as . The corresponding multivariate subsequence of length starting at position is then .
The output of an anomaly detection algorithm applied to a time series is commonly expressed as an anomaly score sequence , where denotes the anomaly score associated with point . Higher scores typically indicate a higher degree of abnormality.
When available, the ground-truth annotation is represented by a label sequence , where indicates whether point is normal () or anomalous (). These labels are used exclusively for evaluation purposes and are not accessible to the detection algorithms during training or inference when in unsupervised settings.
Anomaly detection in time series is predominantly studied as an unsupervised learning problem, as prior information about anomalies or even normal behavior is often unavailable in real-world scenarios. Moreover, most practical applications involve multivariate time series. Consequently, throughout the remainder of this paper, we restrict our study to unsupervised anomaly detection methods operating on multivariate time series.
2.1. Time Series Anomaly Detection (TSAD)
A large panel of TSAD methods have been proposed in the literature (Boniol et al., 2024). The latter can be grouped into three main categories, (i) distance-based, (ii) density-based, and (iii) prediction-based.
2.1.1. Distance-based
These methods utilize distances between subsequences to detect anomalies. We can identify three sub-categories: (i) discord-based, i.e., methods that focus on the analysis of subsequences for the purpose of detecting anomalies in time series, mainly by employing nearest neighbor distances among subsequences (Yeh et al., 2016; Senin et al., 2015; Yankov et al., 2008); (ii) proximity-based, i.e., methods measuring the density of the neighborhood of particular subsequences (Breunig et al., 2000); and (iii) clustering-based, i.e., methods using the distance to a cluster partition (Boniol et al., 2021a; Boniol et al., 2021b).
2.1.2. Density-based
These methods detect anomalies by evaluating the density of the points or subsequences into a specific representation space. We identify four sub-categories: (i) distribution-based such as the HBOS (Goldstein and Dengel, 2012) method, which estimates the distribution of each feature using histograms and detects anomalies as points that fall into low-density regions ; (ii) graph-based in which a method, such as Series2Graph (Boniol and Palpanas, 2020), converts the time series into graphs to facilitate the detection of anomalies ; (iii) tree-based such as Isolation Forest (Liu et al., 2008) which groups points or subsequences into different trees ; and (iv) encoding-based methods, such as PCA that projects data onto a lower dimensional subspace to identify the deviations from the dominant structure.
2.1.3. Prediction-based
These methods detect anomalies using prediction errors to a given self-supervised task. We can identify two sub-categories: (i) forecasting-based, such as recurrent (Malhotra et al., 2015) or convolutional neural network (Munir et al., 2019) that use the past values as input, predict the following one, and use the forecasting error as the anomaly score; and (ii) reconstruction-based, such as AutoEncoder approaches (Sakurada and Yairi, 2014) trained to reconstruct the time series and that use the reconstruction error as an anomaly score.
2.2. Streaming Anomaly Detection
There are two main challenges when it comes to streaming TSAD: the unknown incoming data, whose distribution can be drastically different from the historical data and may force an update of the model, and the computational constraints: limited memory and short execution time in case of real-time analysis (Pevný, 2016; Guha et al., 2016; Sathe and Aggarwal, 2018; Tan et al., 2011).
2.2.1. Concept Drift
In non-stationary environments, the underlying data distribution can change over time, resulting in what is known as Concept Drift. In the seminal work (Gama et al., 2014), it is formally defined as occurring between two time points and if the joint distribution of the input variables and the labels changes, such that: However , concept drift was initially introduced in the context of supervised learning (Gama et al., 2014), where both the labels and the data are available. It enabled the distinction between virtual drift (changes in the incoming data distribution ) and real concept drift (changes in ).
In unsupervised settings, is unavailable, hence we cannot directly measure . We follow the common assumption in streaming literature (Vázquez et al., 2023; Cao et al., 2025) that significant changes in serve as a proxy for concept drift. A shift in means that the statistical patterns learned during training no longer match the incoming data. In this paper, we consider a concept drift to occur when the distribution differs between historical and new data.
2.2.2. Update Mechanism (UM)
To handle concept drift, streaming models contain an update mechanism. We identify two categories: (i) numerical and (ii) structural updates. In the former, models set their internal geometry during the training phase (using random projections for models such as LODA (Pevný, 2016) and xStream (Manzoor et al., 2018), or space partitioning for RSHash (Sathe and Aggarwal, 2018), HSTree (Tan et al., 2011) and SDOstream (Hartl et al., 2020)). The update will merely consist in adjusting frequency counts or mass profile in these fixed regions. As such, HSTree walks down the fixed trees until the new point reaches a leaf and raises the mass of every node on the path. Structural updates, on the other hand, correspond to evolving structures that adapt to concept drifts. RRCF (Guha et al., 2016) inserts and deletes nodes for each new arrival while MCOD (Kontaki et al., 2011) manages dynamical clustering within sliding windows. This process, although better for capturing complex shifts, results in higher computational cost.
2.2.3. Memory Management (MM)
Three strategies exist: (i) tumbling memory, (ii) sliding memory and (iii) soft forgetting. The former includes methods that compare a reference window with a current building window. Once the current window matches the size of the reference window, it is tumbled as the new reference window, and a new current window is built incrementally. The latter is used in xStream (Manzoor et al., 2018) and HSTree (Tan et al., 2011). In contrast, models employing sliding memory operate under a First In First Out paradigm, removing the oldest data when new data are added, whether as individual points (e.g., SWKNN (Ramaswamy et al., 2000), MCOD (Kontaki et al., 2011), RRCF (Guha et al., 2016), RSHash (Sathe and Aggarwal, 2018)) or small batches (e.g., LEAP (Cao et al., 2014)). Finally, soft forgetting regroups less abrupt memory management. SDOstream (Hartl et al., 2020) employs an aging mechanism, assigning a decreasing weight to observations so that historical data has a diminishing impact on current evaluations. MemStream (Bhatia et al., 2022) maintains a fixed-size memory of normal points, adding a new point only if it is sufficiently normal and discarding the oldest entry.
2.3. Static vs. Online vs. Streaming Methods
We now distinguish TSAD methods based on data availability for anomaly score computation and model updates ability. The latter constitutes the major difference between (i) Static, (ii) Online, (iii) Streaming methods, depicted in Figure 1 and defined as follows.
Definition 0 (Static Method).
A static method computes the anomaly score at timestamp using the entire multivariate time series . Formally, the anomaly score at timestamp may depend on any observation for all dimensions and all timestamps . There is no temporal ordering constraint.
Definition 0 (Online Method).
An online method assumes access to an initial batch for some . After this initial batch, the model parameters are fixed. For any timestamp , the anomaly score is computed using only and with and . No model updates are performed..
Definition 0 (Streaming Method).
A streaming method processes the time series sequentially and allows model updates over time. At timestamp , the anomaly score is computed using observations available up to time , i.e., with and . After computing , the model may update its internal state or parameters.
These notions are closely related to those introduced in a recent streaming survey (Correia et al., 2024). In practice, most methods in TSAD literature (cf. Section 2.1) are static methods, assuming access to the full time series. Nevertheless, many of these approaches can naturally be deployed in an online setting, by performing training on a initial batch , followed by inference on incoming observations without further model updates. This observation suggests that static-online boundary is often methodological rather than intrinsic, and that static formulations do not prevent practical online use.
2.4. Gap in Existing Evaluations
Despite extensive research on both static and streaming TSAD, current benchmarks exhibit the following three major limitations:
2.4.1. Streams are Time Series
Data streams consist of ordered observations collected over time. However, the literature on stream outlier detection, largely treats streams as generic sequences of data points (Vázquez et al., 2023), arguing that computational requirements fundamentally differ. Thus, time-series–specific characteristics are often not explicitly considered. This abstraction has led to a separation between the literature on streaming and TSAD, where TSAD methods are typically overlooked in streaming evaluations.
2.4.2. Anomalies are not Only Points
Anomalies are not restricted to isolated observations but also manifest as subsequences
(Keogh et al., 2005; Boniol et al., 2024; Chandola et al., 2009), often referred to as collective anomalies (Chandola et al., 2009).
Detecting such anomalous patterns over contiguous time intervals is a central challenge in the TSAD literature (Boniol et al., 2024). Nevertheless, this notion is sometimes overlooked even within TSAD, especially for multivariate time series, where the complexity of jointly modeling temporal and cross-dimensional structure can lead to a focus on point-wise anomalies. This disconnect is even sharper in streaming outlier literature, which prioritizes identifying
point anomalies (Angiulli and Fassetti, 2007; Ramaswamy et al., 2000; Kontaki et al., 2011; Hartl et al., 2020; Manzoor et al., 2018; Sathe and Aggarwal, 2018) over anomalous contiguous sequences.
2.4.3. Is Concept Drift even Real?
While concept drift is inherent in real-world time series and its treatment is recognized as a central methodological challenge in online and streaming anomaly detection, existing anomaly detection benchmarks provide limited evidence of its occurrence in real time series. In particular, there is a lack of diverse real datasets with explicitly identified concept drift regarding anomaly detection. Many benchmarks rely on synthetically generated drift patterns (Cao et al., 2025; Vázquez et al., 2023), or employ real-world datasets without conducting a dedicated drift analysis (Vázquez et al., 2023; Salles et al., 2025). As noted by (Cao et al., 2025), current streaming datasets typically focus either on anomaly detection or on concept drift, but rarely address both simultaneously. As a result, the practical impact of concept drift on TSAD methods’ performance remains difficult to assess under realistic evaluation settings.
2.5. Objectives and Research Questions
Based on the above-mentioned evaluation gaps for Streaming anomaly detection, the objective of this paper is to propose a comprehensive benchmark and experimental evaluation, aiming to answer the following research questions:
- (Q1)
Do TSAD methods significantly perform better in static settings rather than in online settings?
- (Q2)
Are streaming approaches more relevant than online approaches in streaming settings?
- (Q3)
Are streaming approaches more appropriate than online ones in case of concept drift?
- (Q4)
Do streaming methods balance accuracy and efficiency best?
3. StrAD: Our Proposed Benchmark
To address these questions, we introduce StrAD, a benchmark designed to enable a systematic evaluation of unsupervised and multivariate anomaly detection methods in streaming settings. StrAD evaluates SOTA streaming models and a broad spectrum of TSAD methods on real-world data. Overall StrAD contributions are:
| General | Drift | |||||
| Dataset | Category (Field) | # TS | # TS | ratio | Type | |
| Genesis | Sensor (Robotics) | 1 | 18 | 0/1 | 0% | |
| MITDB | Medical | 13 | 13 | 0/13 | 0% | |
| PSM | Facility | 1 | 25 | 0/1 | 0% | |
| SVDB | Medical | 31 | 2 | 0/31 | 0% | |
| MSL | Sensor (Aerospace) | 16 | 16 | 0/16 | 0% | |
| represented in TSB-drift | ||||||
| Daphnet | Human Activity | 1 | 9 | 1/1 | 11% | |
| GHL | Sensor (Industry) | 25 | 19 | 23/25 | 15% | |
| SMD | Facility | 22 | 38 | 15/22 | 13% | |
| LTDB | Medical | 5 | 2 | 1/5 | 67% | |
| TAO | Environment | 13 | 3 | 8/13 | 54% | |
| OPP. | Human Activity | 8 | 248 | 7/8 | 23% | |
| CreditCard | Finance | 1 | 29 | 1/1 | 3% | |
| CATSv2 | Sensor (Dynamic System) | 6 | 17 | 5/6 | 20% | |
| SMAP | Sensor (Telemetry) | 27 | 25 | 5/27 | 4% | |
| SWaT | Sensor (Cybersecurity) | 2 | 59 | 1/2 | 2% | |
| GECCO | Sensor (Water quality) | 1 | 9 | 1/1 | 11% | |
| Exathlon | Facility | 27 | 21 | 5/27 | 5% | |
3.1. TSB-drift: a Concept-Drift Subset
To the best of our knowledge, no current streaming benchmark explicitly evaluates performance of anomaly detection on real-world data with identified concept drift. To address this gap, we introduce TSB-drift, a curated subset of real-world time series exhibiting clear concept drifts. This subset is constructed from the TSB-AD benchmark (Liu and Paparrizos, 2024).
3.1.1. TSB-AD-M as a starting point
TSB-AD-M consists in 200 multivariate real-world time series from 17 different datasets as presented in Table 1. Those datasets are recurrent in TSAD (Wagner et al., 2023; Audibert et al., 2022; Sylligardos et al., 2025; Wu and Keogh, 2022) and were selected to avoid identified flaws in widely used datasets (Wu and Keogh, 2022; Liu and Paparrizos, 2024; Wagner et al., 2023). TSB-AD-M is divided into two subsets: TSB-AD-M-Tuning, which contains 20 time series dedicated to hyperparameter tuning, and TSB-AD-M-Eval, which includes the remaining 180 time series used for performance evaluation. Each time series is accompanied by a predefined training segment, ensuring reproducibility and fair comparison across methods. We use the latter as the initial batch for online and streaming methods.
3.1.2. Building TSB-drift
We derive TSB-drift from TSB-AD-M by quantifying distributional changes within each series. The goal is to identify time series with significant changes in at least one dimension, ensuring a high-confidence set of drifting series.
(Step a): Batch subdivision. Each multivariate time series is subdivided into consecutive batches of equal size, as illustrated in Figure 2(a). The batch size is the training size , so that for the -th dimension: , with This choice retains the training batch as a historical reference in a streaming context. It balances computational efficiency with the need to preserve the training batch as a historical reference in a streaming context.
(Step b): Measuring distributional change. To quantify distribution changes in the data distribution across batches, we use the Jensen-Shannon Divergence which is a symmetric and bounded variant of the Kullback-Leibler divergence . Formally, let and be two batches. We estimate their respective empirical distributions and by binning the data over the support and obtain as follows:
| (1) |
| (2) |
In the above equation, we have . By computing for every pair of batches across all dimensions (Figure 2(b)), we obtain divergence matrices .
(Step c): Aggregating across dimensions. The divergence information across dimensions is aggregated into a single drift matrix , where each cell is as shown in Figure 2(c). This max-pooling approach ensures that a significant drift occurring even in a single dimension is captured.
(Step d): Selecting series with strong drift. We finally rank series primarily by the maximum of and secondarily by the global mean of to favor long-lasting drifts. As shown in Figure 2(d), a subset of series displays substantial drifts. We retain the top 75 time series, forming a high-confidence set of series with significant drifts.
3.1.3. Drift Characterisation
The matrices obtained for each dimension of the time series allow us to identify recurrent drift patterns through the database as theoretically defined in (Gama et al., 2014; Vázquez et al., 2023). As illustrated in Table 2, diagonal heatmaps correspond to continuous drifts (C), block matrices to change points (CP) and cobbled heatmaps to periodic drifts (P). We name random-walks (RW) the drifts where we could find no distinctive patterns. Finally, dark heatmaps show the absence of drift.
3.1.4. Synthetic Validation
We validate our methodology using a controlled synthetic suite. We create multivariate time series containing each identified drift patterns, as well as a No Drift baseline. We also implement a Virtual Drift scenario to ensure that our approach distinguishes real concept drifts from simple noise fluctuations. More details are available in Appendix B.1. Our methodology successfully detect all drifts. Moreover, it correctly categorize the Virtual Drift as a low magnitude shift () that does not reach our selection values ().
| Matrix Pattern | Time Series | Concept Drift | Dataset (id TS, d) | Occurrence Examples |
|
|
Continuous (C) | SWaT id: 2 dim: AIT201 | Sensor Degradation |
|
|
Change Point (CP) | Exathlon id: 15 dim: 6 | Dynamic Allocation |
|
|
Periodic drift (P) | CATSv2 id: 1 dim: asin2 | Seasonality |
|
|
Random Walk (RW) | OPP. id: 7 dim: TAG2X | Physiological Variability |
|
|
No drift | MITDB id: 3 dim: V5 | Medical (ECG) |
3.2. TSAD methods in StrAD
StrAD incorporates a diverse set of time-series anomaly detection (TSAD) methods (Static and Online) and Streaming algorithms. This selection enables a systematic comparison across different modeling paradigms, computational profiles and adaptation strategies.
| Acronym | Method | Type | Complexity |
| Distance-based | |||
| LOF | LOF (Breunig et al., 2000) | Proximity | |
| KNN | -NN (Hawkins, 1980) | Proximity | |
| KMAD | -Means (Hawkins, 1980) | Clustering | |
| CBLOF | CBLOF (He et al., 2003) | Clustering | |
| Density-based | |||
| IF | Isolation Forest (Liu et al., 2008) | Tree | |
| MCD | MCD (Rousseeuw, 1984) | Distribution | |
| HBOS | HBOS (Goldstein and Dengel, 2012) | Distribution | |
| SVM | OCSVM (Schölkopf et al., 1999) | Distribution | |
| PCA | PCA (Shyu et al., 2003) | Encoding | |
| RPCA | RobustPCA (Paffenroth et al., 2018) | Encoding | |
| Prediction-based | |||
| CNN | CNN (Munir et al., 2019) | Forecasting | |
| LSTM | LSTMAD (Malhotra et al., 2015) | Forecasting | |
| AT | AnomalyTransformer (Xu et al., 2022) | Reconstruction | |
| AE | AutoEncoder (Sakurada and Yairi, 2014) | Reconstruction | |
| TrAD | TranAD (Tuli et al., 2022) | Reconstruction | |
| TN | TimesNet (Wu et al., 2023) | Reconstruction | |
| USAD | USAD (Audibert et al., 2020) | Reconstruction | |
| OA | OmniAnomaly (Su et al., 2019) | Reconstruction | |
| FITS | FITS (Xu et al., 2024) | Reconstruction | |
3.2.1. Static/Online TSAD
Table 3 shows our representative set of TSAD approaches. The selection balances well-established classical techniques with state-of-the-art deep learning models to compare diverse modeling paradigms and computational profiles.
More specifically, we include Proximity-based (LOF (Breunig et al., 2000), KNN
(Hawkins, 1980)), Clustering-based (KMeansAD (Hawkins, 1980), CBLOF (He et al., 2003)), Distribution-based (MCD (Rousseeuw, 1984), HBOS (Goldstein and Dengel, 2012), OCSVM (Schölkopf et al., 1999)), Encoding-based (PCA
(Aggarwal, 2015), RobustPCA (Paffenroth et al., 2018)) and Tree-based methods (IForest (Liu et al., 2008)). These approaches remain highly relevant in modern benchmarks (Liu and Paparrizos, 2024; Audibert et al., 2022; Wagner et al., 2023).
Moreover, their relatively low computational overhead make them particularly pertinent in streaming settings. As such, they provide essential reference points for assessing whether more elaborate models yield significant performance improvements.
The Prediction-based category covers a wide range of methods, including Recurrent and Convolutional methods (LSTMAD (Malhotra et al., 2015), CNN (Munir et al., 2019)), AutoEncoders (AE (Sakurada and Yairi, 2014), USAD (Audibert et al., 2020), OmniAnomaly (Su et al., 2019)), Transformers (Anomaly Transformer (Xu et al., 2022), TranAD (Tuli et al., 2022)) and two Frequency-based encoding (TimesNet (Wu et al., 2023), FITS (Xu et al., 2024)).
The complexity described in Table 3 refers to the inference complexity (i.e., score computation), where is the time series dimensionality, the time series length, and the window-size. Training cost is ignored since all models are trained offline. In streaming settings, inference latency is the most relevant metric for deployment.
| Acronym | Method | UM (Sec 2.2.2) | MM (Sec 2.2.3) | Complexity |
| Numerical | ||||
| LODA | LODA (Pevný, 2016) | Projections | Tumbling Window | |
| xS | xStream (Manzoor et al., 2018) | Projections | Tumbling Window | |
| RSH | RSHash (Sathe and Aggarwal, 2018) | Partitioning | Sliding Window (Point) | |
| HST | HSTree (Tan et al., 2011) | Partitioning | Tumbling Window | |
| SDOs | SDOstream (Hartl et al., 2020) | Partitioning | Soft Forgetting (Aging) | |
| Structural | ||||
| RRCF | RRCF (Guha et al., 2016) | Tree | Sliding Window (Point) | |
| MCOD | MCOD (Kontaki et al., 2011) | Clustering | Sliding Window (Point) | |
| LEAP | LEAP (Cao et al., 2014) | Proximity | Sliding Window (Batch) | |
| SKNN | SWKNN (Ramaswamy et al., 2000) | Proximity | Sliding Window (Point) | |
| MemS | MemStream (Bhatia et al., 2022) | Encoding | Soft Forgetting (Selective) | |
3.2.2. Streaming TSAD
The streaming methods selected in Table 4 represent the state-of-the-art in streaming TSAD, covering a wide range of adaptation and memory strategies.
The Numerical category comprises Projection-based (LODA (Pevný, 2016), xStream (Manzoor et al., 2018)) and Partitioning-based (RSHash (Sathe and Aggarwal, 2018), HSTree (Tan et al., 2011), SDOstream (Hartl et al., 2020)). These methods achieve high efficiency by maintaining simple statistics within predefined regions of the data space, making them particularly suitable for streaming scenarios.
The structural category features methods with
dynamically evolving internal representations.
This group includes Tree-based (RRCF
(Guha et al., 2016)), Clustering-based (MCOD (Kontaki et al., 2011)), Proximity-based (LEAP (Cao et al., 2014), SWKNN (Ramaswamy et al., 2000)) and Encoding-based mechanisms (MemStream (Bhatia et al., 2022)).
Finally, our selection also balances different memory management philosophies, ranging from rigid Tumbling and Sliding windows, to more nuanced forgetting mechanisms.
3.3. From TSAD to Streaming Settings
We aim to compare both Online and Streaming methods in the same benchmark. Thus, we define in this section a common training and normalization procedure.
3.3.1. Initial Batch
If needed, models in Table 3 and 4 are trained on an initial subset of the time series, considered as the historical data in a streaming context. The inference of the anomaly score will then be based on the incoming data, which, depending on the nature of the model, can be a point or a sliding window. There is no retraining for Online models, as opposed to Streaming models, which are based on intrinsic update and/or forgetting mechanism applicable on any incoming data.
3.3.2. Input Normalization
In this paper, we apply z-score standardization as a preprocessing if needed. The mean and standard deviation are computed exclusively from the initial batch and then applied to the incoming stream. The latter is chosen for two reasons: (i) It ensures model agnosticism: local normalization on a per-window basis lack the temporal context to calculate meaningful statistics. (ii) It prevents anomalies from skewing the normalization parameters, a common risk in window-based scaling.
4. Empirical Evaluation
We now discuss the experimental results addressing the research questions introduced in Section 2.5. We provide in Appendix A the hyper-parameter calibration. All implementations are available in our repository StrAD.
Implementation details. We use the implementation of the TSB-AD benchmark (Liu and Paparrizos, 2024) for the Static/Online methods, in Python. As for the Streaming methods, we use MCOD and LEAP’s original Java implementations, and the C++ implementations of xStream, SWKNN and SDOstream using the dSalmon (Hartl et al., 2024) library. RRCF, HSTree and RSHash are added using the Python library Pysad (Yilmaz and Kozat, 2020). Finally, we adapt MemStream’s (Bhatia et al., 2022) code to our pipeline.
Evaluation Measure. The choice of evaluation metric is critical, as many metrics have been shown to exhibit biases (Liu and Paparrizos, 2024; Sørbø and Ruocco, 2024; Wagner et al., 2023). This study focuses on parameter-free evaluation measures, such as the Area Under the Curve (AUC) and the Volume Under the Surface (VUS) (Boniol et al., 2025). Moreover, as we are not interested in detecting precursors but the anomalies themsleves, and since allowing delay tolerance after the true event contradicts the streaming purpose, VUS’s temporal tolerance is not the most appropriate for this study. Additionally, the Area Under a Receiver Operating Characteristic Curve (AUC-ROC) metric tends to produce overly optimistic results on highly imbalanced datasets (Sørbø and Ruocco, 2024), which is typically the case in anomaly detection. Consequently, the Area Under the Precision-Recall Curve (AUC-PR) is preferred in this study.
4.1. Static vs. Online TSAD models (Q1)
As shown in Figure 3, we start by comparing Static and Online methods. While the overall performances may appear similar in Figure 3 (1a), performances vary per method (Figure 3 (1b)). Figure 3 (2) breaks down the mean performance of the four best performing models in each data category of Table 1. We focus on the top performers to ensure a ’Best-in-Class’ comparison, effectively evaluating the theoretical upper bounds of each paradigm. We observe strong differences: environment data, which encompasses the TAO dataset, has a large amount of point anomalies throughout the time series. Local anomaly detection, enabled by the online context, better highlights the anomalies, and therefore explains the substantial gap in performance. In contrast, medical data exhibit high regularity and low noise levels. Static methods can leverage the entire time series, making them particularly effective for sequence anomalies.
We then measure in Figure 4 the performance shifts from Static to Online. We observe three groups: (i) significant improvement, (ii) stable performance, and (iii) performance drop. In the first group, in green, which goes from LOF (Figure 4 (2a)) to KNN, we find the Proximity-based methods, as well as the two Prediction-based methods that use Fourier transform to encode the incoming data. For Proximity-based methods, the Static environment often suffers from the multiple similar anomalies problem, where anomalies mask one another due to their global density. The local context provided by the Online environment eliminates this problem, resulting in substantial gains of performance. As for TimesNet and FITS, in static environment, the Fourier transform is calculated over the entire time series, meaning the spectral representation can be polluted by non-stationarity. By shifting to an online window, these models better capture the local periodicities, making them more sensitive to temporal deviations that would otherwise be diluted.
Amongst the methods that under-perform in Online settings, Density-based methods drop the most, such as IForest shown in Figure 4 (2c). In static context, these models benefit from a comprehensive representation of the data’s distribution. The Online setting freezes the distribution of the training batch as reference. With concept drift, this frozen representation becomes obsolete.
Finally, several models present no significant change. For deep learning-based methods (such as AutoEncoder in Figure 4 (2b)), the transition from static to Online is almost transparent, as their architecture is already window-based by design. Similarly, semi-supervised methods, such as OCSVM and MCD, use the training subset to define the boundaries of a normal space. Thus, Static and Online settings are essentially the same for these approaches.
4.2. Online vs. Streaming TSAD models (Q2)
We now compare accuracy performances of Online methods versus Streaming methods in streaming context. Figure 5 (1a) presents counter intuitive results: on average, Online methods significantly outperform Streaming methods. To ensure this performance gap is not merely due to the initial training batch size, we conduct an ablation study. We evaluate the performance across reduced training sizes of 75% and 50% of their original lengths. The results confirm that Online dominance is structural rather than data-dependent: Online methods maintain a significant lead across all settings, presenting on average a 64% gain at full training size, 49% gain at 75% training size, and 47% gain even when the training data is cut in half. More details are available in Appendix C. This assessment is reinforced when looking at the models’ individual performances in Figure 5 (1b). Only SWKNN is part of the top half of the benchmark models, and amongst the bottom ten, seven of those methods are Streaming. This performance disparity remains agnostic of the application domain, as shown in Figure 5 (2). The top performing Online methods beat or equal the top Streaming methods in every domain.
We propose several hypotheses to explain these findings. Our primary one lies with the architectural target of Streaming algorithms. As explained in Section 2.4, most of them are designed for point outlier detection, meaning they aim to identify isolated instances that fall outside a local density or distance threshold. Therefore, collective anomalies in real-world data are missed.
To assess the plausibility of this hypothesis, we obtain the critical difference diagrams of the four best models of each category on TSB-AD-M (Figure 5(3a)) and on a subset containing only point anomalies (Figure 5(3b)). Figure 5(3a) highlights a significant difference of performance between the best Online model, CNN, and the best Streaming model, SWKNN. Moreover, there are three Online models before SWKNN. Conversely, when it comes to detecting point outliers, Figure 5(3b) shows no significant difference between the eight methods. Streaming models are more competitive regarding point-wise anomalies than sequences anomalies. A second explanation for the difference between Streaming and Online is the model complexity. Streaming methods are by essence designed for efficient update, relying on lightweight, incremental algorithms. Conversely, some Online methods are based on high capacity architectures which can better capture subtle temporal dependencies.
4.3. Evaluation on TSB-drift (Q3)
In this section, we evaluate the robustness of Online and Streaming methods through drifts identified in TSB-drift. Figure 6 (1) presents the mean performance of every method on TSB-drift compared to TSB-AD-M without TSB-drift (written as TSB-drift in the rest of the paper). Methods below the identity line (highlighted in red) are less performant on TSB-drift. Nearness to this diagonal can serve as a proxy for drift resilience. In this regard, we observe that a large majority of the Streaming methods are amongst the closest to the diagonal. It is further highlighted in Figure 6 (1a) and (1b), presenting the mean loss when evaluated on TSB-drift compared to TSB-drift. 70% of the Streaming methods (ranging from LEAP to MemStream) show lower performance degradation than 70% of the Online methods (from FITS to USAD). And even the less robust Streaming model, HSTree, is less impacted than 7 Online models. It is further highlighted in Figure 6 (2), showing the mean performance of the top 4 models of each category (CNN, USAD, TimesNet and KNN for Online models, and SWKNN, MemStream, SDOstream and MCOD for Streaming ones) on the time series presenting each identified type of drift (presented in Section 3.1.3). We can see that in three out of the four types identified (change point, random walk and periodic), there is almost no performance gap. Interestingly, in the case of continuous drift, Online methods maintain a lead, though the gap is halved compared to the non-drifted time series. The resilience to continuous drifts for the top four Online models compared to their Streaming counterparts can be explained by the fact that it is a slow drift. If the models (in particular CNN, USAD and TimesNet) learn a temporal pattern rather than absolute values, they can still recognize the shape despite the drift.
Nevertheless, this robustness to drift does not translate into absolute superiority. The highest-performing methods remain overwhelmingly Online. The resilience shown by the Streaming methods does not bridge the substantial gap highlighted in Section 4.2.
4.4. Efficiency of Online vs. Streaming (Q4)
We now evaluate the computational efficiency of Online versus Streaming methods. Figure 7(1) depicts AUC-PR versus average throughput (average number of inference computations per second) for all time series in TSB-AD. Methods in the top right corner are slower but more accurate, while those in the bottom left corner are the fastest but show the worst accuracy. The top right corner is the ideal sector. We observe a Pareto frontier, highlighted by the red dotted line, that illustrates a fundamental Performance-Efficiency Trade-off. The Streaming models on the Pareto frontier (SDOstream and LEAP) confirm what has been previously underlined: by focusing on high temporal efficiency, those methods have sacrificed TSAD performance on complex data. Conversely, accurate methods are mainly Online methods with deep architectures, whose structural complexity limits iteration speed.
Figure 7(2) focuses on the Pareto frontier methods, depicting throughput versus the number of dimension. LEAP is represented by a discontinued line, because it does not return an anomaly score per point, but rather per small batch (of size 32). The bins in Figure 7(2) are not strictly equivalent (respectively 58, 45, 54 and 23 time series per bin). As expected, increasing dimensionality negatively impacts the computational efficiency of the methods. Despite this, even the slowest frontier method, CNN, maintains a mean throughput of 770 point scored per second. Thus, this method remains applicable for high-precision monitoring, like electrocardiogram (Goldberger et al., 2000), or industrial vibration monitoring, that can operate in the 500-800Hz (Lei et al., 2013). However in high velocity domains, like network intrusion detection (García-Teodoro et al., 2009), high-speed fiber optics links operate at Gbps rates, making such deep models wholly unsuitable.
Finally, we assess the runtime stability in Figure 7(3). We show two boxplots per model for the standard deviation of the inference time across TSB-drift and TSB-drift. The red full lines represent the median of all models on TSB-drift, while the red dotted line is on TSB-drift. Globally, Online methods exhibit higher stability by one order of magnitude. This is expected, as the updating process causes computational instability. Furthermore, concept drifts slightly exacerbate this instability for Streaming methods, as the two medians are almost equal for the Online models, but the TSB-drift median is distinctly higher for Streaming methods in Figure 7(3b).
From a practical deployment perspective, these findings outline clear architectural boundaries governed by operational constraints. While streaming methods offer minimal memory footprints and high throughput ideal for resource-constrained contexts or high-velocity data streams, their lightweight nature inherently limits capacity. Conversely, online methods trade higher computational and memory overhead for structural capacity, making them better suited for environments where prediction accuracy on complex, collective anomalies outweighs strict sub-millisecond latency bounds.
5. Conclusion
We now conclude on the key findings of our experimental evaluation on our novel benchmark StrAD and discuss the implications of our work for future research.
Why Do Online Methods Perform Better? Online methods mitigate anomaly camouflage inherent in static settings by limiting temporal context. This helps proximity-based models avoid multiple similar anomalies problem and allows Fourier-based models to focus on dominant frequencies without long-term non-stationarity. Additionally, their higher model capacity enables better temporal dependency modeling than lightweight Streaming approaches.
Limitations of Current Streaming TSAD. The most significant limitation is the legacy focus on point-wise outliers. Real-world anomalies often consist in collective anomalies, manifesting as pattern changes rather than simple value spikes. Most native Streaming algorithms are built on lightweight distance or density metrics designed to identify isolated points and therefore fail to capture such patterns.
Implications for Future Research. The observed performance gap indicates that native Streaming TSAD still underexplores collective anomalies. Therefore, in a streaming world, we should stand still. But for how long? The performance gap identified in our benchmark suggests that native Streaming TSAD still has the entire spectrum of collective anomalies left to explore. However, rather than merely simplifying deep models for incremental updates, a more promising frontier may lie in Streaming Automated Anomaly Detection. While AutoAD (Bahri et al., 2022; Liu et al., 2025; Sylligardos et al., 2025) has proven effective in Static settings by improving performance and easing model selection for non-experts, its extension to streaming scenarios remains largely unexplored. The latter is a promising research direction for Streaming TSAD.
Acknowledgements.
Funding This work was sponsored by the ANRT-CIFRE Grant No. 2025/0400 as part of the PhD thesis of the first author.References
- Outlier analysis. Springer. Cited by: §3.2.1.
- Anomaly detection in smart grids: a survey from cybersecurity perspective. In 3rd International Conference On Smart Grid And Renewable Energy (SGRE), United States (English). External Links: Document Cited by: §1.
- Detecting distance-based outliers in streams of data. In Proceedings of the Sixteenth ACM Conference on Information and Knowledge Management, CIKM 2007, Lisbon, Portugal, November 6-10, 2007, M. J. Silva, A. H. F. Laender, R. A. Baeza-Yates, D. L. McGuinness, B. Olstad, Ø. H. Olsen, and A. O. Falcão (Eds.), pp. 811–820. External Links: Link, Document Cited by: §1, §2.4.2.
- USAD: unsupervised anomaly detection on multivariate time series. In KDD ’20: The 26th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, Virtual Event, CA, USA, August 23-27, 2020, R. Gupta, Y. Liu, J. Tang, and B. A. Prakash (Eds.), pp. 3395–3404. External Links: Link, Document Cited by: §3.2.1, Table 3.
- Do deep neural networks contribute to multivariate time series anomaly detection?. Pattern Recognit. 132, pp. 108945. External Links: Link, Document Cited by: §3.1.1, §3.2.1.
- AutoML: state of the art with a focus on anomaly detection, challenges, and research directions. Int. J. Data Sci. Anal. 14 (2), pp. 113–126. External Links: Link, Document Cited by: §5.
- MemStream: memory-based streaming anomaly detection. In WWW ’22: The ACM Web Conference 2022, Virtual Event, Lyon, France, April 25 - 29, 2022, F. Laforest, R. Troncy, E. Simperl, D. Agarwal, A. Gionis, I. Herman, and L. Médini (Eds.), pp. 610–621. External Links: Link, Document Cited by: §1, §2.2.3, §3.2.2, Table 4, §4.
- VUS: effective and efficient accuracy measures for time-series anomaly detection. VLDB J. 34 (3), pp. 32. External Links: Link, Document Cited by: §4.
- Unsupervised and scalable subsequence anomaly detection in large data series. VLDB J. 30 (6), pp. 909–931. External Links: Link, Document Cited by: §2.1.1.
- Dive into time-series anomaly detection: A decade review. CoRR abs/2412.20512. External Links: Link, Document, 2412.20512 Cited by: §1, §2.1, §2.4.2.
- Series2Graph: graph-based subsequence anomaly detection for time series. Proc. VLDB Endow. 13 (11), pp. 1821–1834. External Links: Link Cited by: §2.1.2.
- SAND: streaming subsequence anomaly detection. Proc. VLDB Endow. 14 (10), pp. 1717–1729. External Links: Link, Document Cited by: §2.1.1.
- LOF: identifying density-based local outliers. In Proceedings of the 2000 ACM SIGMOD International Conference on Management of Data, May 16-18, 2000, Dallas, Texas, USA, W. Chen, J. F. Naughton, and P. A. Bernstein (Eds.), pp. 93–104. External Links: Link, Document Cited by: §2.1.1, §3.2.1, Table 3.
- Scalable distance-based outlier detection over high-volume data streams. In IEEE 30th International Conference on Data Engineering, Chicago, ICDE 2014, IL, USA, March 31 - April 4, 2014, I. F. Cruz, E. Ferrari, Y. Tao, E. Bertino, and G. Trajcevski (Eds.), pp. 76–87. External Links: Link, Document Cited by: §1, §2.2.3, §3.2.2, Table 4.
- Revisiting streaming anomaly detection: benchmark and evaluation. Artif. Intell. Rev. 58 (1), pp. 8. External Links: Link, Document Cited by: Appendix A, §1, §1, §2.2.1, §2.4.3.
- Anomaly detection: A survey. ACM Comput. Surv. 41 (3), pp. 15:1–15:58. External Links: Link, Document Cited by: §2.4.2.
- Online model-based anomaly detection in multivariate time series: taxonomy, survey, research challenges and future directions. CoRR abs/2408.03747. External Links: Link, Document, 2408.03747 Cited by: §2.3.
- A survey on concept drift adaptation. ACM Comput. Surv. 46 (4), pp. 44:1–44:37. External Links: Link, Document Cited by: §1, §2.2.1, §3.1.3.
- Anomaly-based network intrusion detection: techniques, systems and challenges. Computers and Security 28 (1), pp. 18–28. External Links: ISSN 0167-4048, Document, Link Cited by: §1, §4.4.
- PhysioBank, physiotoolkit, and physionet : components of a new research resource for complex physiologic signals. Circulation 101, pp. E215–20. External Links: Document Cited by: §1, §4.4.
- Histogram-based outlier score (hbos): a fast unsupervised anomaly detection algorithm. In KI, Cited by: §2.1.2, §3.2.1, Table 3.
- Robust random cut forest based anomaly detection on streams. In Proceedings of the 33nd International Conference on Machine Learning, ICML 2016, New York City, NY, USA, June 19-24, 2016, M. Balcan and K. Q. Weinberger (Eds.), JMLR Workshop and Conference Proceedings, Vol. 48, pp. 2712–2721. External Links: Link Cited by: §2.2.2, §2.2.3, §2.2, §3.2.2, Table 4.
- SDOstream: low-density models for streaming outlier detection. In 28th European Symposium on Artificial Neural Networks, Computational Intelligence and Machine Learning, ESANN 2020, Bruges, Belgium, October 2-4, 2020, pp. 661–666. External Links: Link Cited by: §1, §1, §2.2.2, §2.2.3, §2.4.2, §3.2.2, Table 4.
- DSalmon: high-speed anomaly detection for evolving multivariate data streams. pp. 153–169. External Links: ISBN 978-3-031-48884-9, Document Cited by: §4.
- Identification of outliers. Monographs on Applied Probability and Statistics, Springer. External Links: Link, Document, ISBN 978-94-015-3996-8 Cited by: §3.2.1, Table 3, Table 3.
- Discovering cluster-based local outliers. Pattern Recognit. Lett. 24 (9-10), pp. 1641–1650. External Links: Link, Document Cited by: §3.2.1, Table 3.
- HOT SAX: efficiently finding the most unusual time series subsequence. In Proceedings of the 5th IEEE International Conference on Data Mining (ICDM) 2005), 27-30 November 2005, Houston, Texas, USA, pp. 226–233. External Links: Link, Document Cited by: §1, §2.4.2.
- Continuous monitoring of distance-based outliers over data streams. In Proceedings of the 27th International Conference on Data Engineering, ICDE 2011, April 11-16, 2011, Hannover, Germany, S. Abiteboul, K. Böhm, C. Koch, and K. Tan (Eds.), pp. 135–146. External Links: Link, Document Cited by: §1, §1, §2.2.2, §2.2.3, §2.4.2, §3.2.2, Table 4.
- A review on empirical mode decomposition in fault diagnosis of rotating machinery. Mechanical Systems and Signal Processing 35 (1), pp. 108–126. External Links: ISSN 0888-3270, Document, Link Cited by: §1, §4.4.
- Isolation forest. In Proceedings of the 8th IEEE International Conference on Data Mining (ICDM 2008), December 15-19, 2008, Pisa, Italy, pp. 413–422. External Links: Link, Document Cited by: §2.1.2, §3.2.1, Table 3.
- TSB-autoad: towards automated solutions for time-series anomaly detection [e, A & B]. Proc. VLDB Endow. 18 (11), pp. 4364–4379. External Links: Link, Document Cited by: §5.
- The elephant in the room: towards A reliable time-series anomaly detection benchmark. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, A. Globersons, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. M. Tomczak, and C. Zhang (Eds.), External Links: Link Cited by: Appendix A, §1, §1, §3.1.1, §3.1, §3.2.1, §4, §4.
- Long short term memory networks for anomaly detection in time series. In 23rd European Symposium on Artificial Neural Networks, ESANN 2015, Bruges, Belgium, April 22-24, 2015, External Links: Link Cited by: §2.1.3, §3.2.1, Table 3.
- XStream: outlier detection in feature-evolving data streams. In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, KDD 2018, London, UK, August 19-23, 2018, Y. Guo and F. Farooq (Eds.), pp. 1963–1972. External Links: Link, Document Cited by: §1, §1, §2.2.2, §2.2.3, §2.4.2, §3.2.2, Table 4.
- DeepAnT: A deep learning approach for unsupervised anomaly detection in time series. IEEE Access 7, pp. 1991–2005. External Links: Link, Document Cited by: §2.1.3, §3.2.1, Table 3.
- A meta-level analysis of online anomaly detectors. VLDB J. 32 (4), pp. 845–886. External Links: Link, Document Cited by: Appendix A, §1.
- Robust PCA for anomaly detection in cyber networks. CoRR abs/1801.01571. External Links: Link, 1801.01571 Cited by: §3.2.1, Table 3.
- TSB-UAD: an end-to-end benchmark suite for univariate time-series anomaly detection. Proc. VLDB Endow. 15 (8), pp. 1697–1711. External Links: Link, Document Cited by: §1.
- Loda: lightweight on-line detector of anomalies. Mach. Learn. 102 (2), pp. 275–304. External Links: Link, Document Cited by: §2.2.2, §2.2, §3.2.2, Table 4.
- Efficient algorithms for mining outliers from large data sets. In Proceedings of the 2000 ACM SIGMOD International Conference on Management of Data, May 16-18, 2000, Dallas, Texas, USA, W. Chen, J. F. Naughton, and P. A. Bernstein (Eds.), pp. 427–438. External Links: Link, Document Cited by: §1, §2.2.3, §2.4.2, §3.2.2, Table 4.
- Least median of squares regression. Journal of the American Statistical Association 79 (388), pp. 871–880. External Links: Document Cited by: §3.2.1, Table 3.
- Anomaly detection using autoencoders with nonlinear dimensionality reduction. In Proceedings of the MLSDA 2014 2nd Workshop on Machine Learning for Sensory Data Analysis, Gold Coast, Australia, QLD, Australia, December 2, 2014, A. Rahman, J. D. Deng, and J. Li (Eds.), pp. 4. External Links: Link, Document Cited by: §2.1.3, §3.2.1, Table 3.
- Scalable and accurate online multivariate anomaly detection. Inf. Syst. 131, pp. 102524. External Links: Link, Document Cited by: §1, §1, §2.4.3.
- Subspace histograms for outlier detection in linear time. Knowl. Inf. Syst. 56 (3), pp. 691–715. External Links: Link, Document Cited by: §1, §1, §2.2.2, §2.2.3, §2.2, §2.4.2, §3.2.2, Table 4.
- Support vector method for novelty detection. In Advances in Neural Information Processing Systems 12, [NIPS Conference, Denver, Colorado, USA, November 29 - December 4, 1999], S. A. Solla, T. K. Leen, and K. Müller (Eds.), pp. 582–588. External Links: Link Cited by: §3.2.1, Table 3.
- Time series anomaly discovery with grammar-based compression. In Proceedings of the 18th International Conference on Extending Database Technology, EDBT 2015, Brussels, Belgium, March 23-27, 2015, G. Alonso, F. Geerts, L. Popa, P. Barceló, J. Teubner, M. Ugarte, J. V. den Bussche, and J. Paredaens (Eds.), pp. 481–492. External Links: Link, Document Cited by: §2.1.1.
- A novel anomaly detection scheme based on principal component classifier. In Foundations and New Directions of Data Mining Workshop at ICDM 2003, Cited by: Table 3.
- Robust anomaly detection for multivariate time series through stochastic recurrent neural network. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, KDD 2019, Anchorage, AK, USA, August 4-8, 2019, A. Teredesai, V. Kumar, Y. Li, R. Rosales, E. Terzi, and G. Karypis (Eds.), pp. 2828–2837. External Links: Link, Document Cited by: §3.2.1, Table 3.
- MSAD: A deep dive into model selection for time series anomaly detection. VLDB J. 34 (6), pp. 72. External Links: Link, Document Cited by: §3.1.1, §5.
- Navigating the metric maze: a taxonomy of evaluation metrics for anomaly detection in time series. Data Min. Knowl. Discov. 38 (3), pp. 1027–1068. External Links: Link, Document Cited by: §4.
- Fast anomaly detection for streaming data. In IJCAI 2011, Proceedings of the 22nd International Joint Conference on Artificial Intelligence, Barcelona, Catalonia, Spain, July 16-22, 2011, T. Walsh (Ed.), pp. 1511–1516. External Links: Link, Document Cited by: §2.2.2, §2.2.3, §2.2, §3.2.2, Table 4.
- TranAD: deep transformer networks for anomaly detection in multivariate time series data. Proc. VLDB Endow. 15 (6), pp. 1201–1214. External Links: Link, Document Cited by: §3.2.1, Table 3.
- Anomaly detection in streaming data: A comparison and evaluation study. Expert Syst. Appl. 233, pp. 120994. External Links: Link, Document Cited by: Appendix A, §1, §1, §1, §2.2.1, §2.4.1, §2.4.3, §3.1.3.
- TimeSeAD: benchmarking deep multivariate time-series anomaly detection. Trans. Mach. Learn. Res. 2023. External Links: Link Cited by: §3.1.1, §3.2.1, §4.
- TimesNet: temporal 2d-variation modeling for general time series analysis. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023, External Links: Link Cited by: §3.2.1, Table 3.
- Current time series anomaly detection benchmarks are flawed and are creating the illusion of progress (extended abstract). In 38th IEEE International Conference on Data Engineering, ICDE 2022, Kuala Lumpur, Malaysia, May 9-12, 2022, pp. 1479–1480. External Links: Link, Document Cited by: §1, §3.1.1.
- Anomaly transformer: time series anomaly detection with association discrepancy. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022, External Links: Link Cited by: §3.2.1, Table 3.
- FITS: modeling time series with 10k parameters. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024, External Links: Link Cited by: §3.2.1, Table 3.
- Disk aware discord discovery: finding unusual time series in terabyte sized datasets. Knowl. Inf. Syst. 17 (2), pp. 241–262. External Links: Link, Document Cited by: §2.1.1.
- Matrix profile I: all pairs similarity joins for time series: A unifying view that includes motifs, discords and shapelets. In IEEE 16th International Conference on Data Mining, ICDM 2016, December 12-15, 2016, Barcelona, Spain, F. Bonchi, J. Domingo-Ferrer, R. Baeza-Yates, Z. Zhou, and X. Wu (Eds.), pp. 1317–1322. External Links: Link, Document Cited by: §2.1.1.
- PySAD: A streaming anomaly detection framework in python. CoRR abs/2009.02572. External Links: Link, 2009.02572 Cited by: §4.
Appendix A Hyperparameters Calibration
For reproducibility purposes, we keep the hyperparameters proposed by the TSB-AD (Liu and Paparrizos, 2024) benchmark for the static and online configurations. The parameters search ranges for the streaming implementations are determined from previous benchmarks (Cao et al., 2025; Vázquez et al., 2023; Ntroumpogiannis et al., 2023) and recommendations from the original papers. Grid search is done on the Tuning subset of TSB-AD-M and the tuned hyperparameters are chosen based on the best median performance. We present below the search grid in Table 5 and tuned hyperparameters obtained for this study in Table 6.
| Model | Hyperparameter | Search Range |
| RRCF | num_trees | {2, 4, 8, 16, 20, 25} |
| shingle_size | {2, 4, 8} | |
| tree_size | {256, 512} | |
| RSHash | sampling_points | {500, 1000} |
| decay | {0.01, 0.015, 0.02} | |
| num_components | {50, 100, 200, 300} | |
| num_hash_fns | {1, 2, 4} | |
| HSTree | window_size | {50, 100, 200, 250} |
| num_trees | {10, 25, 50, 100} | |
| max_depth | {10, 15} | |
| MemStream | memory_len | {32, 64, 256, 512, 1024} |
| beta | {10, 1, 0.1, 0.01, 0.001} | |
| LEAP | k | {5, 10, 20, 40} |
| R | {0.5, 1.0, 2.0} | |
| slidingWindow | {256, 512} | |
| slide | {32, 64} | |
| MCOD | k | {5, 10, 20, 40} |
| R | {0.5, 1.0, 2.0} | |
| W | {64, 256, 512, 1024} | |
| xStream | window | {64, 128, 256} |
| n_estimators | {50, 100, 200} | |
| n_projections | {50, 100, 200} | |
| depth | {10, 15, 20} | |
| SWKNN | slidingWindow | {64, 128, 256, 512, 1024} |
| k | {5, 10, 15, 20, 30, 40} | |
| SDOstream | k | {200, 500, 1000, 1500} |
| T | {128, 256, 512} | |
| x | {3, 6, 10} |
| Model | Tuned hyperparameters |
| RRCF | {num_trees : 16, shingle_size: 8, tree_size: 512} |
| RSHash | {sampling_points : 1000, decay: 0.01, num_components: 300, num_hash_fns: 2} |
| HSTree | {window_size : 200, num_trees: 25, max_depth: 15} |
| MemStream | {memory_len : 32, beta: 0.001} |
| LEAP | {k : 40, R: 2.0, slidingWindow: 256, slide: 32} |
| MCOD | {k : 5, R: 1.0, W: 1024} |
| xStream | {window : 256, n_estimators: 200, n_projections: 50, depth: 20} |
| SWKNN | {slidingWindow : 64, k: 40} |
| SDOstream | {k : 1000, T: 512, x: 10} |
Appendix B More on TSB-drift
B.1. Synthetic Drift Validation
We create a synthetic suite, designed to manifest Real Concept Drift (where the relationship changes) by defining anomalies relative to an evolving local distribution. The base signal is with and sparse spikes the labeled anomalies (for , where the length of the time series). We validated our methodology across the following scenarios:
- •
Continuous Drift: Linear mean evolution, .
- •
Change Point: An abrupt regime shift at midpoint, if else .
- •
Periodic Drift: Toggling distribution parameters every steps, if is odd else .
- •
Random Walk: Stochastic non-stationarity, .
- •
Virtual Drift (Noise Shift): To ensure our threshold distinguishes between CD and simple noise increases, we simulated a variance shift: at ( constant).
Plots and codes are available on our repository.
B.2. Complex Drift Characterisation
| Matrix Pattern | Time Series | Concept Drift | Dataset (id TS, d) |
|
|
(P) and (CP) | SMD id: 19 dim: 5 |
|
|
Consecutive (CP) | SMD id: 19 dim: 29 |
|
|
Consecutive (P) | SMD id: 22 dim: 18 |
Of the identified patterns illustrated in Table 2, we can ensue more complex combinations that are still visually identifiable with the matrix patterns such as those identified in Table 7. Among those, we have the re-occurrence of the same pattern, like consecutive Change Points, or a change of periodicity. But we also have the succession of different patterns, such as the sudden interruption of a Periodic pattern.
Appendix C Ablation Study
We conduct this study on time series having at least a 1000 points in the initial training size, as some models cannot have an training batch size lower than 500 points, due to their hyperparameters. The study is therefore conducted on 142 times series, representing 15 out of the 17 datasets (the only time series from CreditCard and the 13 from TAO have an initial training batch size of 500 points and could not be included). The results are summarized in Table 8;
| Training Size | Mean Online | Mean Streaming | Relative Online Gain (%) |
| Full (100%) | 0.2807 | 0.1713 | +63.91 |
| 75% | 0.2742 | 0.1847 | +48.51 |
| 50% | 0.2701 | 0.1836 | +47.10 |