arXiv is now an independent nonprofit! Learn more
License: CC BY 4.0
arXiv:2609.39257v1 [cs.LG] 30 Sep 2026

From Benchmarks to Production:
Transferring Time Series Anomaly Detection Methods for Electricity Production Monitoring

Nicolas Vautier1, Paul Caron2, Nardi Xhepi2, Félicie Bizeul1, Manel Boumghar1, Christophe Degouy2, Paul Boniol3 Affiliation: 1 EDF Lab Paris Saclay, 2 EDF, DOAAT, 3 Inria, Ecole Normale Supérieure (PSL), CNRS
1,2 firstname.lastname@edf.fr, 3 paul.boniol@inria.fr
Abstract

Accurate forecasting of electricity production is essential for maintaining the operational efficiency and strategic planning of energy utilities. In industrial settings, such forecasts are generated daily to ensure supply–demand balance and optimal management of production assets. However, the increasing complexity of modern power systems and data flows poses significant challenges for ensuring the reliability and consistency of these forecasts. This paper addresses the problem of anomaly detection in short-term production forecasts at EDF, formulated as identifying atypical intra-day patterns that may signal data quality issues or operational irregularities. We introduce TAMIS, a scalable and interpretable system that analyzes daily production time series to automatically detect anomalous days based on deviations from historical patterns learned from past data. Designed for human-in-the-loop workflows, TAMIS surfaces top-ranked anomalies through an automated daily newsletter, enabling efficient expert review and continuous monitoring. An extensive experimental evaluation on real-world industrial data demonstrates that TAMIS achieves the best accuracy–efficiency trade-off compared to baseline methods. To foster further research and reproducibility, we publicly release the anonymized application datasets used in our study.

I Introduction

Time series data play a central role in a wide range of industrial applications [1, 2, 3], from predictive maintenance and fault detection [4] to resource optimization and demand forecasting [5]. A common challenge across these domains is the reliable detection of unusual patterns or behaviors that deviate from historical data (most commonly called anomaly or outlier in the literature [6, 7, 8]). Anomaly detection in time series is particularly critical in environments where real-time decisions must be made based on automatically generated forecasts or measurements. Despite significant progress in machine learning and statistical methods for time series analysis [9, 10], the practical deployment of anomaly detection systems in industry remains difficult due to specific constraints that often arise in operational settings.

This paper is motivated by the real-world context of energy production planning at EDF, one of the largest electricity producers in Europe. In this setting, large volumes of heterogeneous time series are generated daily to forecast the operation of power plants and storage systems over a 24-hour horizon. Each forecast is represented as a fixed-length vector capturing 48 half-hourly values for the upcoming day. These forecasts are critical inputs to balancing supply and demand, planning energy dispatch, and engaging with energy markets. Ensuring their reliability is thus essential. However, due to the scale and complexity of the data, it is not feasible for domain experts to manually review all forecasts, which necessitates an automated, trustworthy anomaly detection solution. More precisely, our industrial constraints are as follows: First, anomaly detection must be online, in the sense that it evaluates whether the most recent day-ahead forecast is abnormal based on historical data up to that point. Second, the solution must be interpretable and human-centered, as detection results are analyzed daily by energy experts who require concise, actionable explanations. Third, the method must be scalable, capable of handling large volumes of time series, with low latency, and limited hardware resources. These constraints rule out many existing approaches that rely on deep learning models, exhaustive offline training.

Fig. 1: Examples of time series in our proposed JO dataset (anomalies are highlighted in red).

To address these challenges, we propose TAMIS, a lightweight and interpretable anomaly detection system designed for production-grade deployment. Each daily forecast is scored for abnormality using a feature-based approach. The results are automatically ranked and integrated into a daily report. Only the most abnormal cases are returned, allowing experts to review them and make decisions accordingly.

Beyond proposing a practical solution, this work also provides insights into the broader journey of moving from academic research on time series anomaly detection to an operational system. We discuss how methodological advances translate into real-world applicability, and how constraints such as hardware efficiency, scalability, and expert usability shape the design of a deployable framework. In this sense, the paper not only introduces a new system, but also illustrates the path from research concepts to an industry-ready solution.

Finally, we evaluate our system on a large, real-world dataset collected from EDF’s internal production forecasts over multiple years. This dataset is diverse in both physical nature and statistical behavior, reflecting the complexity of our use case. We perform a rigorous experimental study to assess the scalability and detection accuracy of different algorithmic design within our framework. Furthermore, we release a curated and anonymized version of this dataset as a public benchmark to encourage future research in time series anomaly detection. Overall, our contributions are as follows:

  • •

    Novel Industrial Architecture: We analyze the requirements of time series anomaly detection in our industrial setting and propose a novel architecture tailored to these practical operational constraints (Sec. II).

  • •

    Designed-for-Purpose Components: We introduce a lightweight detector, ASHES, and TAMISℱ\text{TAMIS}_{\mathcal{F}}, a feature set that forms a core component of our framework (Sec. V).

  • •

    An End-to-end Pipeline: We present TAMIS, a novel end-to-end anomaly detection framework for electrical production time series that balances interpretability, accuracy, and scalability (Sec. V).

  • •

    Open Data: We release three open-access datasets of real, labeled time series, which are among the first of their kind in the energy sector (cf. Sec. VI-A1).

  • •

    Extensive Evaluation: We conduct an extensive experimental evaluation demonstrating that TAMIS outperforms state-of-the-art baselines while maintaining a superior accuracy-efficiency trade-off (Sec. VI).

II Industrial context and problem formulation

Symbol Description
Δ∈ℕ\Delta\in\mathbb{N} Sampling interval (3030 minutes).
M∈ℕM\in\mathbb{N} Number of samples per day (M=48M=48).
𝑪d∈ℝM\boldsymbol{C}_{d}\in\mathbb{R}^{M} Daily window (24-hour profile) for day dd
cm∈ℝc_{m}\in\mathbb{R} mm-th value of 𝑪d\boldsymbol{C}_{d}
𝑭d∈ℝi\boldsymbol{F}_{d}\in\mathbb{R}^{i} Feature set for day dd
𝒇j∈ℝ\boldsymbol{f}_{j}\in\mathbb{R} jj-th value of 𝑭d\boldsymbol{F}_{d}
𝑻∈ℝn×M\boldsymbol{T}\in\mathbb{R}^{n\times M} Time series of nn consecutive daily windows
Hj∈ℝn{H_{j}}\in\mathbb{R}^{n} Historical values for the jj-th feature, for a given time series
n∈ℕn\in\mathbb{N} Number of days (i.e., windows) in the time series 𝑻\boldsymbol{T}.
i∈ℕi\in\mathbb{N} Number of features 𝑻\boldsymbol{T}.
𝒜\mathcal{A} Anomaly detection function.
DD A unique anomaly detection model (i.e., detector).
ℳ\mathcal{M} An ensembling or selection method.
S∈[0,1]S\in[0,1] Anomaly score.
TABLE I: Summary of main notations used in this paper.

EDF, France’s leading electricity producer, is committed to building a carbon-neutral energy future through a mix of production sources. A critical enabler of this mission is the ability to continuously balance energy supply and demand, a task that relies heavily on accurate forecasts of both production and consumption. Maintaining this balance in the face of dynamic operational conditions and unforeseen events is essential not only for risk mitigation but also for value optimization.

Time series data lie at the heart of this operational ecosystem. These series capture a wide variety of measurements, including electrical power output, reservoir levels, flow rates, turbine states, production targets, and derived variables such as adjustment margins. These series differ in scale, dynamics, semantics, and physical meaning, making manual inspection infeasible. Each day, large volumes of new data are ingested and logged, underscoring the need for automated, scalable monitoring tools to detect and flag abnormal behavior.

Figure 1 illustrates examples of time series and anomalies of interest in our study. Overall, time series from EDF datasets contains heterogeneous types of time series and anomalies. Figure 1(a) illustrates a global point anomaly (i.e., a single value that deviates from the global distribution of values within the time series), Figure 1(b) depicts a contextual point anomaly (i.e., a point that deviates from the distribution of values within a given window of the time series), and Figure 1(c) shows a Collective anomaly (i.e., abnormal sequence of values that, if taken independently, would be considered normal).

A key aspect of this work is the operationalization of the detection within a human-in-the-loop framework. To bridge the gap between model output and practical action, a mailing list is implemented. Each day, the system automatically scores and ranks all potential anomalies. Subsequently, a summary report is distributed to a list of experts. This report is intentionally concise, presenting only the highest-ranked anomalies to mitigate alert fatigue and focus expert analysis on the most pressing issues. Each entry in the report is intended to include key contextual information, such as the anomaly score and timestamp, enabling experts to efficiently triage, validate, and prioritize potential anomalies for in-depth investigation.

II-A Problem Definition

Refer to caption
Fig. 2: Raw-based Time series anomaly detection methods in the literature [9]

As mentioned in the previous section, our objective is to identify anomalies in time series, a well-studied problem in the literature for decades [9]. Anomaly detection in this context is framed as an online, last forecasted window detection task, where the objective is to assess whether the most recently observed forecasted value in a series deviates significantly from historical behavior. Unlike retrospective analyses (the most common tackled problem in the literature [9]), the goal is not to detect previously missed events but to identify anomalies as they emerge, thus enabling timely interventions. This real-time requirement aligns with EDF’s operational needs, where undetected anomalies in key indicators could cascade into significant planning errors or missed opportunities.

We call a window the 24-hour profile of a variable cc sampled every Δ=30\Delta=30 minutes. Formally, a window 𝑪d\boldsymbol{C}_{d} for a given day dd is defined as 𝑪d=(c1,…,cM)∈ℝM\boldsymbol{C}_{d}=\bigl(c_{1},\ldots,c_{M}\bigr)\in\mathbb{R}^{M}, where M=48M=48 and cmc_{m} denotes the measurements for the mm-th half-hour interval of day dd. Therefore, a time series TT is an ordered set of windows defined as follows:

𝑻=(𝑪d1,…,𝑪dn)∈ℝn×M\boldsymbol{T}=\bigl(\boldsymbol{C}_{d_{1}},\ldots,\boldsymbol{C}_{d_{n}}\bigr)\in\mathbb{R}^{n\times M}

with nn being the total number of days in the time series 𝑻\boldsymbol{T}. In practice, the daily measurements 𝑪d\boldsymbol{C}_{d} is forecasted from the previous day d−1d-1. The primary objective of this work is to perform anomaly detection on such a daily production composing a given time series 𝑻\boldsymbol{T}. An anomaly is defined as the overall intra-day sequence (i.e., window) of points of 𝑻\boldsymbol{T} itself, which deviates significantly from what is considered normal or expected behavior for EDF’s power generation system for a day. This normality is established based on patterns learned from historical data (including typical intra-day profiles for various day types). In practice, we consider 3 years of historical data. Formally, in our industrial context, we define an anomaly detection function as follows:

Problem 1 (Anomaly Detection).

We define an anomaly detection function 𝒜\mathcal{A} as 𝒜:ℝn×M×ℝM⟶[0,1]\mathcal{A}:\mathbb{R}^{n\times M}\times\mathbb{R}^{M}\longrightarrow[0,1], that takes as input a time series 𝐓∈ℝn×M\boldsymbol{T}\in\mathbb{R}^{n\times M} and the last day 𝐂dn∈ℝM\boldsymbol{C}_{d_{n}}\in\mathbb{R}^{M}, and returns a real-valued anomaly score SS assessing how anomalous the last day is, given the historical context. A higher score indicates a higher likelihood of the last day being anomalous.

The function 𝒜\mathcal{A} can either correspond to a unique anomaly detection model, denoted by DD, or to an automatic detection solution, denoted by ℳ\mathcal{M}, combining multiple detectors DD. We will describe in the following sections the distinction between a unique detector and an automatic solution. While our problem setting shares core principles with the broader literature on time series anomaly detection, such as the importance of temporal dependencies and seasonality, it presents several practical constraints limiting the applicability of a large panel of existing methods. These constraints are listed below.

  • •

    (C1) Scalability: The practical applicability of an anomaly detection system in an industrial context, such as managing energy grids, is heavily dependent on its computational performance. The system must exhibit low execution time to allow for rapid, proactive interventions.

  • •

    (C2) Limited hardware: While model training can leverage high-performance computing resources, operational deployment must conform to strict hardware constraints. In our context, anomaly detection must run in near-real-time on standard CPU-based infrastructure, without access to GPUs or specialized accelerators.

  • •

    (C3) Interpretability: In an operational context such as energy grid management, the value of an anomaly detection system is tightly coupled with its ability to produce interpretable outputs. Since detected anomalies must be reviewed and acted upon by domain experts (often under tight time constraints) models that function as black boxes are of limited utility.

  • •

    (C4) Data Diversity: In our use case, time series are highly heterogeneous in terms of provenance and structure. First, missing values may have semantic meaning. Rather than removing them, one needs to preserve them to maintain data integrity and retrieve relevant knowledge that could indicate potential anomalies. Then, time series provenance (i.e., from diverse domains and measurement types) might lead to severe Out-of-Distribution scenarios, both in terms of trends and anomaly types.

III Time Series Anomaly Detection:
Foundations and Background

In this section, we review existing time-series anomaly detection methods through the lens of the key operational constraints enumerated in the previous section.

III-A Raw-based time series anomaly detection

In recent years, significant research has been conducted in the field of time-series anomaly detection. Numerous studies and experimental benchmarks have been written to summarize and analyze state-of-the-art methods [11, 7, 12, 13, 14]. These comprehensive evaluations reveal a diverse landscape of methodologies (i.e., detectors DD), which can be broadly classified into several foundational families based on their core operating principles [9]. Figure 2 depicts a large panel of methods grouped in process-centric families according to [9].

One major family consists of prediction-based methods, which train a model on a given self-supervised prediction task. An observation is classified as anomalous if it significantly deviates from the model’s prediction, with the anomaly score derived from the prediction error. Such methods can be reconstruction-based [15] or forecasting-based [16].

A second prominent category is distance-based methods, which operate under the assumption that anomalous points are isolated from the bulk of the data. These techniques identify outliers by measuring the distance of a data point to its nearest neighbors, with large distances indicating abnormality [17].

Closely related are density-based methods, which use a representation of the time series (i.e., tree [18] or graph structure [6]) and compute an anomaly score based on some density criteria of the given representation (i.e., depth of the tree, or centrality in the graph).

While prediction-based methods are appealing due to their ability to model temporal dynamics, they often rely on complex architectures with high computational and memory demands, and may be ill-suited for specific use cases on high-velocity data with real-time deployment on constrained hardware requirements (i.e., failing to meet C1 and C2). Additionally, their black-box nature poses challenges for interpretability (i.e., not respecting C3). Conversely, raw-based approaches such as distance- and density-based methods may offer lighter computation, but they struggle to handle data quality issues like missing values (failing to meet C4) and often lack the contextual transparency needed for non time series practitioners (i.e., not fully respecting C3). These limitations underscore the need for methods explicitly designed with industrial constraints in mind.

III-B Lack of interpretability: Toward Feature-based approaches

Feature-based methods offer a promising avenue for addressing the industrial constraints of interpretability, scalability, and robustness to data quality issues. By transforming raw time series into structured feature vectors (often composed of statistical, structural, or domain-specific descriptors), these approaches enable the use of appropriate density-based algorithms in a lower-dimensional, and meaningful space. Unlike raw-based distance- or density-based methods, feature-based techniques allow better handling of missing values through feature design and provide clear interpretability by exposing which characteristics contribute to the anomaly. As such, they strike a balance between operational feasibility and detection performance, making them particularly well-suited for real-world deployment in our industrial contexts.

In practice, the feature-based methodology [19, 20, 21] entails a two-stage process that transforms the complex temporal problem into a more manageable static outlier-detection task. First, a raw time-series window is converted into a fixed-length vector of numerical features. These features can capture a wide range of properties, including basic statistics (mean, variance), frequency-domain information, and entropy measures. In the second stage, a classical outlier-detection algorithm is applied to the feature space to identify feature vectors that are abnormal relative to the rest of the population [22]. Common algorithms used for this purpose include Isolation Forest [18], Local Outlier Factor (LOF) [23], and One-Class SVM [24].

Refer to caption
Fig. 3: Example of TSFresh and Catch22 features on IOPS [25] time series.

Several tools have been proposed for automatic feature extraction. Among the most widely used are HCTSA [26], TSFresh [21], and Catch22 [27]. HCTSA (Highly Comparative Time Series Analysis) performs massive feature extraction, computing over 7,700 features for each time series. These features span statistical distribution measures (e.g., mean, skewness, outlier ratios), autocorrelation and spectral properties, entropy metrics, and model-based characteristics.

Fig. 4: Time series anomaly detection: from unique detector (a) to supervised ensembling (d).

TSFresh extracts 794 time series features by default and includes an automated feature selection process based on hypothesis testing. This framework supports both exploratory analysis and integration into operational pipelines, making it a flexible option for industrial applications.

Catch22 builds upon HCTSA by identifying a compact subset of 22 features that provide strong classification performance across 93 UCR/UEA datasets [28]. These features are selected to minimize redundancy and maximize diversity, but are specifically tuned for z-normalized time series (zero mean and unit variance) and classification tasks.

While each of these libraries provides valuable insights, they present specific limitations in our context: (i) HCTSA and TSFresh produce high-dimensional feature sets, which are impractical for direct use in real-time anomaly detection. Moreover, TSFresh’s selection is not tailored to anomaly-specific tasks, potentially misaligning with our deployment needs. (ii) Catch22, due to its compact size and design for rapid analysis, appears promising for initial exploration. However, its features were optimized for classification, not anomaly detection. Figure 3 depicts the TSFresh and Catch22 features computed on an IOPS [25] time series. TSfresh contains features that allow the detection of anomalies (labeled in red in Figure 3), while Catch22 features are not sufficient for detecting the same anomalies.

Given these constraints, no existing feature set is suitable for our industrial use case, underscoring the need for feature selection and evaluation tailored to our operational context.

III-C No Universal Best: Toward Automatic Solutions

Recent benchmarking studies have consistently shown that there is no single anomaly detection algorithm that performs best across highly heterogeneous collections of time series [11, 7, 29, 30]. Instead, the performance of individual methods varies significantly depending on the characteristics of the time series (e.g., stationarity, periodicity) and the nature of the anomalies (e.g., point-wise, contextual, or subsequence).

This observation is particularly important in our context, where we must monitor a diverse set of time series with varying behaviors and noise levels. Furthermore, while most research algorithms produce anomaly scores for each time point, our setting requires a single decision at the last timestamp, introducing an additional mismatch between standard benchmarks and operational needs.

To overcome these limitations, we use several automatic solutions (i.e., meta model and ensembling methods ℳ\mathcal{M}). Overall, two main strategies have been proposed: ensembling (supervised or unsupervised) and automatic model selection. Figure 4 illustrates these strategies. Ensembling combines the outputs of multiple detectors, typically by averaging or maximizing their scores, to improve robustness. These methods often outperform individual algorithms on public benchmarks [29], but at the cost of higher computational overhead, which is problematic for real-time applications.

On the other hand, recent AutoML-based approaches have explored model selection as an alternative. These methods aim to select the most appropriate detector for each time series by learning a meta-model that maps extracted features to the best-performing method [31, 32, 29]. While promising, these techniques suffer from limited generalization in out-of-distribution settings [29]. The latter is an essential concern for real-world deployment, where new time series may exhibit previously unseen patterns or behaviors.

Given these limitations, we need to adopt an ensembling strategy that prioritizes robustness in out-of-distribution settings while maintaining low computational cost.

IV Research Questions

In light of the industrial constraints identified and the literature review in the previous section, our proposed solution is guided by several key research questions:

  • •

    R1. How to combine detectors for optimal accuracy–efficiency trade-offs? While ensembling can enhance robustness and detection accuracy, it also increases computational cost (C1). Conversely, model selection offers greater efficiency but may suffer in Out-of-Distribution scenarios (C4).

  • •

    R2. Which feature set to use? A smaller feature set can improve scalability (C1, C2). But, small or non-dedicated features may negatively impact detection accuracy, reducing the system’s practical usefulness (C3, C4).

  • •

    R3. What is the impact of Out-of-Distribution? Identifying the correct automatic solution setting can keep execution time low (C1,C2) while providing greater robustness in Out-of-Distribution scenarios (C4).

These research questions directly inform the design of our TAMIS framework (detailed in the following section), ensuring that it balances operational feasibility with high detection performance. Overall, this study not only describes a data-based pipeline for anomaly detection in production but also presents a roadmap for moving from research findings in data mining and data management to production-based systems.

V Proposed Method: The TAMIS System

Refer to caption
Fig. 5: (1) Training and (2) production pipeline of the proposed solution TAMIS. The end-user interface is illustrated in (3).

In this section, we introduce TAMIS, a lightweight, interpretable anomaly detection system specifically designed for production-grade deployment. As shown in Figure 5, the method involves a training phase composed of several steps: starting from manually labeled datasets (step (a)), a novel feature set, TAMISF, is computed for all available time series (step (b)). Two complementary detectors are applied to these features (step (c)), and their outputs are combined through a supervised ensemble, referred to as SEASA (step (d)). Once trained, this supervised ensemble model is deployed for inference in production (Figure 5(2)).

An important practical consideration is that precomputed features from the previous day are stored in a dedicated database and reused when computing anomaly scores for the following day. This caching mechanism drastically reduces runtime and enables near real-time performance in production.

Finally, the most abnormal time series are presented to the end-user through a visual interface, shown in Figure 5(3). In the following sections, we detail the design of the proposed feature set TAMISF, the two individual detectors, and the supervised ensembling mechanism, SEASA.

V-A TAMISF: A Novel Feature Set

The proposed feature set, TAMISF, aims to capture the most informative and generalizable characteristics of industrial time series for anomaly detection. The feature extraction pipeline begins by aggregating raw measurements into daily windows, denoted as 𝑪d\boldsymbol{C}_{d}, from which a feature vector 𝑭d\boldsymbol{F}_{d} is computed. The final feature set results from the integration of two complementary design strategies: (i) an expert-based feature pool derived from domain knowledge, and (ii) a TSAD-based feature pool identified through synthetic data generation and supervised feature selection.

Feature Description Formula Complexity
Expert-based: Top-10 features from energy-production expert-knowledge
amplitude: maximum amplitude for a given window 𝑪d\boldsymbol{C}_{d}. max⁡(𝑪d)−min⁡(𝑪d)\max(\boldsymbol{C}_{d})-\min(\boldsymbol{C}_{d}) O(n)
val max: Maximum value in 𝑪d\boldsymbol{C}_{d}. max⁡(𝑪d)\max(\boldsymbol{C}_{d}) O(n)
val min: Minimum value in 𝑪d\boldsymbol{C}_{d}. min⁡(𝑪d)\min(\boldsymbol{C}_{d}) O(n)
val mean: Mean of 𝑪d\boldsymbol{C}_{d}. μ⁡(𝑪d)\mu(\boldsymbol{C}_{d}) O(n)
val std: Standard deviation of 𝑪d\boldsymbol{C}_{d}. σ⁡(𝑪d)\sigma(\boldsymbol{C}_{d}) O(n)
missing values: Count of missing (NaN) values in 𝑪d\boldsymbol{C}_{d}. ∑cm∈𝑪d𝕀⁡(cm)\sum_{c_{m}\in\boldsymbol{C}_{d}}\mathbb{I}(c_{m}) O(n)
diff max: Maximum consecutive difference in 𝑪d\boldsymbol{C}_{d}. maxcm∈𝑪d⁡|cm−cm−1|\max_{c_{m}\in\boldsymbol{C}_{d}}|c_{m}-c_{m-1}| O(n)
diff min: Minimum consecutive difference in 𝑪d\boldsymbol{C}_{d}. mincm∈𝑪d⁡|cm−cm−1|\min_{c_{m}\in\boldsymbol{C}_{d}}|c_{m}-c_{m-1}| O(n)
diff mean: Average consecutive differences in 𝑪d\boldsymbol{C}_{d}. 1|𝑪d|−1​∑cm∈𝑪d|cm−cm−1|\frac{1}{|\boldsymbol{C}_{d}|-1}\sum_{c_{m}\in\boldsymbol{C}_{d}}|c_{m}-c_{m-1}| O(n)
diff step1: difference between c0c_{0} and cM′c^{\prime}_{M}, first and last values of 𝑪d\boldsymbol{C}_{d} and 𝑪d−1\boldsymbol{C}_{d-1} respectively. c0−cM′c_{0}-c^{\prime}_{M} O(n)
TSAD-based: Top-10 features from a synthetic time series anomaly detection evaluation
sum reoccurring values: Sum of all values that appear more than once in 𝑪d\boldsymbol{C}_{d}. ∑v∈Vreocc.v\sum_{v\in V_{\text{reocc.}}}v O(n)
Benford correlation: Correlation with the expected Benford’s Law distribution of first digits. corr​(Pobs,PBenford)\text{corr}(P_{\text{obs}},P_{\text{Benford}}) O(n)
Fourier entropy: Shannon entropy of the Power Spectral Density (PSD), with 100 bins. −∑pilog2(pi)-\sum{p}_{i}\log_{2}({p}_{i}) O(n log⁡(n)\log(n))
longest strike above mean: Length of the longest consecutive run of values above the mean (μ\mu). Length of run>μ\text{Length of run}>\mu O(n)
mean change: Mean of the absolute differences between consecutive values. mean​(|Δ​𝑪d|)\text{mean}(|\Delta\boldsymbol{C}_{d}|) O(n)
last location of max: Relative index of the last occurrence of the maximum value. argmaxlast​(𝑪d)/N\text{argmax}_{\text{last}}(\boldsymbol{C}_{d})/N O(n)
variation coefficient: Ratio of the standard deviation (σ\sigma) to the mean (μ\mu). σ/μ\sigma/\mu O(n)
permutation entropy: Complexity measure based on ordinal patterns (dim=5, lag=1). −∑i∈{1,…,D!}pilog2(pi)-\sum_{i\in\{1,\ldots,D!\}}p_{i}\,\log_{2}(p_{i}) O(n)
has duplicate min: Binary indicator for whether the minimum value appears more than once. 𝕀⁡(count​(min⁡(𝑪d))>1)\mathbb{I}(\text{count}(\min(\boldsymbol{C}_{d}))>1) O(n)
longest strike below mean: Length of the longest consecutive run of values below the mean (μ\mu). Length of run<μ\text{Length of run}<\mu O(n)
TABLE II: TAMISF features. Each feature is computed for an entire window (i.e., day dd). Note: 𝕀⁡(p)\mathbb{I}(p) is the indicator function for NaN values. Δ​Cd\Delta{C}_{d} represents first-order differences. Vreocc.V_{\text{reocc.}} is the set of reoccurring values in CDC_{D}.

V-A1 Expert-Based Feature Pool

The initial stage of TAMISF construction relied on domain expertise from energy production specialists. A total of 36 candidate indicators were manually proposed to describe operational behavior and potential deviations. To isolate the most relevant predictors, a structured feature selection protocol was implemented, combining ablation analysis and individual feature evaluation.

Specifically, three model configurations were trained to assess each feature’s contribution: (i) A baseline model trained on the full feature set; (ii) An ablation model trained on all features except the one under consideration (leave-one-out); (iii) An individual model trained on the single feature alone.

Each configuration produced four recall-based performance curves, quantifying the ratio of correctly detected anomalies with respect to both the detection score and rank. By qualitatively and quantitatively analyzing these curves, the ten most informative and discriminative features were retained. These are reported in the expert-based section of Table II.

Nevertheless, two limitations arise with the approach mentioned above. First, since the selection relies on historically labeled anomalies, it may be biased toward previously observed anomaly types, potentially limiting its generalization to unseen fault patterns. Second, the anomaly distribution within the evaluation dataset may not reflect their true operational frequency, introducing sampling bias. As a result, the selected features might overemphasize frequently occurring anomalies at the expense of rare but critical ones.

V-A2 TSAD-Based Feature Pool

To address the scarcity of labeled data and improve cross-domain generalization, we developed a data-driven feature discovery pipeline inspired by time series anomaly detection (TSAD) research. This methodology converts the unsupervised detection problem into a supervised learning framework through synthetic data generation, followed by a two-stage feature selection process.

Synthetic Data Generation and Feature Engineering: A labeled corpus is first synthesized by injecting artificial anomalies into normal time series (from JO dataset described in Section VI-A1). Inspired by recent experimental evaluation studies [11, 7], these anomalies emulate representative fault types, including point perturbations, temporal shifts, and amplitude scaling. The resulting dataset provides explicit ground-truth labels for controlled feature selection. A high-dimensional feature space is then constructed by computing descriptive statistics and signal transformations over each daily window, integrating features from TSFresh and Catch22.

Two-Stage Supervised Feature Selection: From the collected features, we perform the following selection pipeline:

Relevance Filtering: Each feature is ranked according to its discriminative capacity for the synthetic labels, evaluated using p-value and Benjamini Hochberg procedure [33], ANOVA F-test [34], and gradient boosting importance scores [35].

Redundancy Removal: The top-ranked features are clustered based on pairwise correlation to remove redundant predictors. From each cluster, a single representative feature is retained by maximizing its target correlation (via ANOVA F-test score) while minimizing its average inter-feature correlation.

The ten features retained through this supervised selection process constitute the TSAD-based component of TAMISF, as detailed in Table II. We evaluate the relevance of TAMISF in Section VI-C. Note that the proposed feature set is not restricted to use within our anomaly detection framework, TAMIS. Our feature set can be integrated into any anomaly detection system, whether supervised or unsupervised. We evaluate the performance of traditional anomaly detection methods applied on TAMISF in Section VI-B.

V-B Anomaly Detectors Pool

This section describes the two anomaly detectors that compose the core of the TAMIS ensemble. The choice of using only two base detectors is deliberate, as it significantly reduces computational overhead while maintaining high detection accuracy and scalability in production environments.

Let HiH_{i} denote the set of historical feature values for a given time series, and let fi∈𝑭df_{i}\in\boldsymbol{F}_{d} represent the new observation under evaluation. Each detector outputs an anomaly score, which is subsequently combined by the supervised ensemble model described in Section V-C.

V-B1 KDE: Kernel Density Estimation-based anomaly detection

The first detector models the distribution of historical feature values using a univariate Gaussian kernel density estimator (KDE) [36]. Anomalies are identified as points with low estimated density under this model.

Formally, let d^Hi\hat{d}_{H_{i}} denote the kernel density estimate fitted on HiH_{i}. To ensure numerical stability and produce a dimensionless score, the estimated density d^Hi​(fi)\hat{d}_{H_{i}}(f_{i}) is rescaled by the empirical standard deviation of the historical data, σ^​(Hi)\hat{\sigma}(H_{i}). The resulting anomaly score is defined as follows:

DKDE​(fi,Hi)=1−d^Hi​(fi)​σ^​(Hi)\small D_{\text{KDE}}(f_{i},H_{i})=1-\hat{d}_{H_{i}}(f_{i})\,\hat{\sigma}(H_{i}) (1)

This normalization preserves the monotonic relationship between the anomaly score and the estimated density (i.e., higher scores indicate rarer events) while compensating for scale differences across heterogeneous time series. Because DKDED_{\text{KDE}} is a monotone decreasing function of d^Hi​(fi)\hat{d}_{H_{i}}(f_{i}), ranking by DKDED_{\text{KDE}} is equivalent to ranking by −d^Hi-\hat{d}_{H_{i}}.

V-B2 ASHES: Anomaly Scoring-based on Historical ExtremeS

The second detector, denoted ASHES (Anomaly Scoring based on Historical ExtremeS), quantifies how much a new observation fif_{i} extends beyond the historical data extremes. It combines two complementary components: an extremity rank (KK) and a relative amplitude ratio (rr), which are integrated into a single anomaly score.

Extremity Rank (KK): The extremity rank quantifies the position of fif_{i} within the ordered historical distribution. It is defined as the minimum rank of fif_{i} when inserted into HiH_{i} sorted in ascending and descending order. Formally:

K⁡(fi)=min⁡(rankasc​(fi,Hi),rankdesc​(fi,Hi))\small K(f_{i})=\min\big(\text{rank}_{\text{asc}}(f_{i},H_{i}),\,\text{rank}_{\text{desc}}(f_{i},H_{i})\big) (2)

A rank of K=1K=1 indicates that fif_{i} is a new global minimum or maximum. As a practical heuristic, when K>10K>10, the point is considered non-extreme and its anomaly score is set to zero.

Relative Amplitude Ratio (rr): To quantify the magnitude of deviation, we compute a relative amplitude ratio r⁡(fi)r(f_{i}) that measures how much fif_{i} expands the range of the filtered historical set. Let HikH_{i_{k}} denote HiH_{i} with the K−1K-1 largest and K−1K-1 smallest values removed, and let minK=min⁡(Hik)\min_{K}=\min(H_{i_{k}}), maxK=max⁡(Hik)\max_{K}=\max(H_{i_{k}}). Formally, r⁡(fi)r(f_{i}) is defined as follows:

r⁡(fi)=maxK−minKmax⁡(fi,maxK)−min⁡(fi,minK)\small r(f_{i})=\frac{\max_{K}-\min_{K}}{\max(f_{i},\max_{K})-\min(f_{i},\min_{K})} (3)

A value of rr close to 1 indicates that fif_{i} lies within the typical range of the filtered extrema, while smaller values correspond to more substantial deviations from the historical amplitude.

ASHES Anomaly Score: The anomaly score of ASHES integrates both KK and rr to capture the effect of extremity and deviation magnitude. Formally, it is computed as follows:

DASHES​(fi,Hi)=0.5K−1−r×0.5K\small D_{\text{ASHES}}(f_{i},H_{i})=0.5^{K-1}-r\times 0.5^{K} (4)

This exponentially weighted formulation ensures that points with low ranks (high extremity) and large deviations (small rr) receive higher anomaly scores. Consequently, ASHES effectively highlights rare and impactful events that substantially alter the temporal dynamics of the monitored time series.

Overall, while the KDE detector provides a smooth, probabilistic view of data rarity, ASHES focuses on extreme deviations. Their complementary nature (density modeling versus rank-based extremity analysis) motivates their joint use within the supervised ensemble framework of TAMIS.

It is important to note that standard methods like HBOS [37] or LOF [23] treat rarity and amplitude deviation equally. In our industrial context, a value can be rare without being operationally critical if its relative amplitude is low. ASHES is specifically designed to address this by weighing the Extremity Rank (KK) against the Relative Amplitude Ratio (rr), offering a nuanced detection capability that off-the-shelf algorithms lack.

V-C Supervised Ensemble Anomaly Scores Aggregation

The final stage of the proposed approach introduces SEASA (Supervised Ensemble Anomaly Scores Aggregation), a stacking-based ensemble model that integrates the outputs of multiple base anomaly detectors (i.e., KDE and ASHES) into a unified and more accurate prediction. This ensemble mechanism leverages supervised learning to automatically infer optimal weightings and interactions among detector outputs, thereby enhancing overall robustness.

V-C1 Base Anomaly Score Generation

For each daily time series window (i.e., day dd), a feature vector 𝑭d\boldsymbol{F}_{d} is computed using TAMISF feature set. This representation serves as input to the two independent detectors, KDE and ASHES, which each produce a scalar anomaly score reflecting the degree of abnormality for that day. As mentioned in the previous section, these scores capture complementary aspects of the data distribution. The resulting base scores constitute the input features for the SEASA meta-learning stage.

V-C2 Meta-Learner and Prediction Strategy

The second level of SEASA is a meta-learner ℳ\mathcal{M} designed to combine the raw anomaly scores from the base detectors into a final anomaly prediction. In practice, we employ an XGBoost classifier [35] as ℳ\mathcal{M}. Formally, given base detector outputs {DKDE​(fi),DASHES​(fi)}\{D_{\text{KDE}}(f_{i}),D_{\text{ASHES}}(f_{i})\} and corresponding labels yi∈{0,1}y_{i}\in\{0,1\}, the meta-learner learns a mapping as follows:

y^i=ℳ⁡(DKDE​(fi),DASHES​(fi))\small\hat{y}_{i}=\mathcal{M}\big(D_{\text{KDE}}(f_{i}),D_{\text{ASHES}}(f_{i})\big) (5)

where y^i\hat{y}_{i} denotes the final anomaly probability. This supervised aggregation strategy allows the ensemble to adaptively weight detector contributions according to their reliability across different anomaly types.

V-C3 Experimental setup

To obtain an unbiased prediction for every point in our dataset, we utilize a Leave-One-Out Cross-Validation (LOOCV) procedure. For each point being evaluated, the XGBoost model is trained on all other labeled days and then makes a prediction on the single point that was left out. This operation is repeated for the entire dataset, ensuring that every prediction is made on data not seen during the training of that specific model instance. In deployment, the SEASA model is trained once using the available labeled dataset and then applied to new daily measurements.

VI Experiments

This section presents a comprehensive evaluation of our proposed approach, TAMIS, which addresses the research questions outlined in Section IV, and is organized as follows:

  • •

    Overall evaluation: We begin by assessing the global performance of TAMIS (accuracy and throughput) compared to a comprehensive set of baselines. Within this overall evaluation, we analyze different strategies for combining anomaly detectors, focusing on whether it is more advantageous to learn a selection model or to use an ensemble approach. This experiment addresses R1.

  • •

    Impact of the feature set: We then investigate how different feature representations affect performance (i.e., accuracy and efficiency). We compare TAMISF, against two widely used alternatives, TSFresh [21] and Catch22 [27]. This experiment addresses R2.

  • •

    Out-of-distribution and transfer evaluation: Finally, we evaluate the ability of TAMIS to generalize across different types of energy production systems. This experiment addresses R3.

VI-A Experimental Setup

All experiments are conducted on a single node with two Intel Xeon Gold 6234 CPUs at 3.30 GHz. Each node has 16 physical cores and 384 GB of RAM. For reproducibility, we make our implementation publicly available 11 1 Code Available at: https://github.com/VautierNicolas/TAMIS.

VI-A1 Datasets

In our industrial context, time series are highly diverse in nature and come from two areas: thermal and hydraulic, from EDF’s thermal and hydraulic production facilities, respectively. This diversity leads to a wide variety of behaviors. Moreover, several manual annotation campaigns have been conducted to obtain high-quality labels, resulting in a total of 6,931 annotated days with 376 anomalies. We divide our benchmark in the three datasets described in Table III. Overall, these datasets are as follows:

  • •

    THERM dataset: contains time series and labeled windows from the thermal perimeter. More precisely, time series corresponds to (i) marginal costs, penalties, start-up, and operating costs (€MW−1\mathrm{MW}^{-1} or €); (ii) demand, reference, and optimized programs (MW\mathrm{MW}); (iii) operating points for thermal units (MW\mathrm{MW}, gradient, number of modulations); and (iv) gas volumes (m3\mathrm{m^{3}}). In total, this set contains 4515 individual sensors.

  • •

    HYDRAU dataset: contains time series and labeled windows from the hydraulic perimeter. More precisely, time series corresponds to (i) minimum and maximum volumes and trajectories of reservoirs (hm3\mathrm{hm^{3}}); (ii) minimum, maximum, and turbine flow rates of power plants (m3​s−1\mathrm{m^{3}\,s^{-1}}); and (iii) minimum, maximum, and scheduled power output of power plants (MW\mathrm{MW}). In total, this set contains 1740 individual sensors.

  • •

    JO dataset: contains both time series and labeled windows from thermal and hydraulic perimeters; the intersection of the JO, THERM, and HYDRAU is empty. The aim of JO is to represent the initial training database for monitoring systems at EDF. In total, this set contains 3198 individual sensors.

Characteristics JO HYDRAU THERM
Number of sensors 3198 1740 4515
Total number of measurements 219M 267M 103M
Number of Labeled days 5813 540 578
NaN ratio 20.3% 3.5% 27.9%
Total number of NaN 55334 640 7827
Number of Anomalies 307 57 12
TABLE III: Datasets characteristics. NaN ratio: number of labeled days with at least one NaN value.

While existing benchmarks [11, 30, 13] have driven significant progress, our dataset introduces distinct industrial challenges that are under-represented in the literature:

  • •

    Data Quality and Semantic NaNs: Unlike curated benchmarks where missing values are rare, cleaned, or artificially imputed, our dataset retains the original data quality issues inherent to production environments. As shown in Table III, missing values occur at a maximum of 27.9% of the labeled days.

  • •

    Forecast Validation Task: Most public benchmarks focus on monitoring continuous raw sensor streams for retrospective anomaly detection. In contrast, our dataset addresses the validation of short-term production forecasts. This task involves comparing a new value against historical profiles to anticipate future failures, a structure distinct from standard stream monitoring.

  • •

    Physical Heterogeneity: Spanning both thermal and hydraulic domains, our dataset covers a wider range of physical behaviors (e.g., flow rates, temperatures) and economic variables (e.g., marginal costs).

Finally, we release an anonymized version of our datasets 22 2 Datasets: https://github.com/VautierNicolas/TAMIS_dataset. To the best of our knowledge, this collection is the most extensive dataset of energy production time series for anomaly detection, making it a significant contribution of our work. This release aims to encourage future research in this direction.

VI-A2 Baselines

We compare our proposed approach, TAMIS, against several categories of existing methods. We begin by evaluating TAMIS against raw-based detectors, i.e., approaches that operate directly on subsequences of the time series of interest. Next, we compare TAMIS to feature-based detectors that use the same feature set as our solution. Finally, we benchmark TAMIS against a set of automatic solutions. Overall, the baselines considered are the following:

  • •

    Raw-based Detectors: We consider three representative methods: One-Class SVM [24] (OCSVM), Histogram-Based Outlier Score [37] (HBOS), and Local Outlier Factor [23] (LOF). These baselines are chosen based on the industrial constraints (cf. Sec II-A), and for their ability to operate on raw time series or extracted features.

  • •

    Feature-based Detectors: We employ the same three methods (OCSVM, HBOS, and LOF) applied to features. Additionally, we compare with KDE and our proposed feature-based method, ASHES (cf. Sec V-B). As each detector outputs a vector of scores (i.e., one for each feature), the best aggregation method, among min, max, mean, and the meta-learner ℳ\mathcal{M} (cf. Sec V-C), is applied.

  • •

    Automatic solutions: As individual detectors struggle to remain robust on heterogeneous datasets (cf. Sec III-C). We compare TAMIS with two automatic methods, including an unsupervised average ensemble (Avg Ens) and the model selection strategy (MS) introduced in [29].

Fig. 6: Precision-Recall curves of TAMIS against (a) detectors on raw data; (b) detectors on TAMISF feature set; (c) Automatic solutions (Model Selection and Ensembling). (d) Throughput versus AUC-PR for TAMIS against most competitive baselines.

None of the baselines natively handles NaNs. Thus, we impute missing values using both forward and backward fill methods. This strategy is motivated by three considerations: (i) forward/backward filling maintains local temporal consistency, leading to more reliable anomaly scores; (ii) propagating the last (or next) observed value retains the piecewise-constant nature of the signal; and (iii) the imputed values remain within the true range of the observed data.

Finally, we evaluate several variants of our proposed approach. First, we consider a version of TAMIS that incorporates a larger set of feature-based detectors (referred to as TAMIS (all features)), combining OCSVM, HBOS, LOF, KDE, and ASHES. We also examine a more comprehensive configuration combining both feature-based and raw-based detectors (denoted TAMIS (all features and raw)), which, on top of the previously mentioned detectors, includes HBOS, OCSVM, and LOF on raw subsequences.

VI-A3 Evaluation Measures

Anomalies being rare, the Area Under the ROC curve (AUC-ROC) tends to overestimate the accuracy of detectors [30]. We thus consider the area under the Precision-Recall curve (AUC-PR). Beyond accuracy, industrial constraints require low execution time. Thus, we use throughput to measure efficiency and scalability. Note that more recent and robust evaluation measures exist, such as Volume Under the Surface (VUS) [14]. Nevertheless, our problem is to determine whether a given day DD is abnormal (the score produced by our approach is a single value S∈ℝS\in\mathbb{R}). Therefore, our problem is not affected by potential misalignment, justifying the use of VUS.

Fig. 7: Comparision of AUC-PR versus throughput between TAMIS using TAMISF, Catch22 and TSFresh as feature sets. The throughput corresponds to (a) Features computation, (b) Detectors computation, and (c) weights inference.

VI-B Overall Evaluation

Methods JO HYDRAU THERM
Detectors on raw time series
HBOS 0.069 0.297 0.041
LOF 0.228 0.226 0.072
OCSVM 0.138 0.143 0.028
Detectors on (TAMISF features)
HBOS 0.656 0.596 0.191
OCSVM 0.579 0.527 0.156
LOF 0.684 0.523 0.452
KDE 0.764 0.554 0.462
ASHES 0.792 0.571 0.229
Automatic Solutions (on TAMISF features)
Average Ensemble 0.147 0.209 0.033
Model Selection 0.790 0.569 0.157
TAMIS (all features) 0.851 0.683 0.486
TAMIS (all features and raw) 0.856 0.691 0.522
TAMIS 0.837 0.585 0.555
TABLE IV: Accuracy (AUC-PR) of TAMIS and baselines applied on raw data and TAMISF features.

In this section, we evaluate the global performance of TAMIS relative to all baseline configurations. The results are summarized in Table IV and Figure 6.

Raw time series as input: The first block of Table IV demonstrates that directly applying detectors to raw time series leads to poor accuracy. Across all datasets, AUC-PR values of LOF (i.e., the most accurate raw-based detector) remain below 0.23 on JO and HYDRAU, and drop to approximately 0.07 on THERM. This confirms that unprocessed time series lack the discriminative structure required for effective anomaly detection in electrical production systems.

Feature-based detection: Performances increase when detectors operate on TAMISF. Among individual detectors, ASHES achieves the highest AUC-PR on JO (0.792), while maintaining strong results on HYDRAU (0.571) and THERM (0.229). This confirms that the TAMISF feature set provides a more informative representation than raw subsequences.

Automatic aggregation approaches: As shown in Table IV, supervised ensemble methods (i.e., SEASA) significantly outperform both unsupervised ensembles and supervised model selection. The largest ensemble (i.e., TAMIS (all features and raw)) achieves the best overall accuracy (AUC-PR = 0.856), surpassing the best individual detector (ASHES) by +6.40 points. These results highlight the strong complementarity between detectors and the benefits of SEASA.

Accuracy–throughput trade-off: However, Figure 6 reveals that this accuracy improvement comes at the cost of computational efficiency. The largest ensemble has the lowest throughput because of its high inference cost. In contrast, the proposed TAMIS (i.e., using only two detectors, namely KDE and ASHES) achieves an optimal balance between accuracy and scalability. Specifically, it delivers near-best accuracy while being 2.4× faster, establishing TAMIS as the best trade-off solution for large-scale deployment.

Answer to R1: Combining heterogeneous detectors improves anomaly detection accuracy, especially when integrating both raw-data-based and feature-based approaches. Supervised ensembles outperform unsupervised and selection-based strategies, while the proposed TAMIS configuration achieves the best accuracy–efficiency trade-off with only two detectors.
Fig. 8: Comparison of our proposed approach TAMIS versus baselines on different models in-distribution (trained and tested on JO dataset) and two out-of-distribution scenarios. Note that all f. and r. stands for all features and raw.

VI-C Evaluating the Relevance of the Feature Sets

Figure 7 presents the trade-off between accuracy (AUC-PR) and throughput across the three stages of the TAMIS pipeline: (a) feature computation, (b) detector execution, and (c) weight inference. Three feature sets are compared: TAMISF, Catch22, and TsFresh.

Features computation: Across all datasets, TAMISF provides the best trade-off between accuracy and efficiency. It reaches the highest AUC-PR values (≈\approx 0.8–0.9) while sustaining throughputs several times higher than TsFresh and comparable to Catch22. While Catch22 remains computationally lightweight, its AUC-PR saturates around 0.15, revealing limited expressiveness for anomaly detection in electrical production systems. Conversely, TsFresh reaches high accuracy, but at a prohibitively high computational cost.

Pipeline-level efficiency: Figure 7(b) and (c) further demonstrate that the efficiency advantage of TAMISF persists through both the detector computation and ensemble weight inference stages. When all processing stages are considered, TAMISF maintains high accuracy and throughput.

Answer to R2: TAMISF feature set achieves the best accuracy–throughput balance among existing alternatives. It offers near-state-of-the-art detection accuracy at a significantly lower computational cost.

VI-D Toward Production: an Out-of-Distribution test

To assess the robustness and generalization capability of the proposed approach, we compare its performance under in-distribution (ID) and out-of-distribution (OOD) conditions. A baseline model is trained and tested on JO, representing the ID setting, while OOD performance was evaluated by applying the same model to the HYDRAU and THERM datasets, which differ in operational regimes and signal characteristics. Figure 8 reports the AUC-PR results for three evaluation configurations: (i) in-distribution (trained and tested on JO), (ii) OOD 1 (JO →\rightarrow HYDRAU), and (iii) OOD 2 (JO →\rightarrow THERM). We include the best-performing individual detectors (KDE and ASHES) as baselines.

In-distribution performance. All methods perform strongly (AUC-PR ≈\approx 0.8–0.9) in ID settings, demonstrating that both individual detectors and TAMIS successfully capture the statistical structure of the in-distribution data. These results establish a reference for subsequent OOD comparisons.

Cross-domain transfer: JO →\rightarrow HYDRAU. In the first OOD scenario, performance decreases across all methods. KDE shows substantial degradation (AUC-PR ≈\approx 0.2), indicating sensitivity to distributional shifts. In contrast, TAMIS maintains a significantly higher AUC-PR (>0.5>0.5), underscoring its superior ability to generalize across domains with differing noise levels and dynamic properties. Note that in such a scenario, ASHES also shows strong performances, but still suffers from a larger drop than TAMIS between ID and OOD.

Cross-domain transfer: JO →\rightarrow THERM. The second transfer scenario, from JO to THERM, is more challenging. All models experience additional performance loss, but TAMIS again remains the most robust, consistently outperforming KDE and ASHES by a large margin. This result highlights that the ensemble and feature-aggregation mechanisms within TAMIS effectively mitigate domain-specific overfitting.

Relative degradation analysis. The right-hand plots in Figure 8 quantify the relative AUC-PR loss between ID and OOD conditions. TAMIS exhibits only a ∼\sim34% decrease when tested on THERM, whereas ASHES experiences a 53% drop. These findings confirm that the ensemble-based formulation of TAMIS acts as a regularizer, maintaining higher predictive stability under data distribution shifts.

Answer to R3: TAMIS exhibits strong robustness under out-of-distribution conditions, retaining higher accuracy than individual detectors while maintaining low computational cost. The latter confirms the applicability of TAMIS in production, where distribution drift and heterogeneity across time series are expected.

VII Operational Insights and Extensions

This section discusses the operational feedback obtained from applying our approach in practice and outlines extensions toward more general benchmarks.

VII-A Operational feedback

The deployment of our proposed approach, TAMIS, provided valuable insights into the human-in-the-loop requirements for industrial anomaly detection:

Mitigating Alert Fatigue: Before the introduction of our proposed approach, the main challenge was the volume of alarms. By introducing a ranking mechanism (based on a threshold strategy), TAMIS significantly reduced alert fatigue. Experts reported that the daily report enables more efficient triage by focusing only on the most critical deviations.

The Value of Interpretability: While deep learning approaches often act as black boxes, the choice of explicit features in T​A​M​I​SℱTAMIS_{\mathcal{F}} was critical for adoption. Operational teams validated that these features serve as immediate proxies for physical root causes, enabling faster decision-making.

False Positives and Distribution Shifts: The majority of false positives are triggered by Out-of-Distribution (OOD) events. However, our experiments in Figure 8 demonstrate that TAMIS effectively reduces these errors compared to individual detectors, showing a relative resilience that is crucial for maintaining trust over time.

From Generic to Tailored Monitoring: A common industrial hurdle is the cold start problem, where historical labels are unavailable for a new production unit. Our deployment strategy leverages TAMIS’s transferability to address this. The system is initially deployed with a generic model pre-trained on available corporate data. The visualization interface then serves a dual purpose: (i) it supports daily decision-making and (ii) acts as a data annotation tool. As experts validate or reject alerts in the daily reports, they progressively build a high-quality, domain-specific labeled dataset. This feedback loop allows TAMIS to be iteratively retrained, refining the system’s sensitivity to domain-specific characteristics.

VII-B Generalization on public Benchmarks

While TAMIS is designed to address specific industrial constraints (e.g., missing values, heterogeneity), it is crucial to verify that its performance is not solely due to the specificity of our use case and datasets. To assess the generalizability of our approach, we evaluated TAMIS on TSB-UAD [11], a comprehensive, heterogeneous benchmark across domains. We compared TAMIS against the strongest baselines identified in a recent evaluation of model selection for time series anomaly detection [29]. More specifically, we consider the Best Detector (i.e., NormA), the Unsupervised Average Ensemble, and the Best Model Selection (Best MS) strategy identified in [29].

Figure 9 reports the Volume Under the Surface (VUS-PR) accuracy [14] of TAMIS versus the baselines on TSB-UAD. We observe that TAMIS (VUS-PR ≈\approx 0.38) significantly outperforms the Average Ensemble and the Best Detector. Moreover, TAMIS outperforms the Best Model Selection strategy (VUS-PR≈\approx 0.35). This result confirms that TAMIS captures fundamental anomaly characteristics that generalize well to diverse domains beyond energy production.

Fig. 9: TAMIS vs Best Detector, Avg Ens, the Best Model Selection strategy and the theoretical Oracle on TSB-UAD [11].

VIII Conclusion

This paper introduces TAMIS, a lightweight and interpretable anomaly detection system tailored for large-scale industrial time-series monitoring. Through extensive experiments, we addressed three core research questions: detector combination, feature-space design, and cross-domain robustness. More specifically, we demonstrate that TAMIS, effectively balances accuracy and efficiency, with the compact two-detector configuration achieving the best trade-off for scalable deployment, while maintaining strong performances in out-of-distribution conditions. Moreover, the proposed feature set, TAMISF, outperforms existing alternatives such as Catch22 and TsFresh, delivering a superior accuracy–throughput balance. To encourage future research in this direction, we release an anonymized version of our benchmark. The latter is among the most extensive datasets of energy production time series for anomaly detection. Finally, future work will focus on: (i) extending TAMIS to online learning; (ii) adapting TAMIS to multivariate time series.

References

  • [1] T. Palpanas and V. Beckmann (2019) Report on the first and second interdisciplinary time series analysis workshop (itisa). SIGMOD Rec. 48 (3), pp. 36–40. Cited by: §I.
  • [2] K. Uehara and M. Shimada (2002) Extraction of primitive motion and discovery of association rules from human motion data. Progress in Discovery Science: Final Report of the Japanese Dicsovery Science Project, pp. 338–348. Cited by: §I.
  • [3] M. Bach-Andersen, B. Rømer-Odgaard, and O. Winther (2017) Flexible non-linear predictive models for large-scale wind turbine diagnostics. Wind Energy 20 (5), pp. 753–764. Cited by: §I.
  • [4] P. Boniol, M. Meftah, E. Remy, B. Didier, and T. Palpanas (2023) DCNN/dcam: anomaly precursors discovery in multivariate time series with deep convolutional neural networks. Data-Centric Engineering 4, pp. e30. External Links: Document Cited by: §I.
  • [5] A. Petralia, P. Boniol, P. Charpentier, and T. Palpanas (2025) Few Labels are All You Need: A Weakly Supervised Framework for Appliance Localization in Smart-Meter Series . In 2025 IEEE 41st International Conference on Data Engineering (ICDE), Vol. , Los Alamitos, CA, USA, pp. 4386–4399. External Links: ISSN 2375-026X, Document, Link Cited by: §I.
  • [6] P. Boniol and T. Palpanas (2020) Series2Graph: graph-based subsequence anomaly detection for time series. Proc. VLDB Endow. 13 (12), pp. 1821–1834. Cited by: §I, §III-A.
  • [7] S. Schmidl, P. Wenig, and T. Papenbrock (2022) Anomaly detection in time series: a comprehensive evaluation. Proc. VLDB Endow. 15 (9), pp. 1779–1797. External Links: ISSN 2150-8097, Link, Document Cited by: §I, §III-A, §III-C, §V-A2.
  • [8] P. Boniol, J. Paparrizos, T. Palpanas, and M. J. Franklin (2021) SAND: streaming subsequence anomaly detection. Proc. VLDB Endow. 14 (10), pp. 1717–1729. Cited by: §I.
  • [9] P. Boniol, Q. Liu, M. Huang, T. Palpanas, and J. Paparrizos (2024) Dive into time-series anomaly detection: a decade review. External Links: 2412.20512, Link Cited by: §I, Fig. 2, Fig. 2, §II-A, §III-A.
  • [10] J. Paparrizos, P. Boniol, Q. Liu, and T. Palpanas (2025) Advances in time-series anomaly detection: algorithms, benchmarks, and evaluation measures. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2, KDD ’25, New York, NY, USA, pp. 6151–6161. External Links: ISBN 9798400714542, Link, Document Cited by: §I.
  • [11] J. Paparrizos, Y. Kang, P. Boniol, R. S. Tsay, T. Palpanas, and M. J. Franklin (2022) TSB-uad: an end-to-end benchmark suite for univariate time-series anomaly detection. Proc. VLDB Endow. 15 (8), pp. 1697–1711. External Links: ISSN 2150-8097, Link, Document Cited by: §III-A, §III-C, §V-A2, §VI-A1, Fig. 9, Fig. 9, §VII-B.
  • [12] P. Boniol, J. Paparrizos, Y. Kang, T. Palpanas, R. S. Tsay, A. J. Elmore, and M. J. Franklin (2022) Theseus: navigating the labyrinth of time-series anomaly detection. Proc. VLDB Endow. 15 (12), pp. 3702–3705. External Links: ISSN 2150-8097, Link, Document Cited by: §III-A.
  • [13] P. Wenig, S. Schmidl, and T. Papenbrock (2022) TimeEval: a benchmarking toolkit for time series anomaly detection algorithms. Proc. VLDB Endow. 15 (12), pp. 3678–3681. External Links: ISSN 2150-8097, Link, Document Cited by: §III-A, §VI-A1.
  • [14] J. Paparrizos, P. Boniol, T. Palpanas, R. S. Tsay, A. Elmore, and M. J. Franklin (2022) Volume under the surface: a new accuracy evaluation measure for time-series anomaly detection. Proc. VLDB Endow. 15 (11), pp. 2774–2787. External Links: ISSN 2150-8097, Link, Document Cited by: §III-A, §VI-A3, §VII-B.
  • [15] M. Sakurada and T. Yairi (2014) Anomaly detection using autoencoders with nonlinear dimensionality reduction. In Proceedings of the MLSDA 2014 2nd Workshop on Machine Learning for Sensory Data Analysis, MLSDA’14, pp. 4–11. Cited by: §III-A.
  • [16] P. Malhotra, L. Vig, G. M. Shroff, and P. Agarwal (2015) Long Short Term Memory Networks for Anomaly Detection in Time Series. In ESANN, Cited by: §III-A.
  • [17] E. M. Knorr and R. T. Ng (1998) Algorithms for mining distance-based outliers in large datasets. In Proceedings of the 24rd International Conference on Very Large Data Bases, VLDB ’98, San Francisco, CA, USA, pp. 392–403. External Links: ISBN 1558605665 Cited by: §III-A.
  • [18] F. T. Liu, K. M. Ting, and Z. Zhou (2008) Isolation forest. In 2008 Eighth IEEE International Conference on Data Mining, Vol. , pp. 413–422. External Links: Document Cited by: §III-A, §III-B.
  • [19] S. Tafazoli, Y. Lu, R. Wu, T. V. A. Srinivas, H. Dela Cruz, R. Mercer, and E. Keogh (2024) C22MP: the marriage of catch22 and the matrix profile creates a fast, efficient and interpretable anomaly detector. Knowl. Inf. Syst. 66 (8), pp. 4789–4823. External Links: ISSN 0219-1377, Link, Document Cited by: §III-B.
  • [20] A. Bonifati, F. D. Buono, F. Guerra, and D. Tiano (2022) Time2Feat: learning interpretable representations for multivariate time series clustering. Proc. VLDB Endow. 16 (2), pp. 193–201. Cited by: §III-B.
  • [21] M. Christ, N. Braun, J. Neuffer, and A. W. Kempa-Liehr (2018) Time series feature extraction on basis of scalable hypothesis tests (tsfresh – a python package). Neurocomput. 307 (C), pp. 72–77. External Links: ISSN 0925-2312, Link, Document Cited by: §III-B, §III-B, 2nd item.
  • [22] R. Chalapathy and S. Chawla (2019) Deep learning for anomaly detection: A survey. CoRR abs/1901.03407. External Links: Link, 1901.03407 Cited by: §III-B.
  • [23] M. M. Breunig, H. Kriegel, R. T. Ng, and J. Sander (2000) LOF: identifying density-based local outliers. SIGMOD Rec. 29 (2), pp. 93–104. External Links: ISSN 0163-5808, Link, Document Cited by: §III-B, §V-B2, 1st item.
  • [24] B. Schölkopf, R. Williamson, A. Smola, J. Shawe-Taylor, and J. Platt (1999) Support vector method for novelty detection. In Proceedings of the 13th International Conference on Neural Information Processing Systems, NIPS’99, Cambridge, MA, USA, pp. 582–588. Cited by: §III-B, 1st item.
  • [25] http://iops.ai/dataset_detail/?id=10. Cited by: Fig. 3, Fig. 3, §III-B.
  • [26] B. Fulcher and N. Jones (2017) Hctsa : a computational framework for automated time-series phenotyping using massive feature extraction. Cell Systems 5, pp. . External Links: Document Cited by: §III-B.
  • [27] C. H. Lubba, S. S. Sethi, P. Knaute, S. R. Schultz, B. D. Fulcher, and N. S. Jones (2019) Catch22: canonical time-series characteristics: selected through highly comparative time-series analysis. Data Min. Knowl. Discov. 33 (6), pp. 1821–1852. External Links: ISSN 1384-5810, Link, Document Cited by: §III-B, 2nd item.
  • [28] H. A. Dau, A. Bagnall, K. Kamgar, C. M. Yeh, Y. Zhu, S. Gharghabi, C. A. Ratanamahatana, and E. Keogh (2019) The ucr time series archive. External Links: 1810.07758, Link Cited by: §III-B.
  • [29] E. Sylligardos, P. Boniol, J. Paparrizos, P. Trahanias, and T. Palpanas (2023) Choose wisely: an extensive evaluation of model selection for anomaly detection in time series. Proc. VLDB Endow. 16 (11), pp. 3418–3432. External Links: ISSN 2150-8097, Link, Document Cited by: §III-C, §III-C, §III-C, 3rd item, §VII-B.
  • [30] Q. Liu and J. Paparrizos (2025) The elephant in the room: towards a reliable time-series anomaly detection benchmark. In Proceedings of the 38th International Conference on Neural Information Processing Systems, NIPS ’24, Red Hook, NY, USA. External Links: ISBN 9798331314385 Cited by: §III-C, §VI-A1, §VI-A3.
  • [31] Y. Zhao, R. A. Rossi, and L. Akoglu (2021) Automating outlier detection via meta-learning. External Links: 2009.10606, Link Cited by: §III-C.
  • [32] M. Goswami, C. Challu, L. Callot, L. Minorics, and A. Kan (2023) Unsupervised model selection for time-series anomaly detection. External Links: 2210.01078, Link Cited by: §III-C.
  • [33] Y. Benjamini and Y. Hochberg (1995) Controlling the false discovery rate: a practical and powerful approach to multiple testing. Journal of the Royal Statistical Society. Series B (Methodological) 57 (1), pp. 289–300. External Links: ISSN 00359246, Link Cited by: §V-A2.
  • [34] R. A. Fisher (1970) Statistical methods for research workers. In Breakthroughs in statistics: Methodology and distribution, pp. 66–70. Cited by: §V-A2.
  • [35] T. Chen and C. Guestrin (2016) XGBoost: a scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, KDD ’16, pp. 785–794. External Links: Link, Document Cited by: §V-A2, §V-C2.
  • [36] E. Parzen (1962) On Estimation of a Probability Density Function and Mode. The Annals of Mathematical Statistics 33 (3), pp. 1065 – 1076. External Links: Document, Link Cited by: §V-B1.
  • [37] M. Goldstein and A. R. Dengel (2012) Histogram-based outlier score (hbos): a fast unsupervised anomaly detection algorithm. Cited by: §V-B2, 1st item.