arXiv is now an independent nonprofit! Learn more
License: CC BY 4.0
arXiv:2609.04749v1 [cs.AI] 04 Sep 2026

DCFA: Dual-view Causal-inspired Attribution for Failure Reasoning in LLM-based Multi-agent Systems

Zehao Wang Affiliation: College of Intelligence and Computing, Tianjin University, Tianjin, China Affiliation: School of New Media and Communication, Tianjin University, Tianjin, China    Lanjun Wang ††thanks: ˜˜Corresponding Author. Affiliation: School of New Media and Communication, Tianjin University, Tianjin, China    Shilong Jin Affiliation: College of Intelligence and Computing, Tianjin University, Tianjin, China Affiliation: School of New Media and Communication, Tianjin University, Tianjin, China    Junjie Chen Affiliation: College of Intelligence and Computing, Tianjin University, Tianjin, China    Yanghua Xiao Email: {wzhrslh, jinshilong_2025}@gmail.com{wanglanjun, junjiechen}@tju.edu.cn, shawyh@fudan.edu.cn Affiliation: College of Computer Science and Artificial Intelligence, Fudan University, Shanghai, China Affiliation:  Shanghai Key Laboratory of Data Science, Shanghai, China
Abstract

Large language model (LLM)-based multi-agent systems have experienced rapid growth in recent years. Despite their promise, such systems remain fragile, frequently exhibiting reasoning and coordination errors that can lead to system-level failures. Failure attribution in such systems relies on tracing natural language interactions among agents to identify the decisive error, which refers to the earliest action whose correction can reverse system failure. There are two key challenges: 1) Shallow attribution: Existing methods often capture only minor deviations, such as incomplete retrievals or formatting errors, which verification mechanisms can correct, while missing the decisive cause of system failure. 2) Contextual degradation: As the length of the system traces increases, the model’s reasoning ability rapidly deteriorates. To address these challenges, we propose DCFA, a training-free framework for failure attribution. DCFA integrates a global module that constructs structured causal-inspired dependency graphs from system traces to identify the initial decisive error, and a local module that applies local counterfactual-inspired reasoning to refine causal-inspired attribution. Experiments on the Who&When benchmark across six LLMs show that DCFA improves step-level accuracy by up to 8.27% over state-of-the-art baselines. Code is available in the public repository (https://github.com/wzhSteve/DCFA).

Refer to caption
Figure 1: Illustration of failure attribution in an LLM-based multi-agent system. The system trace is represented as a textual record composed of interactive messages generated by agents.

1 Introduction

With the emergence of large language models (LLMs), multi-agent systems (MAS) built upon LLM backends have rapidly evolved Becattini et al. (2025); Ronanki (2025). These systems enable autonomous collaboration among specialized agents for tasks such as code generation, knowledge search, and data analysis Huang et al. (2024); Pan et al. (2024); Li et al. (2024); Majdoub et al. (2025). Despite the impressive progress in orchestration and tool integration, LLM-based multi-agent systems remain susceptible to coordination and reasoning failures Cemri et al. (2026); Hammond et al. (2025); Shen et al. (2026). Errors such as misalignment in agents’ collaborations or improper tool usage can propagate across different agents along the system trace, leading to cascading reasoning errors that ultimately derail entire workflows Cemri et al. (2026); He et al. (2025). This fragility highlights the urgent need for failure attribution in LLM-based multi-agent systems Cemri et al. (2026); Hammond et al. (2025), which refers to the process of identifying the earliest actions whose correction could reverse system failure, known as decisive errors.

Recent efforts have begun to explore failure attribution in LLM-based multi-agent systems using LLMs as diagnostic detectors Zhang et al. (2025); Zhang et al. (2026). Existing approaches can be broadly classified into fine-tuning-based Zhang et al. (2026) and instruction-based Zhang et al. (2025); Cemri et al. (2026) paradigms. The fine-tuning-based approach, exemplified by AgenTracer Zhang et al. (2026), constructs specialized datasets to train dedicated models for failure analysis. In contrast, instruction-based methods Zhang et al. (2025); Cemri et al. (2026) rely on carefully designed prompts to guide LLMs in analyzing multi-agent system traces and identifying the decisive errors. However, both paradigms face inherent limitations. Fine-tuning-based approaches offer strong controllability but require extensive annotation and repeated training, resulting in high costs. Instruction-based methods avoid these costs under a training-free constraint by leveraging LLMs’ general reasoning ability, yet they often fail to reliably interpret long, multi-step, and unstructured system traces, leading to inaccurate failure attribution.

In this work, we study failure attribution in LLM-based multi-agent systems under the training-free setting to improve the extraction of decisive errors from unstructured multi-step traces. We identify two fundamental challenges: 1) Shallow attribution: Existing LLM-based attribution methods often focus on minor, recoverable deviations (e.g., incomplete retrieval or formatting errors Cemri et al. (2026); Banerjee et al. (2025)) that do not determine the final outcome. In contrast, the decisive error refers to the earliest action whose correction would reverse the system failure, truly driving the system to fail. For example, an agent may fabricate missing content when critical information is lost during the initial collection or retrieval process. Unlike missing information that can be subsequently verified and retrieved through downstream interactions, the fabrication error is difficult to verify and thus persists throughout the reasoning chain, misleading the system’s subsequent decisions. 2) Contextual degradation: As trace length and complexity increase, attribution performance degrades sharply. For instance, the Who&When benchmark Zhang et al. (2025) reports that the step-level attribution accuracy of state-of-the-art methods decreases substantially with growing trace length with averaging around 15% for traces of length 10, but dropping to nearly 0% for those approaching length 100.

To overcome these challenges, we propose a training-free framework, named the Dual-view Causal-inspired Failure Attribution (DCFA). DCFA comprises two key modules: the Global Causal-inspired Attribution (GCA) module and the Local Counterfactual-inspired Enhancement (LCE) module, which jointly perform global reasoning over a causal-inspired dependency structure and local counterfactual-inspired validation. To move attribution beyond surface-level deviations, GCA constructs a structured causal-inspired dependency graph and performs reasoning along it to uncover deeper dependency relationships underlying system failure. GCA extracts candidate minor deviations from the interaction system trace and builds a step-level causal-inspired dependency graph centered on these candidates. By performing global reasoning over the causal-inspired dependency graph across the entire system trace, GCA generates an initial hypothesis of the decisive error. To mitigate long-context degradation, LCE refines the global hypothesis through localized counterfactual-inspired reasoning. Starting from the hypothesized error, LCE performs a bidirectional search over the causal-inspired dependency graph to identify local dependency chains through which the error propagates, and evaluates how local corrections propagate through the reasoning chain and affect the final outcome via counterfactual-inspired evaluation. By evaluating these locally induced outcome changes, LCE progressively identifies the interaction that most plausibly constitutes the decisive error. In summary, our contributions are fourfold:

  • •

    We propose DCFA, a training-free framework for failure attribution in LLM-based multi-agent systems, integrating global reasoning with local counterfactual-inspired validation.

  • •

    To address shallow attribution, we construct causal-inspired dependency graphs to model interaction-level dependencies, enabling LLMs to move beyond minor deviations and identify decisive errors.

  • •

    To mitigate contextual degradation, we introduce localized counterfactual-inspired refinement that focuses reasoning on critical dependency chains, thereby improving attribution accuracy on long traces.

  • •

    Experiments on the Who&When benchmark with six LLMs show consistent gains, improving step-level accuracy by up to 8.27% over the state-of-the-art baselines.

2 Related Works

2.1 LLM-based Multi-Agent Systems and System Failure

Recent advances in LLM-based agents have driven rapid progress in multi-agent systems (MAS), enabling complex, long-horizon tasks through task decomposition and inter-agent communication Yang et al. (2026); Sun et al. (2025). Prior work has proposed systematic taxonomies of agent architectures Wu et al. (2023); Fourney et al. (2024), identifying components such as role assignment, planning, external memory, and feedback mechanisms, which aim to bolster the problem-solving capabilities of LLM-based multi-agent systems.

Despite their promise, LLM-based multi-agent systems remain brittle Zhang et al. (2025); Cemri et al. (2026). Empirical studies reveal frequent coordination failures and cascading reasoning errors, particularly in long or highly interactive execution traces Liu et al. (2024b). These observations highlight the need for methods to analyze the decisive error of system failure.

2.2 Failure Attribution in LLM-based Multi-agent Systems

Failure attribution in multi-agent systems aims to identify the earliest mistake whose correction could reverse the overall system failure, referred to as the decisive error Cemri et al. (2026); Fourney et al. (2024). Unlike prior work such as AgentErrorBench with AgentDebug Zhu et al. (2025), which locates mistakes that trigger cascading reasoning distortions, or TRAIL Deshpande et al. (2025), which identifies and analyzes all errors in trajectories, focusing on the decisive error directly targets the most actionable point for repairing task failure.

Recent approaches employ LLMs directly as diagnostic tools, including fine-tuning-based methods Zhang et al. (2026) and instruction-based methods Zhang et al. (2025); Cemri et al. (2026). Fine-tuned models offer strong control but require costly annotation and retraining. Instruction-based methods instead analyze execution traces via prompting. For instance, ECHO Banerjee et al. (2025) is equipped with hierarchical context representations and multi-perspective consensus to improve attribution performance, while A2P West et al. (2025) is designed to perform causal-guided attribution through prompt-based instruction.

Despite their effectiveness, existing methods rely heavily on LLM intrinsic reasoning, often resulting in shallow attribution and degraded performance on long traces. To address these limitations, we propose a unified framework that combines structured reasoning over a causal-inspired dependency graph with localized dependency-chain analysis to improve failure attribution.

2.3 Causal Reasoning and Counterfactual Analysis with LLMs

Recent work has explored LLMs for causal reasoning and counterfactual-inspired analysis, showing that they can extract events, propose candidate causal relations, and simulate hypothetical interventions from unstructured text Cheng et al. (2025); Luo et al. (2024). To improve robustness, several studies combine LLM-derived causal hypotheses with algorithmic validation or graph-based aggregation Tong et al. (2024); Ban et al. (2025). Others treat LLMs as informative priors for causal-inspired dependency graph discovery, leveraging their encoded commonsense knowledge to guide structure learning Darvariu et al. (2024); Jiralerspong et al. (). Counterfactual simulation has also been used to probe the sensitivity and faithfulness of LLM reasoning chains Tutek et al. (2025); Yu et al. (2025). These advances motivate our use of causal-inspired scaffolding and localized counterfactual-inspired validation for failure attribution in LLM-based multi-agent systems. Discussions of related work are deferred to Appendix J.

3 Problem Formulation

To respond to a user query, a multi-agent system generates a system trace including a sequence of interactions denoted as 𝒯={τ1,τ2,…,τn}\mathcal{T}=\{\tau_{1},\tau_{2},\dots,\tau_{n}\}, where each interaction τt=(at,ct)\tau_{t}=(a_{t},c_{t}) consists of the agent identifier ata_{t} and its corresponding content ctc_{t}. The MAS produces a final output Y^\widehat{Y}, which is expected to match the ground-truth outcome YY. A system failure occurs when the generated output deviates from the expected one, i.e., Y^≠Y\widehat{Y}\neq Y. Such failures may arise from one or multiple erroneous or misleading interactions within 𝒯\mathcal{T}.

The objective of failure attribution is to identify the decisive error responsible for the system failure. When a single erroneous interaction accounts for the failure, the decisive error corresponds to the interaction whose correction can recover the correct outcome. When multiple errors jointly contribute to the failure and no single correction fully recovers the correct outcome, we define the decisive error as the interaction whose correction yields the largest reduction in the system failure.

Formally, a failure attribution model ℳ\mathcal{M} is defined as:

ℳ:(𝒯,Y^)↦τ∗,\mathcal{M}:(\mathcal{T},\widehat{Y})\mapsto\tau^{*}, (1)

where τ∗\tau^{*} denotes the ground-truth decisive error, i.e., the interaction whose correction either recovers the correct outcome when a single-step correction suffices, or maximally reduces system failure when multiple errors jointly contribute to the failure.

Refer to caption
Figure 2: Overview of DCFA. DCFA integrates global causal-inspired attribution and local counterfactual-inspired enhancement to identify and refine the decisive error by reasoning over a causal-inspired dependency graph constructed from the system trace.

4 Methodology

In this section, we introduce the proposed DCFA framework (Fig. 2), which consists of the GCA (Sec. 4.1) and LCE (Sec. 4.2) modules. To overcome shallow attribution, GCA reasons over full system traces to uncover latent dependency relationships, constructing a causal-inspired dependency graph and identifying an initial hypothesis of the decisive error from a global perspective. Then, to further mitigate contextual degradation, the LCE module locally refines this hypothesis through counterfactual-inspired evaluation, using LLM-mediated approximate interventions to assess how correcting candidate interactions may alter the final outcome. This strategy first establishes a global causal-inspired dependency structure and then locally validates its critical links, enabling more interpretable failure attribution for unstructured and long system traces.

4.1 Global Causal-inspired Attribution Module

Interactions in system traces influence system outcomes through chained dependencies, requiring attribution methods that capture both minor deviations and how their effects propagate through the trace. Directly analyzing unstructured traces, LLMs often overemphasize salient deviations without sufficiently considering their downstream dependencies and influence on the final outcome. To address this, GCA detects candidate deviations, constructs a directed causal-inspired dependency graph that links them to the full trace, and leverages graph-aware global reasoning to generate an initial hypothesis of the decisive error.

4.1.1 Minor Deviation Identification

Minor deviations are directly observable but typically benign faults in system traces, such as malformed outputs or incomplete information retrieval Cemri et al. (2026); Hammond et al. (2025). These deviations are often recoverable through built-in verification, tool re-invocation Barta et al. (2025); Zhou et al. (2025). However, minor deviations are not necessarily decisive errors, where they may serve as early symptoms of deeper failures.

Based on this motivation, GCA first extracts a set of candidate minor deviations that represent plausible manifestations of incorrect behaviors. Rather than prematurely determining which interaction constitutes the decisive error, GCA aims to retain all potentially relevant interactions for subsequent structured reasoning and refinement over the causal-inspired dependency graph.

Specifically, a designed process MDI​(⋅)\text{MDI}(\cdot) is utilized to detect the interactions of minor deviations and produce short natural-language explanations for each identified minor deviation:

(𝒯e,ℛe)=MDI​(𝒯),(\mathcal{T}^{e},\mathcal{R}^{e})=\text{MDI}\big(\mathcal{T}\big), (2)

where 𝒯e={τ1e,τ2e,…}\mathcal{T}^{e}=\{\tau^{e}_{1},\tau^{e}_{2},\dots\} is the set of candidate interactions that are identified as minor deviations, ℛe={r1e,r2e,…}\mathcal{R}^{e}=\{r_{1}^{e},r_{2}^{e},\dots\} are the textual reasons associated with each candidate.

4.1.2 Dependency Graph Construction

Having identified the candidate minor deviations and their corresponding textual reasons, GCA constructs a structured causal-inspired dependency graph 𝒢\mathcal{G} that encodes plausible dependency relationships among interactions in the system trace 𝒯\mathcal{T}. The construction is centered on the detected minor deviations 𝒯e\mathcal{T}^{e} and their associated reasoning cues ℛe\mathcal{R}^{e}, ensuring that the resulting structure reflects interpretable, error-centric reasoning pathways.

Formally, the causal-inspired dependency graph is denoted as 𝒢=(𝒱,ℰ)\mathcal{G}=(\mathcal{V},\mathcal{E}), where 𝒱\mathcal{V} is the set of step indicators tt corresponding to all interactions, and ℰ\mathcal{E} is the set of directed edges representing putative dependency relationships. Specifically, an edge e∈ℰe\in\mathcal{E} is denoted as (p,q)(p,q), indicating a directed dependency from interaction τp\tau_{p} to τq\tau_{q}, where τp\tau_{p} precedes and potentially influences τq\tau_{q}.

Dependency Condition Evaluation.

A directed edge e:τp→τqe:\tau_{p}\rightarrow\tau_{q} is admitted into ℰ\mathcal{E} only if it satisfies a set of structured dependency criteria Zeng et al. (2026); Goldberg (2019). For each ordered pair (τp,τq)(\tau_{p},\tau_{q}) with step indices p<qp<q, GCA evaluates three complementary properties: temporality, necessity, and sufficiency.

Temporality enforces the directional ordering constraint that a preceding interaction must occur before the downstream interaction:

Temp(τp→τq)=𝕀[p<q].\mathrm{Temp}(\tau_{p}\rightarrow\tau_{q})=\mathbb{I}[p<q]. (3)

Necessity tests whether the τq\tau_{q} would fail to occur if the upstream step τp\tau_{p} is removed, using a counterfactual-inspired judgment function JJ:

Nec⁡(τp→τq)=𝕀⁡(J⁡(do⁡(¬τp))=No).\mathrm{Nec}(\tau_{p}\rightarrow\tau_{q})=\mathbb{I}\!\left(J(\mathrm{do}(\neg\tau_{p}))=\text{No}\right). (4)

Sufficiency assesses whether enforcing τp\tau_{p} is sufficient to plausibly induce the outcome at τq\tau_{q}:

Suff⁡(τp→τq)=𝕀⁡(J⁡(do⁡(τp))=Yes).\mathrm{Suff}(\tau_{p}\rightarrow\tau_{q})=\mathbb{I}\!\left(J(\mathrm{do}(\tau_{p}))=\text{Yes}\right). (5)

An edge is added only when the temporal, necessity, and sufficiency conditions are all satisfied. This conservative rule filters out spurious correlations, retaining causally supported dependencies.

Causal-inspired Structured Reasoning.

Given the structured dependency criteria, GCA performs causal-inspired structured reasoning over the system trace 𝒯\mathcal{T}, focusing on the identified minor deviations τe∈𝒯e\tau^{e}\in\mathcal{T}^{e} to construct the causal-inspired dependency graph 𝒢\mathcal{G}. The graph is constructed by incorporating these criteria into instruction sets ℐc\mathcal{I}^{c}:

𝒢=CGC⁡(𝒯,𝒯e,ℛe,ℐc),\mathcal{G}=\mathrm{CGC}\big(\mathcal{T},\mathcal{T}^{e},\mathcal{R}^{e},\mathcal{I}^{c}\big), (6)

where CGC⁡(⋅)\mathrm{CGC}(\cdot) denotes the dependency graph construction procedure.

4.1.3 Hypothesis Generation

Based on the causal-inspired dependency graph 𝒢\mathcal{G} and the candidate minor deviation set 𝒯e\mathcal{T}^{e}, GCA performs a global analysis to propose an initial hypothesis of the decisive error of the system failure. The hypothesis is produced by the hypothesis generation process HG​(⋅)\text{HG}(\cdot) on a compact evidence bundle that includes: the ground truth outcome 𝒴^\widehat{\mathcal{Y}}, the entire system trace 𝒯\mathcal{T}, the candidate minor deviations 𝒯e\mathcal{T}^{e}, the corresponding textual reasons ℛe\mathcal{R}^{e} to 𝒯e\mathcal{T}^{e}, and the constructed graph 𝒢\mathcal{G}. Formally, the hypothesis generation is formulated as follows:

(τ⋆,r⋆)=HG​(𝒴^,𝒯,𝒯e,ℛe,𝒢),(\tau^{\star},r^{\star})=\text{HG}\Big(\widehat{\mathcal{Y}},\,\mathcal{T},\,\mathcal{T}^{e},\,\mathcal{R}^{e},\,\mathcal{G}\Big), (7)

where the output τ⋆\tau^{\star} is the hypothesis of the decisive error, and r⋆r^{\star} is the corresponding reason.

4.2 Local Counterfactual-inspired Enhancement Module

Practical system traces are often long and dense Zhang et al. (2025), which exacerbates performance degradation in failure attribution based on full traces. To address this limitation, we introduce the LCE module, which refines τ⋆\tau^{\star} by focusing on a compact local dependency neighborhood. LCE performs a bidirectional search over the causal-inspired dependency graph and uses counterfactual-inspired evaluation with LLM-mediated approximate interventions to assess how correcting candidate interactions may alter the final outcome. By iteratively evaluating local outcome changes, LCE identifies the interaction that most plausibly constitutes the decisive error.

4.2.1 Bidirectional Dependency Neighborhood Search

The bidirectional search aims to identify a minimal set of interactions that provide relevant evidence for explaining the system failure and to localize the decisive error by tracing dependency propagation across upstream and downstream interactions. Starting from the index i⋆i^{\star} of the global hypothesis τ⋆\tau^{\star}, LCE initializes the pivot as p(0)=i⋆p^{(0)}=i^{\star} and explores the causal-inspired dependency graph 𝒢\mathcal{G} in both forward and backward directions.

The forward search traces downstream dependencies to capture how the hypothesized error may propagate, while the backward search examines upstream dependencies to identify prior misreasoning or missing premises. In both directions, interactions are selected greedily based on counterfactual-inspired evaluation, which assesses how modifying a candidate interaction may affect the final outcome. The union of the forward and backward results forms a locally refined dependency subchain centered on τ⋆\tau^{\star}.

Counterfactual-inspired Evaluation

Central to LCE is a counterfactual-inspired evaluation function that assesses the potential causal importance of an interaction by measuring the change in outcome alignment when that interaction is hypothetically corrected instead.

We first construct a fully corrected trace by applying a correction function Correct​(⋅)\text{Correct}(\cdot) conditioned on the ground-truth outcome Y^\widehat{Y}, the original trace 𝒯\mathcal{T}, extracted erroneous interactions 𝒯e\mathcal{T}^{e}, their corresponding reasons ℛe\mathcal{R}^{e}, and the causal-inspired dependency graph 𝒢\mathcal{G}:

𝒯^=Correct​(Y^,𝒯,𝒯e,ℛe,𝒢),\widehat{\mathcal{T}}=\text{Correct}\left(\widehat{Y},\mathcal{T},\mathcal{T}^{e},\mathcal{R}^{e},\mathcal{G}\right), (8)

where 𝒯^={τ^1,…,τ^N}\widehat{\mathcal{T}}=\{\widehat{\tau}_{1},\dots,\widehat{\tau}_{N}\} and τ^i\widehat{\tau}_{i} denotes the corrected version of τi\tau_{i}.

Using 𝒯^\widehat{\mathcal{T}}, we construct a counterfactual-inspired trace for each interaction τi\tau_{i} by combining the corrected prefix up to step ii with the original suffix from step i+1i\!+\!1, which is denoted as 𝒯~(i)={τ^1,…,τ^i,τi+1,…,τN}\widetilde{\mathcal{T}}^{(i)}=\{\widehat{\tau}_{1},\dots,\widehat{\tau}_{i},\tau_{i+1},\dots,\tau_{N}\}. This represents a scenario in which only the first ii interactions have been corrected. The counterfactual-inspired trace is then evaluated via a simulation process:

Y(i)=Simulation​(𝒯~(i)).Y^{(i)}=\text{Simulation}\bigl(\widetilde{\mathcal{T}}^{(i)}\bigr). (9)

We measure the semantic alignment between the outcome Y(i)Y^{(i)} and the ground truth Y^\widehat{Y} using a pretrained encoder Enc⁡(⋅)\mathrm{Enc}(\cdot) and cosine similarity:

𝒮⁡(τi)=Cos⁡(Enc⁡(Y(i)),Enc⁡(Y^)).\mathcal{S}(\tau_{i})=\mathrm{Cos}\bigl(\mathrm{Enc}(Y^{(i)}),\mathrm{Enc}(\widehat{Y})\bigr). (10)

The marginal attributional impact of correcting τi\tau_{i} is defined as

Δ​𝒮​(τi)=𝒮⁡(τi)−𝒮⁡(τi−1),\Delta\mathcal{S}(\tau_{i})=\mathcal{S}(\tau_{i})-\mathcal{S}(\tau_{i-1}), (11)

which measures the incremental improvement in outcome alignment obtained by correcting the ii-th interaction after all preceding interactions have been corrected.

Bidirectional Greedy Search

Using the counterfactual-inspired scores, LCE performs a greedy bidirectional search over the causal-inspired dependency graph 𝒢\mathcal{G}.

In the forward search, given the pivot p(l−1)p^{(l-1)} at iteration ll, we collect all direct successors denoted as 𝒯eff(l−1)={τj∣(p(l−1),j)∈ℰ}\mathcal{T}^{(l-1)}_{\text{eff}}=\{\tau_{j}\mid(p^{(l-1)},j)\in\mathcal{E}\}, where ℰ\mathcal{E} denotes the edge set of 𝒢\mathcal{G}. The successor with the largest positive marginal contribution is selected:

τj(l)=arg⁡maxτj∈𝒯eff(l−1)​Δ​𝒮​(τj).\tau_{j^{(l)}}=\arg\max_{\tau_{j}\in\mathcal{T}^{(l-1)}_{\text{eff}}}\Delta\mathcal{S}(\tau_{j}). (12)

If Δ​𝒮​(τj(l))>0\Delta\mathcal{S}(\tau_{j^{(l)}})>0, the forward candidate set, denoted as 𝒞fwd\mathcal{C}_{\text{fwd}}, is expanded and the pivot updated as j(l)j^{(l)}. Otherwise, the forward search terminates.

Symmetrically, the backward search traces incoming edges to identify upstream causes denoted as 𝒯cau(l−1)={τk∣(k,p(l−1))∈ℰ}\mathcal{T}^{(l-1)}_{\text{cau}}=\{\tau_{k}\mid(k,p^{(l-1)})\in\mathcal{E}\}. The predecessor with the strongest positive effect is selected:

τk(l)=arg⁡maxτk∈𝒯cau(l−1)​Δ​𝒮​(τk).\tau_{k^{(l)}}=\arg\max_{\tau_{k}\in\mathcal{T}^{(l-1)}_{\text{cau}}}\Delta\mathcal{S}(\tau_{k}). (13)

If Δ​𝒮​(τk(l))>0\Delta\mathcal{S}(\tau_{k^{(l)}})>0, the backward candidate set, denoted as 𝒞bwd\mathcal{C}_{\text{bwd}}, is expanded and the pivot updated as k(l)k^{(l)}. This process continues until no further positive contributions are found.

These two searches yield a local dependency neighborhood that captures both antecedent causes and propagated effects surrounding τ⋆\tau^{\star}.

4.2.2 Decisive Error Decision

The forward and backward candidate sets are merged with the global hypothesis to form the final local candidate set 𝒯c={τi∣i∈{i⋆}∪𝒞fwd∪𝒞bwd}\mathcal{T}^{c}=\{\tau_{i}\mid i\in\{i^{\star}\}\cup\mathcal{C}_{\text{fwd}}\cup\mathcal{C}_{\text{bwd}}\}. Each element in 𝒯c\mathcal{T}^{c} corresponds to an interaction whose counterfactual-inspired correction yields a positive improvement in outcome alignment.

Unlike GCA, which analyzes the full trace, LCE operates only on the interaction subset 𝒯c\mathcal{T}^{c}, avoiding context degradation. This evaluation is formalized by a decisive-error decision function:

τ∗=DED​(Y^,𝒯c,𝒢),\tau^{*}=\text{DED}\bigl(\widehat{Y},\mathcal{T}^{c},\mathcal{G}\bigr), (14)

where τ∗\tau^{*} denotes the final identified decisive error.

5 Experimental Design

5.1 Dataset

We evaluate DCFA on the Who&When benchmark Zhang et al. (2025), which is currently the only widely adopted benchmark specifically designed for failure attribution in LLM-based multi-agent systems and serves as the common evaluation testbed for existing baselines West et al. (2025); Banerjee et al. (2025). The dataset contains 184 multi-agent system traces spanning both Algorithm-generated and Hand-crafted settings. Specifically, 126 traces are generated using CaptainAgent-based systems Wu et al. (2023), and 58 traces are manually curated from systems such as Magnetic-One Fourney et al. (2024). The tasks cover web-navigation and general reasoning scenarios derived from AssistantBench Yoran et al. (2024) and GAIA Mialon et al. (2023). Each trace is annotated with fine-grained failure labels, including the decisive error and a natural-language explanation. Details are provided in Appendix A.

5.2 Baselines

We compare DCFA against five representative baselines that reflect different LLM-based failure attribution paradigms: All-at-Once, Step-by-Step, and Binary-Search from the Who&When benchmark Zhang et al. (2025), as well as A2P West et al. (2025) and ECHO Banerjee et al. (2025). All methods rely on LLMs for attribution and are evaluated under identical inference settings. We test both open-source and commercial models, including Qwen3-Coder-30B, Qwen3-235B, DeepSeek-R1-32B, DeepSeek-R1-671B, GPT-5, and Gemini-2.5-pro. Details are deferred to Appendix B.

Model Algorithm-generated Hand-crafted
All-at-Once Step-by-Step Binary-Search A2P ECHO DCFA All-at-Once Step-by-Step Binary-Search A2P ECHO DCFA
Qwen3-Coder-30B 11.90 20.63 13.49 12.70 23.02 37.30 1.72 10.34 6.67 8.62 13.79 15.52
DeepSeek-R1-32B 22.22 29.37 30.16 19.84 30.95 42.06 5.17 15.52 6.90 6.90 18.97 20.69
Qwen3-235B 30.95 25.40 21.43 29.37 36.50 43.65 3.45 15.52 12.07 3.45 17.24 22.41
DeepSeek-R1-671B 16.67 26.98 32.54 15.87 40.48 53.17 3.45 12.07 5.17 6.90 20.69 22.41
GPT-5 15.87 27.78 31.75 15.87 31.75 47.62 3.45 13.79 10.34 13.79 15.52 24.14
Gemini-2.5-pro 26.98 34.13 25.40 32.54 30.95 46.03 5.17 15.52 6.90 10.34 13.79 18.97
Average 20.77 27.72 25.13 21.03 32.28 44.79 3.74 13.46 8.33 8.33 16.67 20.69
Table 1: Step-level Accuracy Comparison on Algorithm-generated and Hand-crafted datasets.

5.3 Implementation Details

In DCFA, the GCA stage is performed by commercial LLMs using the same configurations as the baselines. The LCE stage is executed with a locally deployed Qwen-Coder-30B model. The judgment function JJ used for dependency condition evaluation in GCA is implemented via structured prompting, where the LLM is queried to perform binary judgments. All LLMs are run with the temperature set to 0 to ensure deterministic inference. All baselines follow identical preprocessing pipelines and prompt configurations as specified in their original papers and released codebases to ensure fair comparison. As ECHO does not provide an official implementation, we reimplement it based on the descriptions in its paper.

We evaluate attribution quality using Step-level Accuracy, which directly measures precise failure localization and avoids the inflation effects of agent-level metrics. Additional implementation and evaluation details are provided in Appendix C.

5.4 Results and Analysis

5.4.1 Overall Performance

Table 1 reports step-level accuracy of DCFA and five baselines on the Algorithm-generated and Hand-crafted datasets. Compared with the strongest baseline on each dataset, DCFA achieves average improvements of 12.51% and 4.02%, respectively. On the Algorithm-generated dataset, DeepSeek-R1 achieves the highest accuracy, likely benefiting from reasoning-oriented training that aligns with DCFA’s structured causal-inspired dependency graph construction Guo et al. (2025), while GPT remains competitive on longer traces due to its strong long-context reasoning Leon (2025). In contrast, most baselines rely primarily on prompt-induced implicit reasoning, making their attributions less sensitive to the latent dependency relations in multi-agent interactions. For instance, ECHO focuses on observable errors such as tool failures or formatting issues. While these explicit patterns can improve detection accuracy, the resulting deviations are often minor rather than decisive. By tracing such surface symptoms back to their underlying causes, DCFA enables more precise and causally grounded failure attribution. Examples are presented in the case study (Sec. 5.4.4), with additional analysis of GCA-induced causal-inspired dependency graphs in Appendix F.

Model Algorithm-generated Hand-crafted
w/o LCE DCFA w/o LCE DCFA
Qwen3-Coder-30B 33.33 37.30 12.07 15.52
DeepSeek-R1-32B 40.48 42.06 15.52 20.69
Qwen3-235B 42.06 43.65 20.69 22.41
DeepSeek-R1-671B 53.17 53.17 18.97 22.41
GPT-5 44.44 47.62 22.41 24.14
Gemini-2.5-pro 43.65 46.03 13.79 18.97
Table 2: Ablation study of the LCE module.

5.4.2 Ablation on LCE

Since LCE operates on the causal-inspired dependency graph and the hypothesis of the decisive error produced by GCA, we compare GCA-only with DCFA. Table 2 reports an ablation study of DCFA with and without LCE. Adding LCE consistently improves step-level accuracy on both datasets, with 3.45% average improvement on the Hand-crafted dataset and 1.93% average improvement on the Algorithm-generated dataset.

Notably, GCA alone exhibits a noticeable performance drop on the Hand-crafted dataset, where traces are substantially longer (average length =51.6=51.6) than those in the Algorithm-generated dataset (average length =8.7=8.7). In contrast, LCE remains effective even when global attribution degrades under long-context conditions. This robustness arises because LCE identifies the local dependency chains with the greatest corrective impact through bidirectional search and counterfactual-inspired evaluation, capturing both upstream causes and downstream effects of the error localized by GCA. By focusing the LLM on higher-level and more specific causal-inspired structures, LCE mitigates context degradation and improves attribution precision.

5.4.3 Performance on Varying Trace Lengths

We evaluate DCFA on the Hand-crafted dataset using DeepSeek-R1-671B and 32B, comparing it with all baselines across five context-length levels defined by the Who&When benchmark Zhang et al. (2025).

As shown in Fig. 3(a) and Fig. 3(b), DCFA outperforms all baselines across nearly all levels, with the largest gains observed at Level 1 and maintained through Level 5. One exception occurs at Level 2 with DeepSeek-R1-32B, where DCFA slightly underperforms ECHO. This may be due to the reduced reliability of weaker LLMs in following the structured dependency reasoning required by DCFA, which can introduce additional variance in LLM-mediated evaluation. Nevertheless, DCFA remains robust overall, achieving strong performance across varying trace lengths.

(a)
(b)
Figure 3: Step-level accuracy of DCFA and baselines across trace lengths using DeepSeek-R1 series LLMs.

5.4.4 Case Study on Causal-inspired Dependency Graph Search

We analyze a system trace from a hand-crafted multi-agent system (ID: 49.json), where three agents (Orchestrator, WebSurfer, and Assistant) collaboratively solve an Unlambda debugging query (see details in Appendix H). The execution trace 𝒯={τ1,…,τ15}\mathcal{T}=\{\tau_{1},\dots,\tau_{15}\} results in an incorrect prediction of “k”, whereas the correct missing character is the backtick “`”.

Inference Phase of GCA.

GCA first identifies multiple erroneous or misleading interactions that are potentially involved in the failure. Specifically, it detects a minor deviation at τ8\tau_{8}, where WebSurfer returns incomplete operator definitions, and further identifies τ9\tau_{9}, where the Orchestrator propagates the partial information without further validation. These errors form a dependency chain and jointly contribute to the subsequent failure, rather than constituting independent failure sources. Based on the resulting causal-inspired dependency graph, GCA localizes the relevant error region for further refinement, as shown in Figure 4.

Refinement Phase via LCE.

LCE then conducts localized, bidirectional counterfactual-inspired reasoning over the error region identified by GCA. By examining the corrective effects of candidate interactions along the dependency graph, LCE determines that τ12\tau_{12} has the greatest impact on mitigating the system failure. At τ12\tau_{12}, the Assistant fabricates a solution based on incomplete semantics, which further propagates the accumulated errors toward the incorrect final outcome. Although correcting τ8\tau_{8} or τ9\tau_{9} can mitigate the error propagation, neither correction is sufficient to fully resolve the failure. In contrast, correcting τ12\tau_{12} provides the largest reduction in the failure, thereby identifying it as the decisive error. The refinement process is illustrated in Figure 4.

Figure 4: LCE refinement on the causal-inspired dependency graph of trace 49.json. Orange dashed arrows indicate search paths. Forward and backward traversal evaluates the corrective effects of candidate interactions, isolating τ12\tau_{12} as the decisive error. Yellow: minor deviation (τ8\tau_{8}); Orange: GCA-identified error (τ9\tau_{9}); Red: decisive error (τ12\tau_{12}).
Decisive Error and DCFA Attribution.

This case exemplifies the multi-error setting in our problem formulation, where multiple errors jointly contribute to the system failure while no single correction among the earlier errors fully recovers the correct outcome. Specifically, τ8\tau_{8} and τ9\tau_{9} contribute to the error propagation and can be corrected to mitigate the failure, whereas τ12\tau_{12} has the greatest corrective effect and is therefore identified as the decisive error. By combining global causal-inspired dependency search with local counterfactual-inspired refinement, DCFA successfully distinguishes contributing errors from the decisive error and accurately attributes the system failure to τ12\tau_{12}. Additional details are provided in Appendix H.

5.4.5 Computational Cost

We analyze DCFA’s computational cost by decomposing it into GCA and LCE. GCA relies on commercial LLM API calls, while LCE is executed locally and adds only inference-time overhead. Following Zhang et al. (2025), output-token costs are ignored unless stated otherwise.

GCA performs two full-context LLM calls: deviation detection with dependency graph construction, and decisive-error refinement. Since outputs are small relative to the input context, the total cost can be approximated as

CostGCA≈2​C+2​n¯​L,\text{Cost}_{\text{GCA}}\approx 2C+2\overline{n}L, (15)

where CC is the prompt overhead, n¯\overline{n} the average trace length, and LL the token length per step. Notably, the strongest baseline ECHO requires the same number of full-context calls, resulting in identical computational complexity, while GCA alone already yields substantial accuracy improvements.

The LCE module runs locally and introduces additional inference overhead. Its complexity can be approximated as

CostLCE≈m​C+m​n¯​L,\text{Cost}_{\text{LCE}}\approx mC+m\overline{n}L, (16)

where m≈4.5m\approx 4.5 is the average number of bidirectional neighborhood expansions. Although LCE increases runtime, it further improves attribution accuracy, especially on longer trajectories.

Table 3 summarizes the approximate runtime, token consumption, and attribution accuracy of all compared methods on the Who&When benchmark. DCFA without LCE incurs a cost comparable to ECHO while achieving substantially higher accuracy (30.0530.05 vs. 24.4824.48). Adding LCE increases the runtime from 32s to 300s and token consumption from 13k to 32k, but further improves accuracy to 32.7432.74. Detailed runtime statistics, token usage, and comprehensive comparisons with baseline methods are provided in Appendix E.

Method Runtime (s) Tokens (k) Accuracy
All-at-Once 15 6 12.26
Search-by-Step 61 9 20.59
Binary-Search 25 12 16.73
A2P 15 6 14.68
ECHO 31 12 24.48
DCFA w/o LCE 32 13 30.05
DCFA 300 32 32.74
Table 3: Approximate runtime, token usage, and attribution accuracy on the Who&When benchmark.

6 Conclusion

In this work, we propose DCFA, a training-free framework for failure attribution in LLM-based multi-agent systems. Combining global causal-inspired dependency graph analysis with local counterfactual-inspired reasoning, DCFA identifies decisive errors and performs robustly across traces of varying lengths. Experiments show it improves step-level attribution accuracy by up to 8.27% over state-of-the-art baselines and remains effective on challenging long-context traces.

Limitations

Dependence on LLM Capability.

DCFA relies on the reasoning inference capabilities of the underlying large language models. When system traces become very long, LLMs may struggle to maintain coherent dependency representations, which can reduce the accuracy of global causal-inspired dependency graph construction. Similarly, when using smaller or less capable models for local counterfactual-inspired refinement, the precision of decisive error identification may be constrained.

Scalability to Long or Complex Traces.

Although the global-local combination in DCFA improves robustness to moderately long execution traces, extremely long or highly branching multi-agent interactions can still pose challenges. In such cases, both global causal-inspired attribution and local counterfactual-inspired reasoning may become less reliable, limiting applicability to systems with deeply nested or prolonged interaction patterns.

Attribution of Intrinsic LLM Reasoning Errors.

Compared with the decisive errors that reflect failures within the MAS workflow. Errors originating from inherent LLM hallucinations may be difficult to attribute accurately.

Ethical Considerations

This research is intended as a diagnostic tool to support failure analysis in LLM-based multi-agent systems, rather than an automated mechanism for judgment or accountability. Its attribution results depend on the reasoning behavior of underlying language models and may be imperfect, especially when failures stem from intrinsic model errors or ambiguous interactions. Over-reliance on such automated explanations could lead to misinterpretation of responsibility if used without human oversight. Therefore, we highlight that the proposed DCFA should be applied with caution in high-stakes settings and used to assist, not replace, human analysis of system failures.

Acknowledgment

This work was supported in part by the National Natural Science Foundation of China under Grants 62572346 and 62322208.

References

  • Ban et al. (2025) T. Ban, L. Chen, D. Lyu, X. Wang, Q. Zhu, Q. Tu, and H. Chen Integrating large language model for improved causal discovery. IEEE Transactions on Artificial Intelligence. Cited by: §J.3, §2.3.
  • Banerjee et al. (2025) A. Banerjee, A. Nair, and T. Borogovac Where did it all go wrong? a hierarchical look into multi-agent error attribution. In NeurIPS 2025 Workshop on Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling, Cited by: §J.2, §J.2, Appendix B, §1, §2.2, §5.1, §5.2.
  • Barta et al. (2025) Z. Barta, B. Nagy, and L. Gulyás Measuring the robustness of multi-agent reinforcement learning systems under partial agent failure. In Proceedings of the Intelligent Robotics FAIR 2025, pp. 58–63. Cited by: §4.1.1.
  • Becattini et al. (2025) M. Becattini, R. Verdecchia, and E. Vicario SALLMA: a software architecture for llm-based multi-agent systems. In 2025 IEEE/ACM International Workshop New Trends in Software Architecture (SATrends), pp. 5–8. Cited by: §1.
  • Cemri et al. (2026) M. Cemri, M. Z. Pan, S. Yang, L. A. Agrawal, B. Chopra, R. Tiwari, K. Keutzer, A. Parameswaran, D. Klein, K. Ramchandran, et al. Why do multi-agent llm systems fail?. Advances in Neural Information Processing Systems 38. Cited by: §J.1, §J.2, §J.2, §1, §1, §1, §2.1, §2.2, §2.2, §4.1.1.
  • Cheng et al. (2025) Q. Cheng, Z. Zeng, X. Hu, Y. Si, and Z. Liu A survey of event causality identification: taxonomy, challenges, assessment, and prospects. ACM Computing Surveys 58 (3), pp. 1–37. Cited by: §J.3, §2.3.
  • Cohrs et al. (2025) K. Cohrs, E. Diaz, V. Sitokonstantinou, G. Varando, and G. Camps-Valls Large language models for causal hypothesis generation in science. Machine Learning: Science and Technology 6 (1), pp. 013001. Cited by: §J.3.
  • Darvariu et al. (2024) V. Darvariu, S. Hailes, and M. Musolesi Large language models are effective priors for causal graph discovery. arXiv preprint arXiv:2405.13551. Cited by: §J.3, §2.3.
  • Deshpande et al. (2025) D. Deshpande, V. Gangal, H. Mehta, J. Krishnan, A. Kannappan, and R. Qian Trail: trace reasoning and agentic issue localization. arXiv preprint arXiv:2505.08638. Cited by: Appendix A, §J.2, §2.2.
  • Fourney et al. (2024) A. Fourney, G. Bansal, H. Mozannar, C. Tan, E. Salinas, F. Niedtner, G. Proebsting, G. Bassman, J. Gerrits, J. Alber, et al. Magentic-one: a generalist multi-agent system for solving complex tasks. arXiv preprint arXiv:2411.04468. Cited by: Appendix A, §J.1, §J.2, §2.1, §2.2, §5.1.
  • Goldberg (2019) L. R. Goldberg The book of why: the new science of cause and effect: by judea pearl and dana mackenzie, basic books (2018). isbn: 978-0465097609.. Taylor & Francis. Cited by: §4.1.2.
  • Guo et al. (2025) D. Guo, D. Yang, H. Zhang, J. Song, P. Wang, Q. Zhu, R. Xu, R. Zhang, S. Ma, X. Bi, et al. Deepseek-r1: incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948. Cited by: §5.4.1.
  • Hammond et al. (2025) L. Hammond, A. Chan, J. Clifton, J. Hoelscher-Obermaier, A. Khan, E. McLean, C. Smith, W. Barfuss, J. Foerster, T. Gavenčiak, et al. Multi-agent risks from advanced ai. arXiv preprint arXiv:2502.14143. Cited by: §1, §4.1.1.
  • He et al. (2025) X. He, D. Wu, Y. Zhai, and K. Sun SentinelAgent: graph-based anomaly detection in multi-agent systems. arXiv preprint arXiv:2505.24201. Cited by: §1.
  • Huang et al. (2024) Y. Huang, F. Cheng, F. Zhou, J. Li, J. Gong, H. Yang, Z. Fan, C. Jiang, S. Xue, and F. Chen Romas: a role-based multi-agent system for database monitoring and planning. arXiv preprint arXiv:2412.13520. Cited by: §1.
  • Imai and Nakamura (2026) K. Imai and K. Nakamura Causal inference with generative artificial intelligence: application to texts as treatments. Journal of the American Statistical Association (just-accepted), pp. 1–27. Cited by: §J.3.
  • [17] T. Jiralerspong, X. Chen, Y. More, V. Shah, and Y. Bengio Efficient causal graph discovery using large language models. In ICLR 2024 Workshop: How Far Are We From AGI, Cited by: §J.3, §2.3.
  • Leon (2025) M. Leon GPT-5 and open-weight large language models: advances in reasoning, transparency, and control. Information Systems, pp. 102620. Cited by: §5.4.1.
  • Li et al. (2024) X. Li, S. Wang, S. Zeng, Y. Wu, and Y. Yang A survey on llm-based multi-agent systems: workflow, infrastructure, and challenges. Vicinagearth 1 (1), pp. 9. Cited by: §1.
  • Li et al. (2022) Y. Li, L. Meng, L. Chen, L. Yu, D. Wu, Y. Zhou, and B. Xu Training data debugging for the fairness of machine learning software. In Proceedings of the 44th International Conference on Software Engineering, pp. 2215–2227. Cited by: Appendix I.
  • Liu et al. (2024a) C. Liu, W. Xiang, and B. Wang Identifying while learning for document event causality identification. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 3815–3827. Cited by: §J.3.
  • Liu et al. (2024b) X. Liu, H. Yu, H. Zhang, Y. Xu, X. Lei, H. Lai, Y. Gu, H. Ding, K. Men, K. Yang, et al. Agentbench: evaluating llms as agents. In International Conference on Learning Representations, Vol. 2024, pp. 52989–53046. Cited by: §J.1, §2.1.
  • Liu et al. (2025) X. Liu, P. Xu, J. Wu, J. Yuan, Y. Yang, Y. Zhou, F. Liu, T. Guan, H. Wang, T. Yu, et al. Large language models and causal inference in collaboration: a comprehensive survey. Findings of the Association for Computational Linguistics: NAACL 2025, pp. 7668–7684. Cited by: §J.3, §J.3.
  • Luo et al. (2024) K. Luo, T. Zhou, Y. Chen, J. Zhao, and K. Liu Open event causality extraction by the assistance of llm in task annotation, dataset, and method. In Proceedings of the Workshop: Bridging Neurons and Symbols for Natural Language Processing and Knowledge Graphs Reasoning (NeusymBridge)@ LREC-COLING-2024, pp. 33–44. Cited by: §J.3, §2.3.
  • Majdoub et al. (2025) Y. Majdoub, E. B. Charrada, and H. Touati Towards adaptive software agents for debugging. In Proceedings of the 33rd ACM International Conference on the Foundations of Software Engineering, pp. 636–640. Cited by: §1.
  • Mialon et al. (2023) G. Mialon, C. Fourrier, T. Wolf, Y. LeCun, and T. Scialom Gaia: a benchmark for general ai assistants. In The Twelfth International Conference on Learning Representations, Cited by: Appendix A, Appendix A, §5.1.
  • Pan et al. (2024) Y. Pan, J. Sun, H. Yu, J. Luck, G. Bai, N. Chamara, Y. Ge, and T. Awada Building multi-agent copilot towards autonomous agricultural data management and analysis. In 2024 IEEE International Conference on Big Data (BigData), pp. 4384–4393. Cited by: §1.
  • Pei et al. (2025) C. Pei, Z. Wang, F. Liu, Z. Li, Y. Liu, X. He, R. Kang, T. Zhang, J. Chen, J. Li, et al. Flow-of-action: sop enhanced llm-based multi-agent system for root cause analysis. In Companion Proceedings of the ACM on Web Conference 2025, pp. 422–431. Cited by: §J.1.
  • Ronanki (2025) K. Ronanki Facilitating trustworthy human-agent collaboration in llm-based multi-agent system oriented software engineering. In Proceedings of the 33rd ACM International Conference on the Foundations of Software Engineering, pp. 1333–1337. Cited by: §1.
  • Shen et al. (2026) X. Shen, Q. Zhang, S. Wang, Z. Tan, X. Zhao, L. Yao, V. Tadiparthi, H. N. Mahjoub, E. M. Pari, K. Lee, et al. Metacognitive self-correction for multi-agent system via prototype-guided next-execution reconstruction. In Findings of the Association for Computational Linguistics: ACL 2026, pp. 23320–23337. Cited by: §1.
  • Shrotriya et al. (2025) S. Shrotriya, N. Banu, A. Kulkarni, and S. Aiyangar Navigating the path to ethical & responsible ai integration in health & life sciences with human and machine collaboration. Journal of the Society for Clinical Data Management 5 (1), pp. 1–12. Cited by: Appendix I.
  • Su et al. (2025) Y. Su, H. Zhang, G. Zhang, Y. Wang, Y. Fan, R. Li, and Y. Wang Enhancing event causality identification with llm knowledge and concept-level event relations. In Proceedings of the 31st International Conference on Computational Linguistics, pp. 7403–7414. Cited by: §J.3.
  • Sun et al. (2025) C. Sun, S. Huang, and D. Pompili LLM-based multi-agent decision-making: challenges and future directions. IEEE Robotics and Automation Letters. Cited by: §J.1, §2.1.
  • Tong et al. (2024) S. Tong, K. Mao, Z. Huang, Y. Zhao, and K. Peng Automating psychological hypothesis generation with ai: when large language models meet causal graph. Humanities and Social Sciences Communications 11 (1), pp. 896. Cited by: §J.3, §2.3.
  • Tutek et al. (2025) M. Tutek, F. H. Chaleshtori, A. Marasović, and Y. Belinkov Measuring chain of thought faithfulness by unlearning reasoning steps. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pp. 9946–9971. Cited by: §J.3, §2.3.
  • Verma et al. (2025) V. Verma, S. Acharya, S. Simko, D. Bhardwaj, A. Haghighat, D. Janzing, M. Sachan, Z. Jin, and Y. Yang Causal ai scientist: facilitating causal data science with large language models. In NeurIPS 2025 AI for Science Workshop, Cited by: §J.3.
  • Wang et al. (2024) H. Wang, F. Liu, J. Zhang, D. Roth, and K. Richardson Event causality identification with synthetic control. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp. 1725–1737. Cited by: §J.3.
  • West et al. (2025) A. West, Y. Weng, M. Zhu, Z. Lin, Z. Ning, and Y. Zhang Abduct, act, predict: scaffolding causal inference for automated failure attribution in multi-agent systems. arXiv preprint arXiv:2509.10401. Cited by: §J.2, §J.2, Appendix B, §2.2, §5.1, §5.2.
  • Wu et al. (2023) Q. Wu, G. Bansal, J. Zhang, Y. Wu, B. Li, E. Zhu, L. Jiang, X. Zhang, S. Zhang, J. Liu, et al. Autogen: enabling next-gen llm applications via multi-agent conversation. arXiv preprint arXiv:2308.08155. Cited by: Appendix A, §J.1, §J.2, §2.1, §5.1.
  • Xia et al. (2025) C. S. Xia, Y. Deng, S. Dunn, and L. Zhang Demystifying llm-based software engineering agents. Proceedings of the ACM on Software Engineering 2 (FSE), pp. 801–824. Cited by: §J.1.
  • Yan et al. (2025) B. Yan, Z. Zhou, L. Zhang, L. Zhang, Z. Zhou, D. Miao, Z. Li, C. Li, and X. Zhang Beyond self-talk: a communication-centric survey of llm-based multi-agent systems. arXiv preprint arXiv:2502.14321. Cited by: §J.1.
  • Yang et al. (2026) Y. Yang, H. Chai, S. Shao, Y. Song, S. Qi, R. Rui, and W. Zhang Agentnet: decentralized evolutionary coordination for llm-based multi-agent systems. Advances in Neural Information Processing Systems 38, pp. 107309–107336. Cited by: §J.1, §2.1.
  • Yoran et al. (2024) O. Yoran, S. J. Amouyal, C. Malaviya, B. Bogin, O. Press, and J. Berant Assistantbench: can web agents solve realistic and time-consuming tasks?. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp. 8938–8968. Cited by: Appendix A, Appendix A, §5.1.
  • Yu et al. (2025) L. Yu, D. Chen, S. Xiong, Q. Wu, D. Li, Z. Chen, X. Liu, and L. Pan Causaleval: towards better causal reasoning in language models. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pp. 12512–12540. Cited by: §J.3, §2.3.
  • Zeng et al. (2026) Z. Zeng, Q. Cheng, X. Hu, W. Li, W. Ding, and Z. Liu Zero-shot event causality identification via multisource evidence fuzzy aggregation with large language models. IEEE Transactions on Fuzzy Systems. Cited by: §4.1.2.
  • Zhang et al. (2026) G. Zhang, J. Wang, J. Chen, W. Zhou, K. Wang, and S. Yan Agentracer: who is inducing failure in the llm agentic systems?. In International Conference on Learning Representations, Vol. 2026, pp. 11377–11399. Cited by: §J.2, §1, §2.2.
  • Zhang et al. (2025) S. Zhang, M. Yin, J. Zhang, J. Liu, Z. Han, J. Zhang, B. Li, C. Wang, H. Wang, Y. Chen, and Q. Wu Which agent causes task failures and when? On automated failure attribution of LLM multi-agent systems. In Proceedings of the 42nd International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 267, pp. 76583–76599. Cited by: Appendix A, §J.1, §J.2, Appendix B, Appendix I, §1, §1, §2.1, §2.2, §4.2, §5.1, §5.2, §5.4.3, §5.4.5.
  • Zhou et al. (2025) J. Zhou, J. Chen, Q. Lu, D. Zhao, and L. Zhu SHIELDA: structured handling of exceptions in llm-driven agentic workflows. arXiv preprint arXiv:2508.07935. Cited by: §4.1.1.
  • Zhu et al. (2025) K. Zhu, Z. Liu, B. Li, M. Tian, Y. Yang, J. Zhang, P. Han, Q. Xie, F. Cui, W. Zhang, et al. Where llm agents fail and how they can learn from failures. arXiv preprint arXiv:2509.25370. Cited by: Appendix A, §J.2, §2.2.
  • Zhu et al. (2024) Y. Zhu, Y. He, J. Ma, M. Hu, S. Li, and J. Li Causal inference with latent variables: recent advances and future prospectives. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pp. 6677–6687. Cited by: §J.3.

Appendix A Dataset

We utilize the datasets from the Who&When benchmark Zhang et al. (2025), which comprise both Algorithm-generated and Hand-crafted multi-agent system traces, totaling 184 executions. Specifically, 126 traces are generated algorithmically using multi-agent systems built on CaptainAgent Wu et al. (2023), while 58 traces are manually curated from Hand-crafted systems such as Magnetic-One Fourney et al. (2024). The dataset encompasses a wide range of realistic multi-agent scenarios based on queries from GAIA Mialon et al. (2023) and AssistantBench Yoran et al. (2024).

The dataset covers a broad set of multi-agent tasks, including web-navigation-style challenges derived from AssistantBench Yoran et al. (2024) and general-assistant reasoning tasks inspired by GAIA Mialon et al. (2023). Each trace is annotated with fine-grained failure information, including the responsible agent, the true decisive error, and a natural language explanation of the failure.

The two categories exhibit distinct characteristics in terms of agent participation and trace length. Algorithm-generated traces involve less than 4 agents, with trace lengths ranging from 5 to 10 steps, a mean length of 8.7 steps, and a median length of 10 steps. In contrast, Hand-crafted traces involve between 1 and 5 agents, with trace lengths spanning 5 to 130 steps, a mean length of 51.6 steps, and a median length of 32 steps. The distributions of both datasets, illustrated in Fig. 5, highlight the greater variability present in the Hand-crafted traces.

Figure 5: Violin plot showing the distribution of trace lengths for the Algorithm-generated and Hand-crafted datasets. The width of each violin represents the kernel density estimation of the data distribution, illustrating the frequency of different trace lengths. The central red line indicates the median, while the black triangles mark the minimum and maximum values. Notably, wider regions denote higher concentrations of trace lengths within that range.

Moreover, some benchmarks, such as AgentErrorBench with AgentDebug Zhu et al. (2025) and TRAIL Deshpande et al. (2025), may appear superficially similar to our task setting. However they differ fundamentally from our problem formulation and evaluation objective, making direct comparison inappropriate. AgentErrorBench focuses on human-agent interaction and single-agent trajectories, where the objective is to detect the earliest mistake that initiates cascading reasoning distortions leading to failure. TRAIL is designed for identifying and analyzing all errors in trajectories. In contrast, this work focuses on failure attribution in LLM-based multi-agent systems, where the decisive error is defined as the earliest mistake whose correction could reverse the overall system failure. For example, an upstream deviation (an agent returning incomplete search results) eventually leads to a downstream decisive error (another agent fabricating missing information). AgentErrorBench would typically detect the earlier deviation, which could potentially be mitigated through subsequent agent interactions such as verification or additional tool usage, whereas this work aims to identify the interaction whose counterfactual-inspired correction would directly rescue the final outcome.

Appendix B Baselines

To evaluate the effectiveness of DCFA, we compare it against five representative baseline strategies that differ in how the LLM interacts with the failure traces: All-at-Once, Step-by-Step, and Binary-Search from the Who&When benchmark Zhang et al. (2025), as well as A2P West et al. (2025) and ECHO Banerjee et al. (2025). Each baseline reflects a distinct reasoning paradigm in localizing the failure-responsible step and agent within multi-agent system traces.

  • •

    All-at-Once. The algorithm leverages the LLM to interpret the entire system trace in a single pass and directly infer the decisive error of the system failure.

  • •

    Step-by-Step. The algorithm partitions the complete system trace into individual interactions. At each step, the model determines whether an error has occurred in the current interaction. If an error is detected, the process terminates immediately and the LLM outputs the corresponding step index. Otherwise, the reasoning continues sequentially until the final interaction is reached.

  • •

    Binary-Search. The algorithm begins with the query and the full failure log. It first determines whether the fault occurred in the upper or lower half of the trace. The identified half is then provided for further inspection. This iterative procedure continues until a single interaction is isolated as the decisive error.

  • •

    A2P. The algorithm employs a causal-inspired scaffold that sequentially guides LLMs through abductive hypothesis generation, explicit action-level intervention, and short-horizon counterfactual-inspired simulation to assess whether correcting a given step can reverse a system failure.

  • •

    ECHO. The algorithm introduces a hierarchical error attribution approach for multi-agent systems, explicitly targeting both minor execution deviations and overt errors within long interaction traces. Through multi-level contextual abstraction and consensus-based analysis, ECHO can detect early, subtle deviations that precede downstream failures, enabling fine-grained and reliable error localization.

All the baselines are based on the LLMs to achieve the failure attribution. To comprehensively assess each algorithm, we test them with various LLMs, including both open-source models of different sizes and commercial models. The models tested include the Qwen3 series (Qwen3-Coder-30B and Qwen3-235B), the DeepSeek series (DeepSeek-R1-32B and DeepSeek-R1-671B), GPT-5, and Gemini-2.5-pro. All baselines are evaluated under identical inference conditions to ensure a fair and consistent comparison.

Appendix C Implementation Details

Our approach operates in a training-free manner. In the GCA stage, minor deviation identification, causal-inspired dependency graph construction, and hypothesis generation are performed by the selected commercial LLM, following the same configuration as the baseline setup. In the LCE stage, the bidirectional dependency neighborhood search and decisive error decision are executed using the locally deployed Qwen-Coder-30B model. The encoder used to transform the outcomes generated by the simulation process into a continuous embedding space is implemented using the text-embedding-3-large API provided by OpenAI.

The judgment function JJ is implemented as a prompt-based evaluator operating over multi-agent system traces. Rather than re-executing the entire system, JJ receives a system trace containing agent identities and message contents during in the interactions, and is prompted to assess whether a target interaction would plausibly occur under a counterfactual-inspired intervention

Commercial LLMs are accessed through their respective APIs, while for local experiments, we deploy Qwen-Coder-30B on a workstation equipped with two NVIDIA A800 GPUs. For both commercial and locally deployed LLMs, the temperature is fixed to 0 to ensure deterministic outputs.

All inference procedures of the baseline methods follow identical preprocessing steps and prompt configurations on the Who&When benchmark to ensure fair and reproducible evaluation across different models.

To evaluate attribution quality, we focus on Step-level Accuracy, which measures the proportion of cases in which the exact erroneous step is correctly identified. Notably, we do not focus on Agent-level Accuracy, as in systems with a small number of agents, Agent-level Accuracy can be inflated. This occurs because it overlooks the more precise Step-level localization of failures. Given our emphasis on accurate localization at both the agent and step levels, we exclusively use Step-level Accuracy, rejecting the misleading Agent-level baseline.

Appendix D Robustness Across Independent Runs

Although we set the decoding temperature to 0 for all experiments, minor randomness in LLM inference may still lead to small fluctuations in the outputs. To evaluate the robustness of our results, we repeated the experiments using the three best-performing LLMs and report the mean and standard deviation over three independent runs.

As shown in Table 4, the variance across runs is extremely small for all methods and models. This indicates that the observed performance differences are stable and not caused by stochastic variations in the LLM outputs. In particular, DCFA consistently achieves the highest performance across all settings with minimal variance, demonstrating the robustness of the proposed failure attribution framework.

Method Algorithm-Generated Hand-Crafted
DeepSeek-R1-671B GPT-5 Gemini-2.5-pro DeepSeek-R1-671B GPT-5 Gemini-2.5-pro
All-at-Once 16.93 ±\pm 0.00 16.14 ±\pm 0.00 27.51 ±\pm 0.02 4.02 ±\pm 0.01 3.45 ±\pm 0.00 4.60 ±\pm 0.01
Step-by-Step 27.25 ±\pm 0.03 28.31 ±\pm 0.03 34.92 ±\pm 0.03 11.50 ±\pm 0.01 13.22 ±\pm 0.01 16.10 ±\pm 0.01
Binary-Search 33.86 ±\pm 0.01 31.75 ±\pm 0.03 26.46 ±\pm 0.02 5.75 ±\pm 0.01 10.92 ±\pm 0.01 5.75 ±\pm 0.01
A2P 16.40 ±\pm 0.04 16.40 ±\pm 0.05 31.75 ±\pm 0.05 6.32 ±\pm 0.01 12.64 ±\pm 0.01 9.77 ±\pm 0.01
ECHO 40.21 ±\pm 0.00 31.48 ±\pm 0.03 31.22 ±\pm 0.00 19.54 ±\pm 0.09 13.79 ±\pm 0.01 13.22 ±\pm 0.08
DCFA 53.44 ±\pm 0.01 47.09 ±\pm 0.04 46.69 ±\pm 0.01 21.38 ±\pm 0.04 22.41 ±\pm 0.01 21.84 ±\pm 0.03
Table 4: Robustness evaluation across three independent runs. We report mean performance and standard deviation. The extremely small variance indicates that the results are stable despite potential randomness in LLM inference.

Appendix E Detailed Computational Cost Analysis

E.1 Computational Cost Calculation

This section reports detailed empirical statistics of token usage and runtime for DCFA and baseline methods on the Who&When benchmark. For DCFA, the average prompt overhead is approximately 400 tokens. The average trace length is about 3k tokens for the Algorithm-Generated dataset and 10k tokens for the Hand-Crafted dataset. The GCA stage performs two full-context API calls, each taking roughly 16 seconds, resulting in an average runtime of about 32 seconds per instance. The LCE stage is executed locally using Qwen3-Coder-30B on two NVIDIA A800 GPUs. Each inference processes approximately 6k tokens and takes about 60 seconds. Since bidirectional dependency neighborhood search requires on average 4.5 such calls, the total LCE runtime is approximately 270 seconds. Consequently, the overall runtime of DCFA is around 300 seconds per trajectory.

Baseline methods differ primarily in how they process the interaction trace. All-at-Once analyzes the entire trace using a single LLM call of roughly 6k tokens. Search-by-Step processes one interaction at a time (about 50 tokens per step) and averages around 20 steps per trace, resulting in approximately 9k tokens. Binary-Search iteratively analyzes halves of the remaining trace and typically requires about five calls, leading to roughly 12k tokens. A2P uses a prompt of about 300 tokens and processes the full trace in a single call, while ECHO also uses a prompt of about 300 tokens but performs two full-context calls.

E.2 Trade-off Analysis

Although DCFA introduces additional computational overhead, it yields substantial performance improvements. In terms of computational complexity, the strongest baseline, ECHO, and the GCA component of DCFA both require two full-context LLM calls. Notably, the GCA module alone already achieves clear improvements over ECHO (+10.58% on the Algorithm-Generated dataset and +0.57% on the Hand-Crafted dataset), suggesting that the primary performance gains stem from enhanced reasoning over the trajectory. The LCE module, in contrast, relies on local LLMs and therefore incurs relatively additional cost. It further improves detection accuracy on longer trajectories (e.g., +3.45% on the Hand-Crafted dataset), indicating a favorable trade-off between attribution accuracy and computational efficiency.

E.3 Runtime Cost and Practicality

DCFA is designed as a post-hoc failure attribution framework, consistent with baselines such as ECHO and A2P, rather than as a real-time debugging module. Its primary goal is to analyze the complete system trace after a task failure and identify the decisive errors. Insights from this analysis can guide offline modifications to the MAS, including agent structures, context management strategies, or tool invocation policies, rather than adjusting the system during execution.

For online or real-time scenarios, a different design is required. Incremental reasoning can be applied at each interaction step, leveraging cached causal-inspired dependency subgraphs and restricting counterfactual-inspired evaluations to local neighborhoods. This avoids repeatedly processing the full trace and can substantially reduce runtime overhead, making failure attribution more practical for long-running or continuously operating multi-agent systems.

(a)
(b)
Figure 6: Representative causal-inspired dependency graphs constructed by Gemini-2.5-pro and GPT-5 on the Hand-crafted example 49.json. Differences are most pronounced in the critical failure region τ9\tau_{9}–τ12\tau_{12}.

Appendix F Supplementary Analysis of GCA-Induced Causal-inspired Dependency Graphs

We analyze the structural statistics of the causal-inspired dependency graphs constructed during the GCA stage, with a particular focus on cross-model variability and its implications for decisive error localization. To facilitate a concrete comparison, we further present representative causal-inspired dependency graphs generated by GPT-5 and Gemini-2.5-pro on a selected sample (49.json) from the Hand-crafted dataset.

F.1 Model Capacity and Deviations in Causal-inspired Dependency Graph Edge Density

Table 5 reports the average number of edges per graph across different LLMs on both the Algorithm-generated and Hand-crafted datasets.

Model Algorithm-generated Hand-crafted
Qwen3-Coder-30B 7.89 68.90
DeepSeek-R1-32B 8.24 46.39
Qwen3-235B 9.41 61.52
DeepSeek-R1-671B 9.25 57.09
GPT-5 10.12 55.40
Gemini-2.5-pro 8.76 56.38
Average 8.95 56.61
Table 5: Average number of edges per causal-inspired dependency graph generated during the GCA stage on the Algorithm-generated and Hand-crafted datasets.

A clear stratification emerges when comparing models of different capability. Less capable models, such as Qwen3-Coder-30B and DeepSeek-R1-32B, exhibit pronounced deviations from the overall mean, especially on the Hand-crafted dataset, producing overly dense or sparse causal-inspired dependency graphs (e.g., 68.90 edges for Qwen3-Coder-30B and 46.39 edges for DeepSeek-R1-32B). Such extreme deviations, in either direction, indicate unstable causal induction and low quality of the causal-inspired dependency graph generation, likely due to context degradation in LLMs.

By contrast, higher-capability models, such as GPT-5, Gemini-2.5-pro, and DeepSeek-R1-671B, maintain edge counts closer to the global average, suggesting a stronger ability to suppress irrelevant dependencies and preserve salient relations.

F.2 Implications for Decisive Error Localization.

To qualitatively assess how these structural differences affect downstream reasoning, Figure 6 visualizes the causal-inspired dependency graphs constructed by Gemini-2.5-Pro and GPT-5 on the same hand-crafted example, 49.json. Although both models successfully capture the overall dependency structure, their local structures in the critical failure region differ substantially.

Specifically, the GCA stage using Gemini-2.5-Pro initially hypothesizes τ9\tau_{9} as the decisive error, reflecting a more entangled dependency subgraph in the interval τ9\tau_{9}–τ12\tau_{12}. In contrast, GCA with GPT-5 constructs a more compact and hierarchical causal-inspired dependency structure over the same region, enabling it to identify τ12\tau_{12} as the decisive error.

Importantly, this discrepancy does not prevent DCFA with Gemini-2.5-Pro from ultimately recovering the correct failure source. Through the LCE module in the subsequent DCFA stage, the model performs counterfactual-inspired reasoning over the extracted local dependency subgraph, progressively evaluating and eliminating less relevant upstream candidates and converging on the decisive error τ12\tau_{12}. This observation highlights that while reliable global causal-inspired dependency graphs facilitate earlier localization, the proposed framework remains robust to structural noise through localized counterfactual-inspired refinement.

Appendix G Edge Growth and Long-Trace Scalability

Analysis of the causal-inspired dependency graphs constructed by DCFA shows that edge growth is nearly linear with trajectory length. In realistic multi-agent traces, each interaction typically links only to temporally or semantically adjacent steps, rather than to all preceding interactions. For example, in the Algorithm-generated dataset, traces average 8.7 steps with 8.95 edges, while in the Hand-crafted dataset, traces average 51.6 steps with 56.61 edges. This correspondence indicates that edge growth scales linearly rather than quadratically, mitigating combinatorial complexity.

Maintaining consistent structured reasoning over long traces remains challenging for LLM-based approaches. LCE addresses this challenge by performing localized exploration within the dependency graph, reducing the reasoning scope and keeping the additional search overhead roughly linear in trace length. Alternative cost-efficient strategies could further alleviate scaling issues, including heuristic graph pruning (e.g., limiting search depth or filtering low-confidence edges), sliding-window reasoning without full graph traversal, or hierarchical trace compression prior to graph construction. These approaches offer options for scaling DCFA to longer or more complex trajectories while retaining the benefits of localized dependency reasoning.

Appendix H Details for Case Study on Causal-inspired Dependency Graph Search

To demonstrate the effectiveness of DCFA, we present a case study based on a system trace from the Hand-crafted multi-agent system (ID: 49.json). In this scenario, the user query triggers interactions among three agents: Orchestrator, WebSurfer, and Assistant. The Orchestrator decomposes the task and coordinates the reasoning workflow, the WebSurfer retrieves external information, and the Assistant integrates the gathered evidence to produce the final answer.

Refer to caption
Figure 7: Key interaction steps of incorrect answer generation in trace 49.json.
Stage Identified Decisive Error Agent Supporting Evidence
GCA (Global) τ9\tau_{9} Orchestrator Ignore missing operator details
LCE (Local) τ12\tau_{12} Assistant Fabricate missing content
Table 6: Comparison between GCA and LCE Attributions.

The system trace captures each step of the multi-agent reasoning process, including information collection, context propagation, and intermediate reasoning actions. Specifically, the detail content of the user query is shown in the following box:

User Query In Unlambda, what exact character or text should be added to make the given code output “For penguins”? Answer with the name of the needed character if applicable.
Code: ‘r“““““‘.F.o.r. .p.e.n.g.u.i.n.si

In response, the multi-agent system generates the system trace consisting of fifteen interactions 𝒯={τ1,…,τ15}\mathcal{T}=\{\tau_{1},\dots,\tau_{15}\}. Each represents an interaction from one agent, as illustrated in Fig. 7. The system produces an incorrect answer “k”, whereas the correct missing character should be backtick “`”.

Within this trace, some steps involve minor deviations, such as incomplete retrieval or partial context propagation, which are generally correctable through downstream verification. In contrast, a decisive error occurs when the Assistant fabricates information based on incomplete or ambiguous input, producing content that is irreversible and misguides subsequent reasoning, ultimately leading to an incorrect outcome.

This overview sets the stage for a detailed case study analysis, where we examine how DCFA distinguishes contributory minor deviations from the decisive fabrication error and demonstrates its capability to pinpoint the root cause of failure in multi-agent reasoning.

Inference Phase of GCA

During global analysis, GCA first detects a potential minor deviation at τ8\tau_{8}, where the WebSurfer attempts to retrieve definitions of Unlambda operators but returns incomplete information.

Excerpt from τ8\tau_{8} WebSurfer: Searching for Unlambda operators “.”, “`”, and “r”…

Based on this, GCA builds a causal-inspired dependency graph centered on τ8\tau_{8} (Fig. 4) and ultimately identifies and attributes the system failure to τ9\tau_{9}, where the Orchestrator mistakenly assumes that the retrieved information is complete and proceeds without verification.

Excerpt from τ9\tau_{9} Orchestrator: Progress confirmed. Proceeding to analyze gathered information…
Refinement Phase of LCE.

To refine this hypothesis of GCA, the Local Counterfactual-inspired Enhancement (LCE) module focuses on the neighborhood around τ9\tau_{9}, forming a local subgraph (Fig. 4). It conducts counterfactual-inspired simulations to evaluate whether modifying prior steps would correct the outcome. Forward traversal from τ9\tau_{9} to τ8\tau_{8} shows no effect, while backward traversal identifies τ12\tau_{12} as the decisive cause.

Excerpt from τ12\tau_{12} Assistant: Summarizing Unlambda constructs… The backtick (‘) applies functions, the dot (.) outputs characters, and the “r” operator reads inputs. Based on this, adding “k” may terminate unwanted outputs.

LCE identifies τ12\tau_{12} as the decisive error, where the Assistant fabricates knowledge by falsely claiming that the WebSurfer’s incomplete results included operator definitions. This misinformation misleads subsequent reasoning, ultimately producing the wrong output in τ15\tau_{15}.

Decisive Error and DCFA Attribution.

Table 6 compares the decisive error attributions identified by the GCA and LCE modules. In this trace, τ8\tau_{8} corresponds to the WebSurfer attempting to retrieve essential information about Unlambda operators. Although the retrieved information is incomplete, this step constitutes a minor deviation: it is benign and can be effectively corrected by the system’s subsequent verification mechanisms, such as detecting missing fields and re-invoking tools to retrieve the omitted information. Therefore, τ8\tau_{8} does not constitute a decisive error.

τ9\tau_{9} represents the Orchestrator proceeding and propagating the partial context returned by WebSurfer. While the GCA module flags τ9\tau_{9} as a global-level misjudgment, this step primarily propagates the minor deviation from τ8\tau_{8} rather than introducing a fundamentally new error, making it contributory but not causative.

The decisive error occurs at τ12\tau_{12}, where the Assistant integrates the retrieved information and fabricates operator details, transforming a recoverable incompleteness into an irreversible factual error. This fabricated knowledge directly misguides subsequent reasoning and drives the final incorrect outcome. Unlike τ8\tau_{8}, the error at τ12\tau_{12} cannot be corrected through downstream verification, highlighting its role as the decisive upstream failure.

The DCFA framework identifies this decisive error through a two-stage reasoning process. First, GCA constructs a step-level causal-inspired dependency graph centered on minor deviations, capturing high-level dependencies and propagating influence to generate an initial hypothesis, here pointing to τ9\tau_{9}. Second, LCE performs a local neighborhood analysis around the hypothesis, employing counterfactual-inspired simulations to quantify the impact of each step on the final outcome. By evaluating the potential corrections in context, LCE isolates τ12\tau_{12} as the only step whose modification could prevent the error, thereby pinpointing the true decisive error.

This dual-view strategy, combining global reasoning over a causal-inspired dependency structure with local counterfactual-inspired validation, allows DCFA to transition from broad dependency mapping to precise error localization, effectively distinguishing contributory minor deviations from decisive errors. The approach not only enhances the interpretability of multi-agent reasoning traces but also identifies critical interactions whose correction is most likely to improve the final outcome, demonstrating clear advantages over purely global or local attribution methods.

(a)
(b)
Figure 8: Distribution of step-level deviations between predicted and true decisive errors. DCFA shows more concentrated distributions around zero, indicating substantially reduced bias compared to ECHO.

Appendix I Bias Reduction

To further validate the robustness of our causal-inspired attribution, we examine DCFA’s capability to mitigate bias in decisive error localization compared with baselines. Here, bias refers to the systematic deviation of predicted steps from the ground-truth decisive error, which can obscure the actual failure source and mislead subsequent debugging Li et al. (2022); Shrotriya et al. (2025).

We compare DCFA with the representative baseline, ECHO, on both Algorithm-generated and Hand-crafted traces from the Who&When benchmark Zhang et al. (2025). For each trace, we quantify the deviation between the predicted and true causal-step indices:

dDCFA=|tDCFA∗−t†|,dBaseline=|tBaseline∗−t†|d_{\text{DCFA}}\!=\!\left|t^{*}_{\text{DCFA}}\!-\!t^{{\dagger}}\right|,\quad d_{\text{Baseline}}\!=\!\left|t^{*}_{\text{Baseline}}\!-\!t^{{\dagger}}\right| (17)

where tDCFAt_{\text{DCFA}} and tBaselinet_{\text{Baseline}} are predicted indices, and t†t^{{\dagger}} denotes the ground-truth step. A smaller deviation implies lower bias and more reliable reasoning.

Figure 8 visualizes the deviation distributions for DCFA and ECHO. Across both DeepSeek-R1-671B and Gemini-2.5-pro backends, DCFA exhibits a notably sharper and more centralized distribution around zero. Overall, DCFA exhibits a smaller failure attribution offset than ECHO, indicating more accurate and less biased localization of the true decisive errors. This indicates that DCFA predictions consistently align closer to the ground-truth decisive error. Such improvement stems from the both global and local view causal-inspired enhancement of the DCFA, preventing error accumulation in step-level reasoning. Overall, the results demonstrate that DCFA reduces attribution bias, leading to more stable and interpretable failure diagnostics across synthetic and real-world traces.

Appendix J Related Works

J.1 LLM-based Multi-Agent Systems and System Failure

The recent progress of Large Language Model (LLM)-based agents has accelerated the development of Multi-Agent Systems (MAS), which have emerged as a promising paradigm for addressing complex, multi-step tasks that exceed the capabilities and token limitations of single, monolithic LLMs Yang et al. (2026); Sun et al. (2025); Pei et al. (2025); Xia et al. (2025). The principles underlying MAS, such as task decomposition, inter-agent communication, and emergent collective reasoning. Recent surveys have proposed systematic taxonomies of agent systems architectures Wu et al. (2023); Fourney et al. (2024), identifying core functional components required for effective coordination and execution. These components generally include explicit role assignment, hierarchical or iterative planning mechanisms, structured communication protocols, externalized memory modules, and evaluation and feedback mechanisms Yan et al. (2025).

Despite rapid progress, current agent systems remain limited in robustness, interpretability, and stability Zhang et al. (2025); Cemri et al. (2026). Agents often exhibit brittle coordination, unstable reasoning trajectories, and cascading errors, particularly in long-horizon or highly interactive environments Zhang et al. (2025); Cemri et al. (2026). Benchmarking efforts such as AgentBench Liu et al. (2024b) empirically demonstrate these weaknesses, showing that even state-of-the-art systems struggle to maintain coherence and consistency in collaborative, multi-step tasks. The recurrence of these failures exposes a critical research gap: understanding why and how LLM-based agent systems deviate from their intended execution pathways.

J.2 Failure Attribution in LLM-based Multi-agent Systems

Prior work has studied error analysis in agentic systems. AgentErrorBench with AgentDebug Zhu et al. (2025) locates mistakes that trigger cascading reasoning distortions in human-agent interactions in single-agent settings, whereas TRAIL Deshpande et al. (2025) analyzes errors in execution trajectories in multi-agent settings. These definitions, however, do not directly target the most actionable points for repairing task failures. Meanwhile, unlike single-agent systems with sequential execution and full information access, multi-agent systems involve heterogeneous roles, complex interaction topologies, partial observability, and non-trivial communication, producing intricate and interdependent dependency chains Cemri et al. (2026); Fourney et al. (2024); Wu et al. (2023). Thus, the focus shifts to identifying the earliest mistake whose correction could reverse overall system failure in multi-agent systems. The mistake, termed the decisive error Cemri et al. (2026); Fourney et al. (2024); Wu et al. (2023), represents the key point for restoring task success.

Recent research has begun to explore failure attribution in multi-agent systems using LLMs as diagnostic detectors Zhang et al. (2025); Zhang et al. (2026). Existing methods can be categorized into fine-tuning-based Zhang et al. (2026) and instruction-based Zhang et al. (2025); Cemri et al. (2026) paradigms. The fine-tuning-based approach, exemplified by AgenTracer Zhang et al. (2026), constructs specialized datasets to train dedicated models for failure diagnosis. While this approach offers strong controllability, it also incurs substantial costs for annotation, training, and maintenance. Instruction-based methods Zhang et al. (2025); Cemri et al. (2026), in contrast, rely on carefully designed prompts to guide LLMs in analyzing multi-agent system traces and identifying the decisive errors. For example, the Step-by-Step algorithm in the Who&When benchmark Zhang et al. (2025) formulates failure attribution as a sequential inspection process that progressively checks each interaction. Another algorithm in the benchmark, Binary-Search, adopts a divide-and-conquer strategy, recursively narrowing the trace scope until the decisive error is pinpointed. Additionally, Banerjee et al. Banerjee et al. (2025) propose hierarchical context representations and multi-perspective consensus to improve attribution reliability, while Weng et al. West et al. (2025) introduce A2P, a causal-guided attribution framework powered by the DeepScientist AI system.

Despite these advances, current methods West et al. (2025); Banerjee et al. (2025) rely primarily on the vanilla reasoning capabilities of LLMs, which often restricts their focus to superficial deviations and leads to performance degradation when processing long system traces. These limitations frequently result in inaccurate or incomplete attribution. To address these challenges, we propose a unified framework that incorporates structured reasoning and extracts local dependency chains to enable robust and reliable failure attribution in LLM-based multi-agent systems.

J.3 Causal Reasoning and Counterfactual Analysis via LLM

Recent advances have increasingly explored the use of large language models (LLMs) for causal reasoning and counterfactual-inspired analysis, including discovering, representing, and simulating causal relationships Su et al. (2025); Cheng et al. (2025); Luo et al. (2024). This line of research leverages the linguistic and reasoning capabilities of LLMs to extract and reason about causal structures from unstructured text. These studies show that LLMs can effectively extract events and propose candidate causal links with high recall, serving as powerful front-end modules for event causality analysis Liu et al. (2024a); Wang et al. (2024).

To improve reliability, several works Tong et al. (2024); Cohrs et al. (2025); Ban et al. (2025); Liu et al. (2025) integrate LLM-derived evidence with algorithmic estimators, aggregating multiple causal hypotheses or graph structures and taking their intersection to enhance robustness and filter out spurious relations. This hybrid approach demonstrates that LLMs can serve as high-level causal reasoning modules, while traditional estimators provide quantitative validation. Further research explores the use of LLMs as informative priors for causal graph discovery. Recent studies demonstrate that LLMs encode rich domain and commonsense knowledge that can guide structure learning and improve interpretability Darvariu et al. (2024); Jiralerspong et al. (). By injecting LLM-derived constraints or edge probabilities into graph search algorithms, these methods reduce sample complexity and produce more plausible causal graphs for downstream causal analysis.

Beyond causal structure discovery, causal inference and counterfactual simulation have been explored as mechanisms for testing causal hypotheses Liu et al. (2025); Zhu et al. (2024); Verma et al. (2025). Causal inference provides formal tools for reasoning about the effects of interventions Imai and Nakamura (2026), while counterfactual simulation assesses how altering an event or reasoning step might change subsequent outcomes. Recent works apply such techniques to LLM reasoning chains, using generated interventions to evaluate faithfulness and sensitivity to hypothetical changes in chain-of-thought reasoning Tutek et al. (2025); Yu et al. (2025). These findings suggest that LLMs can both represent causal structures and support the simulation of hypothetical interventions for evaluating outcome changes.

Inspired by these developments, our work adapts these ideas to failure attribution in LLM-based multi-agent systems. Rather than claiming formal causal inference, DCFA employs causal-inspired structured reasoning over system traces. Specifically, we construct a causal-inspired dependency graph to organize inter-agent dependencies and use counterfactual-inspired evaluation with LLM-mediated approximate interventions to assess how correcting candidate interactions may alter the final outcome. This dual-view design provides an interpretable basis for identifying the interaction that most plausibly constitutes the decisive error in long and unstructured multi-agent traces.

Algorithm 1 Global Causal-inspired Attribution (GCA)
 Input: trace 𝒯\mathcal{T}, ground truth 𝒴^\widehat{\mathcal{Y}}
 Output: hypothesis (τ⋆,r⋆)(\tau^{\star},r^{\star}), causal-inspired dependency graph 𝒢\mathcal{G}
 Minor Deviation Identification
 (𝒯e,ℛe)←MDI⁡(𝒯)(\mathcal{T}^{e},\mathcal{R}^{e})\leftarrow\mathrm{MDI}(\mathcal{T})
 Dependency Graph Construction
 𝒢←CGC⁡(𝒯,𝒯e,ℛe)\mathcal{G}\leftarrow\mathrm{CGC}(\mathcal{T},\mathcal{T}^{e},\mathcal{R}^{e})
 Hypothesis Generation
 (τ⋆,r⋆)←HG⁡(𝒴^,𝒯,𝒯e,ℛe,𝒢)(\tau^{\star},r^{\star})\leftarrow\mathrm{HG}(\widehat{\mathcal{Y}},\mathcal{T},\mathcal{T}^{e},\mathcal{R}^{e},\mathcal{G})
 return τ⋆,r⋆,𝒢\tau^{\star},r^{\star},\mathcal{G}

Appendix K Pseudocode Description of DCFA

This section briefly explains the pseudocode of the two modules in the DCFA framework.

Global Causal-inspired Attribution (GCA).

Algorithm 1 summarizes the overall global attribution procedure. The module first identifies candidate minor deviations in the system trace using the deviation identification process MDI⁡(⋅)\mathrm{MDI}(\cdot), which extracts interactions that exhibit abnormal behaviors together with brief textual explanations. Based on these candidates, the causal-inspired dependency graph construction procedure CGC⁡(⋅)\mathrm{CGC}(\cdot) organizes the trace into a directed causal-inspired dependency graph that encodes plausible dependencies between interactions. Finally, the hypothesis generation process HG⁡(⋅)\mathrm{HG}(\cdot) analyzes the trace, the detected deviations, and the constructed causal-inspired dependency graph to produce an initial hypothesis (τ⋆,r⋆)(\tau^{\star},r^{\star}) for the decisive error.

Algorithm 2 Local Counterfactual-inspired Enhancement (LCE)
 Input: trace 𝒯\mathcal{T}, ground truth Y^\widehat{Y}, causal-inspired dependency graph 𝒢\mathcal{G}, hypothesis τ⋆\tau^{\star}
 Output: decisive error τ∗\tau^{*}
 𝒯^←Correct⁡(Y^,𝒯,𝒢)\widehat{\mathcal{T}}\leftarrow\mathrm{Correct}(\widehat{Y},\mathcal{T},\mathcal{G})
 p←index​(τ⋆)p\leftarrow\text{index}(\tau^{\star})
 Forward Search
 while True do
  𝒩←{τj∣(p,j)∈ℰ}\mathcal{N}\leftarrow\{\tau_{j}\mid(p,j)\in\mathcal{E}\}
  evaluate Δ​𝒮​(τj)\Delta\mathcal{S}(\tau_{j}) for τj∈𝒩\tau_{j}\in\mathcal{N}
  p′←arg⁡maxτj∈𝒩​Δ​𝒮​(τj)p^{\prime}\leftarrow\arg\max_{\tau_{j}\in\mathcal{N}}\Delta\mathcal{S}(\tau_{j})
  if Δ​𝒮​(p′)>0\Delta\mathcal{S}(p^{\prime})>0 then
   p←p′p\leftarrow p^{\prime}
  else
   break
  end if
 end while
 p←index​(τ⋆)p\leftarrow\text{index}(\tau^{\star})
 Backward Search
 while True do
  𝒩←{τk∣(k,p)∈ℰ}\mathcal{N}\leftarrow\{\tau_{k}\mid(k,p)\in\mathcal{E}\}
  evaluate Δ​𝒮​(τk)\Delta\mathcal{S}(\tau_{k}) for τk∈𝒩\tau_{k}\in\mathcal{N}
  p′←arg⁡maxτk∈𝒩​Δ​𝒮​(τk)p^{\prime}\leftarrow\arg\max_{\tau_{k}\in\mathcal{N}}\Delta\mathcal{S}(\tau_{k})
  if Δ​𝒮​(p′)>0\Delta\mathcal{S}(p^{\prime})>0 then
   p←p′p\leftarrow p^{\prime}
  else
   break
  end if
 end while
 𝒯c←\mathcal{T}^{c}\leftarrow visited interactions
 τ∗←DED⁡(Y^,𝒯c,𝒢)\tau^{*}\leftarrow\mathrm{DED}(\widehat{Y},\mathcal{T}^{c},\mathcal{G})
 return τ∗\tau^{*}
Local Counterfactual-inspired Enhancement (LCE).

Algorithm 2 refines the global hypothesis by performing a bidirectional greedy search on the causal-inspired dependency graph. The procedure first constructs a corrected trace using Correct⁡(⋅)\mathrm{Correct}(\cdot) to provide a reference trajectory aligned with the ground-truth outcome. Starting from the hypothesis interaction τ⋆\tau^{\star}, the algorithm iteratively explores downstream and upstream neighbors in two search loops. At each step, candidate interactions are evaluated using counterfactual-inspired interventions, and the interaction with the largest positive marginal improvement Δ​𝒮\Delta\mathcal{S} becomes the next pivot. The union of visited interactions forms a compact dependency neighborhood, which is then evaluated by the decision function DED⁡(⋅)\mathrm{DED}(\cdot) to determine the final decisive error.

Appendix L Prompts for DCFA

This section presents the prompts used to implement the DCFA components described in Section 4. The prompts correspond to key reasoning steps in the Global Causal-inspired Attribution (GCA) and Local Counterfactual-inspired Enhancement (LCE) modules. For clarity and reproducibility, we provide representative prompt templates. Actual inputs (e.g., system traces, candidate events, and ground-truth answers) are dynamically inserted during execution.

Prompts for the GCA Module.

The GCA module performs global structured reasoning over the full system trace to generate an initial hypothesis for the decisive error.

First, the model identifies candidate minor deviations and constructs a causal-inspired dependency graph over the system trace. The prompt instructs the model to analyze the interaction events extracted from the system trace, detect potential mistake events, and identify dependency relationships among events based on the criteria of temporality, necessity, and sufficiency defined in Section 4.1. This step jointly performs minor deviation identification and causal-inspired dependency graph construction, producing both the set of candidate deviation events and the directed dependency relationships among them. The prompt used for this step is shown in Figure 9.

After the candidate deviations and causal-inspired dependency graph are obtained, the GCA module performs global reasoning to generate an initial hypothesis of the decisive error. Given the system trace, the set of candidate minor deviations, the explanations for each deviation, and the constructed causal-inspired dependency graph, the model evaluates how deviations propagate through the interaction chain and selects the event that most plausibly explains the final failure outcome. The corresponding prompt is illustrated in Figure 10.

Prompts for the LCE Module.

The LCE refines the hypothesis generated by GCA through localized counterfactual-inspired reasoning, comprising prompt-based steps for error selection.

First, in the counterfactual-inspired evaluation process, the model generates a corrected system trace. Given the original trace, the identified mistake events, and the ground-truth outcome, the prompt instructs the model to correct the erroneous events and causally affected downstream interactions while preserving the original event structure and formatting. This produces a fully corrected trace representing a counterfactual-inspired scenario in which the mistakes are fixed. The prompt is shown in Figure 11.

Next, for each counterfactual-inspired trace constructed during evaluation, the model simulates the corresponding system outcome based on the partially corrected interaction sequence. Starting from the corrected prefix and the remaining original interactions, the model reconstructs the downstream reasoning process and predicts the resulting final outcome. This prompt enables the model to estimate how correcting specific interactions affects the final result. The corresponding prompt is presented in Figure 12.

Finally, after the bidirectional dependency neighborhood search and counterfactual-inspired evaluation produce a set of candidate decisive error events, the model performs a final semantic decision and selection step. Given the candidate events, the causal-inspired dependency graph structure, and the ground-truth outcome, the model determines which interaction most plausibly constitutes the decisive error responsible for the observed system failure. The prompt is shown in Figure 13.

Figure 9: Prompt used for minor deviation identification and causal-inspired dependency graph construction. The model identifies surface-level mistake events as minor deviations, and establishes dependency relations among them according to temporality, necessity, and sufficiency criteria.
Figure 10: Prompt used for global causal-inspired structured reasoning and hypothesis generation. Given the system trace, candidate minor deviations, their explanations, and the constructed causal-inspired dependency graph, the model performs global reasoning over the dependency structure to analyze how errors may propagate and proposes an initial hypothesis for the decisive error.
Figure 11: Prompt used for causal-inspired log correction. The model corrects mistake events and causally affected events using ground-truth information while preserving the original event structure and format.
Figure 12: Prompt used for counterfactual-inspired outcome simulation. Given a partially corrected interaction trace, the model reconstructs the downstream reasoning process and predicts the resulting final outcome.
Figure 13: Prompt used for decisive error identification. The model analyzes candidate mistake events and selects the single event that is causally decisive for the final system failure.