DCFA: Dual-view Causal-inspired Attribution for Failure Reasoning in LLM-based Multi-agent Systems
Abstract
Large language model (LLM)-based multi-agent systems have experienced rapid growth in recent years. Despite their promise, such systems remain fragile, frequently exhibiting reasoning and coordination errors that can lead to system-level failures. Failure attribution in such systems relies on tracing natural language interactions among agents to identify the decisive error, which refers to the earliest action whose correction can reverse system failure. There are two key challenges: 1) Shallow attribution: Existing methods often capture only minor deviations, such as incomplete retrievals or formatting errors, which verification mechanisms can correct, while missing the decisive cause of system failure. 2) Contextual degradation: As the length of the system traces increases, the model’s reasoning ability rapidly deteriorates. To address these challenges, we propose DCFA, a training-free framework for failure attribution. DCFA integrates a global module that constructs structured causal-inspired dependency graphs from system traces to identify the initial decisive error, and a local module that applies local counterfactual-inspired reasoning to refine causal-inspired attribution. Experiments on the Who&When benchmark across six LLMs show that DCFA improves step-level accuracy by up to 8.27% over state-of-the-art baselines. Code is available in the public repository (https://github.com/wzhSteve/DCFA).
1 Introduction
With the emergence of large language models (LLMs), multi-agent systems (MAS) built upon LLM backends have rapidly evolved Becattini et al. (2025); Ronanki (2025). These systems enable autonomous collaboration among specialized agents for tasks such as code generation, knowledge search, and data analysis Huang et al. (2024); Pan et al. (2024); Li et al. (2024); Majdoub et al. (2025). Despite the impressive progress in orchestration and tool integration, LLM-based multi-agent systems remain susceptible to coordination and reasoning failures Cemri et al. (2026); Hammond et al. (2025); Shen et al. (2026). Errors such as misalignment in agents’ collaborations or improper tool usage can propagate across different agents along the system trace, leading to cascading reasoning errors that ultimately derail entire workflows Cemri et al. (2026); He et al. (2025). This fragility highlights the urgent need for failure attribution in LLM-based multi-agent systems Cemri et al. (2026); Hammond et al. (2025), which refers to the process of identifying the earliest actions whose correction could reverse system failure, known as decisive errors.
Recent efforts have begun to explore failure attribution in LLM-based multi-agent systems using LLMs as diagnostic detectors Zhang et al. (2025); Zhang et al. (2026). Existing approaches can be broadly classified into fine-tuning-based Zhang et al. (2026) and instruction-based Zhang et al. (2025); Cemri et al. (2026) paradigms. The fine-tuning-based approach, exemplified by AgenTracer Zhang et al. (2026), constructs specialized datasets to train dedicated models for failure analysis. In contrast, instruction-based methods Zhang et al. (2025); Cemri et al. (2026) rely on carefully designed prompts to guide LLMs in analyzing multi-agent system traces and identifying the decisive errors. However, both paradigms face inherent limitations. Fine-tuning-based approaches offer strong controllability but require extensive annotation and repeated training, resulting in high costs. Instruction-based methods avoid these costs under a training-free constraint by leveraging LLMs’ general reasoning ability, yet they often fail to reliably interpret long, multi-step, and unstructured system traces, leading to inaccurate failure attribution.
In this work, we study failure attribution in LLM-based multi-agent systems under the training-free setting to improve the extraction of decisive errors from unstructured multi-step traces. We identify two fundamental challenges: 1) Shallow attribution: Existing LLM-based attribution methods often focus on minor, recoverable deviations (e.g., incomplete retrieval or formatting errors Cemri et al. (2026); Banerjee et al. (2025)) that do not determine the final outcome. In contrast, the decisive error refers to the earliest action whose correction would reverse the system failure, truly driving the system to fail. For example, an agent may fabricate missing content when critical information is lost during the initial collection or retrieval process. Unlike missing information that can be subsequently verified and retrieved through downstream interactions, the fabrication error is difficult to verify and thus persists throughout the reasoning chain, misleading the system’s subsequent decisions. 2) Contextual degradation: As trace length and complexity increase, attribution performance degrades sharply. For instance, the Who&When benchmark Zhang et al. (2025) reports that the step-level attribution accuracy of state-of-the-art methods decreases substantially with growing trace length with averaging around 15% for traces of length 10, but dropping to nearly 0% for those approaching length 100.
To overcome these challenges, we propose a training-free framework, named the Dual-view Causal-inspired Failure Attribution (DCFA). DCFA comprises two key modules: the Global Causal-inspired Attribution (GCA) module and the Local Counterfactual-inspired Enhancement (LCE) module, which jointly perform global reasoning over a causal-inspired dependency structure and local counterfactual-inspired validation. To move attribution beyond surface-level deviations, GCA constructs a structured causal-inspired dependency graph and performs reasoning along it to uncover deeper dependency relationships underlying system failure. GCA extracts candidate minor deviations from the interaction system trace and builds a step-level causal-inspired dependency graph centered on these candidates. By performing global reasoning over the causal-inspired dependency graph across the entire system trace, GCA generates an initial hypothesis of the decisive error. To mitigate long-context degradation, LCE refines the global hypothesis through localized counterfactual-inspired reasoning. Starting from the hypothesized error, LCE performs a bidirectional search over the causal-inspired dependency graph to identify local dependency chains through which the error propagates, and evaluates how local corrections propagate through the reasoning chain and affect the final outcome via counterfactual-inspired evaluation. By evaluating these locally induced outcome changes, LCE progressively identifies the interaction that most plausibly constitutes the decisive error. In summary, our contributions are fourfold:
- •
We propose DCFA, a training-free framework for failure attribution in LLM-based multi-agent systems, integrating global reasoning with local counterfactual-inspired validation.
- •
To address shallow attribution, we construct causal-inspired dependency graphs to model interaction-level dependencies, enabling LLMs to move beyond minor deviations and identify decisive errors.
- •
To mitigate contextual degradation, we introduce localized counterfactual-inspired refinement that focuses reasoning on critical dependency chains, thereby improving attribution accuracy on long traces.
- •
Experiments on the Who&When benchmark with six LLMs show consistent gains, improving step-level accuracy by up to 8.27% over the state-of-the-art baselines.
2 Related Works
2.1 LLM-based Multi-Agent Systems and System Failure
Recent advances in LLM-based agents have driven rapid progress in multi-agent systems (MAS), enabling complex, long-horizon tasks through task decomposition and inter-agent communication Yang et al. (2026); Sun et al. (2025). Prior work has proposed systematic taxonomies of agent architectures Wu et al. (2023); Fourney et al. (2024), identifying components such as role assignment, planning, external memory, and feedback mechanisms, which aim to bolster the problem-solving capabilities of LLM-based multi-agent systems.
Despite their promise, LLM-based multi-agent systems remain brittle Zhang et al. (2025); Cemri et al. (2026). Empirical studies reveal frequent coordination failures and cascading reasoning errors, particularly in long or highly interactive execution traces Liu et al. (2024b). These observations highlight the need for methods to analyze the decisive error of system failure.
2.2 Failure Attribution in LLM-based Multi-agent Systems
Failure attribution in multi-agent systems aims to identify the earliest mistake whose correction could reverse the overall system failure, referred to as the decisive error Cemri et al. (2026); Fourney et al. (2024). Unlike prior work such as AgentErrorBench with AgentDebug Zhu et al. (2025), which locates mistakes that trigger cascading reasoning distortions, or TRAIL Deshpande et al. (2025), which identifies and analyzes all errors in trajectories, focusing on the decisive error directly targets the most actionable point for repairing task failure.
Recent approaches employ LLMs directly as diagnostic tools, including fine-tuning-based methods Zhang et al. (2026) and instruction-based methods Zhang et al. (2025); Cemri et al. (2026). Fine-tuned models offer strong control but require costly annotation and retraining. Instruction-based methods instead analyze execution traces via prompting. For instance, ECHO Banerjee et al. (2025) is equipped with hierarchical context representations and multi-perspective consensus to improve attribution performance, while A2P West et al. (2025) is designed to perform causal-guided attribution through prompt-based instruction.
Despite their effectiveness, existing methods rely heavily on LLM intrinsic reasoning, often resulting in shallow attribution and degraded performance on long traces. To address these limitations, we propose a unified framework that combines structured reasoning over a causal-inspired dependency graph with localized dependency-chain analysis to improve failure attribution.
2.3 Causal Reasoning and Counterfactual Analysis with LLMs
Recent work has explored LLMs for causal reasoning and counterfactual-inspired analysis, showing that they can extract events, propose candidate causal relations, and simulate hypothetical interventions from unstructured text Cheng et al. (2025); Luo et al. (2024). To improve robustness, several studies combine LLM-derived causal hypotheses with algorithmic validation or graph-based aggregation Tong et al. (2024); Ban et al. (2025). Others treat LLMs as informative priors for causal-inspired dependency graph discovery, leveraging their encoded commonsense knowledge to guide structure learning Darvariu et al. (2024); Jiralerspong et al. (). Counterfactual simulation has also been used to probe the sensitivity and faithfulness of LLM reasoning chains Tutek et al. (2025); Yu et al. (2025). These advances motivate our use of causal-inspired scaffolding and localized counterfactual-inspired validation for failure attribution in LLM-based multi-agent systems. Discussions of related work are deferred to Appendix J.
3 Problem Formulation
To respond to a user query, a multi-agent system generates a system trace including a sequence of interactions denoted as , where each interaction consists of the agent identifier and its corresponding content . The MAS produces a final output , which is expected to match the ground-truth outcome . A system failure occurs when the generated output deviates from the expected one, i.e., . Such failures may arise from one or multiple erroneous or misleading interactions within .
The objective of failure attribution is to identify the decisive error responsible for the system failure. When a single erroneous interaction accounts for the failure, the decisive error corresponds to the interaction whose correction can recover the correct outcome. When multiple errors jointly contribute to the failure and no single correction fully recovers the correct outcome, we define the decisive error as the interaction whose correction yields the largest reduction in the system failure.
Formally, a failure attribution model is defined as:
| (1) |
where denotes the ground-truth decisive error, i.e., the interaction whose correction either recovers the correct outcome when a single-step correction suffices, or maximally reduces system failure when multiple errors jointly contribute to the failure.
4 Methodology
In this section, we introduce the proposed DCFA framework (Fig. 2), which consists of the GCA (Sec. 4.1) and LCE (Sec. 4.2) modules. To overcome shallow attribution, GCA reasons over full system traces to uncover latent dependency relationships, constructing a causal-inspired dependency graph and identifying an initial hypothesis of the decisive error from a global perspective. Then, to further mitigate contextual degradation, the LCE module locally refines this hypothesis through counterfactual-inspired evaluation, using LLM-mediated approximate interventions to assess how correcting candidate interactions may alter the final outcome. This strategy first establishes a global causal-inspired dependency structure and then locally validates its critical links, enabling more interpretable failure attribution for unstructured and long system traces.
4.1 Global Causal-inspired Attribution Module
Interactions in system traces influence system outcomes through chained dependencies, requiring attribution methods that capture both minor deviations and how their effects propagate through the trace. Directly analyzing unstructured traces, LLMs often overemphasize salient deviations without sufficiently considering their downstream dependencies and influence on the final outcome. To address this, GCA detects candidate deviations, constructs a directed causal-inspired dependency graph that links them to the full trace, and leverages graph-aware global reasoning to generate an initial hypothesis of the decisive error.
4.1.1 Minor Deviation Identification
Minor deviations are directly observable but typically benign faults in system traces, such as malformed outputs or incomplete information retrieval Cemri et al. (2026); Hammond et al. (2025). These deviations are often recoverable through built-in verification, tool re-invocation Barta et al. (2025); Zhou et al. (2025). However, minor deviations are not necessarily decisive errors, where they may serve as early symptoms of deeper failures.
Based on this motivation, GCA first extracts a set of candidate minor deviations that represent plausible manifestations of incorrect behaviors. Rather than prematurely determining which interaction constitutes the decisive error, GCA aims to retain all potentially relevant interactions for subsequent structured reasoning and refinement over the causal-inspired dependency graph.
Specifically, a designed process is utilized to detect the interactions of minor deviations and produce short natural-language explanations for each identified minor deviation:
| (2) |
where is the set of candidate interactions that are identified as minor deviations, are the textual reasons associated with each candidate.
4.1.2 Dependency Graph Construction
Having identified the candidate minor deviations and their corresponding textual reasons, GCA constructs a structured causal-inspired dependency graph that encodes plausible dependency relationships among interactions in the system trace . The construction is centered on the detected minor deviations and their associated reasoning cues , ensuring that the resulting structure reflects interpretable, error-centric reasoning pathways.
Formally, the causal-inspired dependency graph is denoted as , where is the set of step indicators corresponding to all interactions, and is the set of directed edges representing putative dependency relationships. Specifically, an edge is denoted as , indicating a directed dependency from interaction to , where precedes and potentially influences .
Dependency Condition Evaluation.
A directed edge is admitted into only if it satisfies a set of structured dependency criteria Zeng et al. (2026); Goldberg (2019). For each ordered pair with step indices , GCA evaluates three complementary properties: temporality, necessity, and sufficiency.
Temporality enforces the directional ordering constraint that a preceding interaction must occur before the downstream interaction:
| (3) |
Necessity tests whether the would fail to occur if the upstream step is removed, using a counterfactual-inspired judgment function :
| (4) |
Sufficiency assesses whether enforcing is sufficient to plausibly induce the outcome at :
| (5) |
An edge is added only when the temporal, necessity, and sufficiency conditions are all satisfied. This conservative rule filters out spurious correlations, retaining causally supported dependencies.
Causal-inspired Structured Reasoning.
Given the structured dependency criteria, GCA performs causal-inspired structured reasoning over the system trace , focusing on the identified minor deviations to construct the causal-inspired dependency graph . The graph is constructed by incorporating these criteria into instruction sets :
| (6) |
where denotes the dependency graph construction procedure.
4.1.3 Hypothesis Generation
Based on the causal-inspired dependency graph and the candidate minor deviation set , GCA performs a global analysis to propose an initial hypothesis of the decisive error of the system failure. The hypothesis is produced by the hypothesis generation process on a compact evidence bundle that includes: the ground truth outcome , the entire system trace , the candidate minor deviations , the corresponding textual reasons to , and the constructed graph . Formally, the hypothesis generation is formulated as follows:
| (7) |
where the output is the hypothesis of the decisive error, and is the corresponding reason.
4.2 Local Counterfactual-inspired Enhancement Module
Practical system traces are often long and dense Zhang et al. (2025), which exacerbates performance degradation in failure attribution based on full traces. To address this limitation, we introduce the LCE module, which refines by focusing on a compact local dependency neighborhood. LCE performs a bidirectional search over the causal-inspired dependency graph and uses counterfactual-inspired evaluation with LLM-mediated approximate interventions to assess how correcting candidate interactions may alter the final outcome. By iteratively evaluating local outcome changes, LCE identifies the interaction that most plausibly constitutes the decisive error.
4.2.1 Bidirectional Dependency Neighborhood Search
The bidirectional search aims to identify a minimal set of interactions that provide relevant evidence for explaining the system failure and to localize the decisive error by tracing dependency propagation across upstream and downstream interactions. Starting from the index of the global hypothesis , LCE initializes the pivot as and explores the causal-inspired dependency graph in both forward and backward directions.
The forward search traces downstream dependencies to capture how the hypothesized error may propagate, while the backward search examines upstream dependencies to identify prior misreasoning or missing premises. In both directions, interactions are selected greedily based on counterfactual-inspired evaluation, which assesses how modifying a candidate interaction may affect the final outcome. The union of the forward and backward results forms a locally refined dependency subchain centered on .
Counterfactual-inspired Evaluation
Central to LCE is a counterfactual-inspired evaluation function that assesses the potential causal importance of an interaction by measuring the change in outcome alignment when that interaction is hypothetically corrected instead.
We first construct a fully corrected trace by applying a correction function conditioned on the ground-truth outcome , the original trace , extracted erroneous interactions , their corresponding reasons , and the causal-inspired dependency graph :
| (8) |
where and denotes the corrected version of .
Using , we construct a counterfactual-inspired trace for each interaction by combining the corrected prefix up to step with the original suffix from step , which is denoted as . This represents a scenario in which only the first interactions have been corrected. The counterfactual-inspired trace is then evaluated via a simulation process:
| (9) |
We measure the semantic alignment between the outcome and the ground truth using a pretrained encoder and cosine similarity:
| (10) |
The marginal attributional impact of correcting is defined as
| (11) |
which measures the incremental improvement in outcome alignment obtained by correcting the -th interaction after all preceding interactions have been corrected.
Bidirectional Greedy Search
Using the counterfactual-inspired scores, LCE performs a greedy bidirectional search over the causal-inspired dependency graph .
In the forward search, given the pivot at iteration , we collect all direct successors denoted as , where denotes the edge set of . The successor with the largest positive marginal contribution is selected:
| (12) |
If , the forward candidate set, denoted as , is expanded and the pivot updated as . Otherwise, the forward search terminates.
Symmetrically, the backward search traces incoming edges to identify upstream causes denoted as . The predecessor with the strongest positive effect is selected:
| (13) |
If , the backward candidate set, denoted as , is expanded and the pivot updated as . This process continues until no further positive contributions are found.
These two searches yield a local dependency neighborhood that captures both antecedent causes and propagated effects surrounding .
4.2.2 Decisive Error Decision
The forward and backward candidate sets are merged with the global hypothesis to form the final local candidate set . Each element in corresponds to an interaction whose counterfactual-inspired correction yields a positive improvement in outcome alignment.
Unlike GCA, which analyzes the full trace, LCE operates only on the interaction subset , avoiding context degradation. This evaluation is formalized by a decisive-error decision function:
| (14) |
where denotes the final identified decisive error.
5 Experimental Design
5.1 Dataset
We evaluate DCFA on the Who&When benchmark Zhang et al. (2025), which is currently the only widely adopted benchmark specifically designed for failure attribution in LLM-based multi-agent systems and serves as the common evaluation testbed for existing baselines West et al. (2025); Banerjee et al. (2025). The dataset contains 184 multi-agent system traces spanning both Algorithm-generated and Hand-crafted settings. Specifically, 126 traces are generated using CaptainAgent-based systems Wu et al. (2023), and 58 traces are manually curated from systems such as Magnetic-One Fourney et al. (2024). The tasks cover web-navigation and general reasoning scenarios derived from AssistantBench Yoran et al. (2024) and GAIA Mialon et al. (2023). Each trace is annotated with fine-grained failure labels, including the decisive error and a natural-language explanation. Details are provided in Appendix A.
5.2 Baselines
We compare DCFA against five representative baselines that reflect different LLM-based failure attribution paradigms: All-at-Once, Step-by-Step, and Binary-Search from the Who&When benchmark Zhang et al. (2025), as well as A2P West et al. (2025) and ECHO Banerjee et al. (2025). All methods rely on LLMs for attribution and are evaluated under identical inference settings. We test both open-source and commercial models, including Qwen3-Coder-30B, Qwen3-235B, DeepSeek-R1-32B, DeepSeek-R1-671B, GPT-5, and Gemini-2.5-pro. Details are deferred to Appendix B.
| Model | Algorithm-generated | Hand-crafted | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| All-at-Once | Step-by-Step | Binary-Search | A2P | ECHO | DCFA | All-at-Once | Step-by-Step | Binary-Search | A2P | ECHO | DCFA | |
| Qwen3-Coder-30B | 11.90 | 20.63 | 13.49 | 12.70 | 23.02 | 37.30 | 1.72 | 10.34 | 6.67 | 8.62 | 13.79 | 15.52 |
| DeepSeek-R1-32B | 22.22 | 29.37 | 30.16 | 19.84 | 30.95 | 42.06 | 5.17 | 15.52 | 6.90 | 6.90 | 18.97 | 20.69 |
| Qwen3-235B | 30.95 | 25.40 | 21.43 | 29.37 | 36.50 | 43.65 | 3.45 | 15.52 | 12.07 | 3.45 | 17.24 | 22.41 |
| DeepSeek-R1-671B | 16.67 | 26.98 | 32.54 | 15.87 | 40.48 | 53.17 | 3.45 | 12.07 | 5.17 | 6.90 | 20.69 | 22.41 |
| GPT-5 | 15.87 | 27.78 | 31.75 | 15.87 | 31.75 | 47.62 | 3.45 | 13.79 | 10.34 | 13.79 | 15.52 | 24.14 |
| Gemini-2.5-pro | 26.98 | 34.13 | 25.40 | 32.54 | 30.95 | 46.03 | 5.17 | 15.52 | 6.90 | 10.34 | 13.79 | 18.97 |
| Average | 20.77 | 27.72 | 25.13 | 21.03 | 32.28 | 44.79 | 3.74 | 13.46 | 8.33 | 8.33 | 16.67 | 20.69 |
5.3 Implementation Details
In DCFA, the GCA stage is performed by commercial LLMs using the same configurations as the baselines. The LCE stage is executed with a locally deployed Qwen-Coder-30B model. The judgment function used for dependency condition evaluation in GCA is implemented via structured prompting, where the LLM is queried to perform binary judgments. All LLMs are run with the temperature set to 0 to ensure deterministic inference. All baselines follow identical preprocessing pipelines and prompt configurations as specified in their original papers and released codebases to ensure fair comparison. As ECHO does not provide an official implementation, we reimplement it based on the descriptions in its paper.
We evaluate attribution quality using Step-level Accuracy, which directly measures precise failure localization and avoids the inflation effects of agent-level metrics. Additional implementation and evaluation details are provided in Appendix C.
5.4 Results and Analysis
5.4.1 Overall Performance
Table 1 reports step-level accuracy of DCFA and five baselines on the Algorithm-generated and Hand-crafted datasets. Compared with the strongest baseline on each dataset, DCFA achieves average improvements of 12.51% and 4.02%, respectively. On the Algorithm-generated dataset, DeepSeek-R1 achieves the highest accuracy, likely benefiting from reasoning-oriented training that aligns with DCFA’s structured causal-inspired dependency graph construction Guo et al. (2025), while GPT remains competitive on longer traces due to its strong long-context reasoning Leon (2025). In contrast, most baselines rely primarily on prompt-induced implicit reasoning, making their attributions less sensitive to the latent dependency relations in multi-agent interactions. For instance, ECHO focuses on observable errors such as tool failures or formatting issues. While these explicit patterns can improve detection accuracy, the resulting deviations are often minor rather than decisive. By tracing such surface symptoms back to their underlying causes, DCFA enables more precise and causally grounded failure attribution. Examples are presented in the case study (Sec. 5.4.4), with additional analysis of GCA-induced causal-inspired dependency graphs in Appendix F.
| Model | Algorithm-generated | Hand-crafted | ||
|---|---|---|---|---|
| w/o LCE | DCFA | w/o LCE | DCFA | |
| Qwen3-Coder-30B | 33.33 | 37.30 | 12.07 | 15.52 |
| DeepSeek-R1-32B | 40.48 | 42.06 | 15.52 | 20.69 |
| Qwen3-235B | 42.06 | 43.65 | 20.69 | 22.41 |
| DeepSeek-R1-671B | 53.17 | 53.17 | 18.97 | 22.41 |
| GPT-5 | 44.44 | 47.62 | 22.41 | 24.14 |
| Gemini-2.5-pro | 43.65 | 46.03 | 13.79 | 18.97 |
5.4.2 Ablation on LCE
Since LCE operates on the causal-inspired dependency graph and the hypothesis of the decisive error produced by GCA, we compare GCA-only with DCFA. Table 2 reports an ablation study of DCFA with and without LCE. Adding LCE consistently improves step-level accuracy on both datasets, with 3.45% average improvement on the Hand-crafted dataset and 1.93% average improvement on the Algorithm-generated dataset.
Notably, GCA alone exhibits a noticeable performance drop on the Hand-crafted dataset, where traces are substantially longer (average length ) than those in the Algorithm-generated dataset (average length ). In contrast, LCE remains effective even when global attribution degrades under long-context conditions. This robustness arises because LCE identifies the local dependency chains with the greatest corrective impact through bidirectional search and counterfactual-inspired evaluation, capturing both upstream causes and downstream effects of the error localized by GCA. By focusing the LLM on higher-level and more specific causal-inspired structures, LCE mitigates context degradation and improves attribution precision.
5.4.3 Performance on Varying Trace Lengths
We evaluate DCFA on the Hand-crafted dataset using DeepSeek-R1-671B and 32B, comparing it with all baselines across five context-length levels defined by the Who&When benchmark Zhang et al. (2025).
As shown in Fig. 3(a) and Fig. 3(b), DCFA outperforms all baselines across nearly all levels, with the largest gains observed at Level 1 and maintained through Level 5. One exception occurs at Level 2 with DeepSeek-R1-32B, where DCFA slightly underperforms ECHO. This may be due to the reduced reliability of weaker LLMs in following the structured dependency reasoning required by DCFA, which can introduce additional variance in LLM-mediated evaluation. Nevertheless, DCFA remains robust overall, achieving strong performance across varying trace lengths.
5.4.4 Case Study on Causal-inspired Dependency Graph Search
We analyze a system trace from a hand-crafted multi-agent system (ID: 49.json), where three agents (Orchestrator, WebSurfer, and Assistant) collaboratively solve an Unlambda debugging query (see details in Appendix H). The execution trace results in an incorrect prediction of “k”, whereas the correct missing character is the backtick “`”.
Inference Phase of GCA.
GCA first identifies multiple erroneous or misleading interactions that are potentially involved in the failure. Specifically, it detects a minor deviation at , where WebSurfer returns incomplete operator definitions, and further identifies , where the Orchestrator propagates the partial information without further validation. These errors form a dependency chain and jointly contribute to the subsequent failure, rather than constituting independent failure sources. Based on the resulting causal-inspired dependency graph, GCA localizes the relevant error region for further refinement, as shown in Figure 4.
Refinement Phase via LCE.
LCE then conducts localized, bidirectional counterfactual-inspired reasoning over the error region identified by GCA. By examining the corrective effects of candidate interactions along the dependency graph, LCE determines that has the greatest impact on mitigating the system failure. At , the Assistant fabricates a solution based on incomplete semantics, which further propagates the accumulated errors toward the incorrect final outcome. Although correcting or can mitigate the error propagation, neither correction is sufficient to fully resolve the failure. In contrast, correcting provides the largest reduction in the failure, thereby identifying it as the decisive error. The refinement process is illustrated in Figure 4.
Decisive Error and DCFA Attribution.
This case exemplifies the multi-error setting in our problem formulation, where multiple errors jointly contribute to the system failure while no single correction among the earlier errors fully recovers the correct outcome. Specifically, and contribute to the error propagation and can be corrected to mitigate the failure, whereas has the greatest corrective effect and is therefore identified as the decisive error. By combining global causal-inspired dependency search with local counterfactual-inspired refinement, DCFA successfully distinguishes contributing errors from the decisive error and accurately attributes the system failure to . Additional details are provided in Appendix H.
5.4.5 Computational Cost
We analyze DCFA’s computational cost by decomposing it into GCA and LCE. GCA relies on commercial LLM API calls, while LCE is executed locally and adds only inference-time overhead. Following Zhang et al. (2025), output-token costs are ignored unless stated otherwise.
GCA performs two full-context LLM calls: deviation detection with dependency graph construction, and decisive-error refinement. Since outputs are small relative to the input context, the total cost can be approximated as
| (15) |
where is the prompt overhead, the average trace length, and the token length per step. Notably, the strongest baseline ECHO requires the same number of full-context calls, resulting in identical computational complexity, while GCA alone already yields substantial accuracy improvements.
The LCE module runs locally and introduces additional inference overhead. Its complexity can be approximated as
| (16) |
where is the average number of bidirectional neighborhood expansions. Although LCE increases runtime, it further improves attribution accuracy, especially on longer trajectories.
Table 3 summarizes the approximate runtime, token consumption, and attribution accuracy of all compared methods on the Who&When benchmark. DCFA without LCE incurs a cost comparable to ECHO while achieving substantially higher accuracy ( vs. ). Adding LCE increases the runtime from 32s to 300s and token consumption from 13k to 32k, but further improves accuracy to . Detailed runtime statistics, token usage, and comprehensive comparisons with baseline methods are provided in Appendix E.
| Method | Runtime (s) | Tokens (k) | Accuracy |
|---|---|---|---|
| All-at-Once | 15 | 6 | 12.26 |
| Search-by-Step | 61 | 9 | 20.59 |
| Binary-Search | 25 | 12 | 16.73 |
| A2P | 15 | 6 | 14.68 |
| ECHO | 31 | 12 | 24.48 |
| DCFA w/o LCE | 32 | 13 | 30.05 |
| DCFA | 300 | 32 | 32.74 |
6 Conclusion
In this work, we propose DCFA, a training-free framework for failure attribution in LLM-based multi-agent systems. Combining global causal-inspired dependency graph analysis with local counterfactual-inspired reasoning, DCFA identifies decisive errors and performs robustly across traces of varying lengths. Experiments show it improves step-level attribution accuracy by up to 8.27% over state-of-the-art baselines and remains effective on challenging long-context traces.
Limitations
Dependence on LLM Capability.
DCFA relies on the reasoning inference capabilities of the underlying large language models. When system traces become very long, LLMs may struggle to maintain coherent dependency representations, which can reduce the accuracy of global causal-inspired dependency graph construction. Similarly, when using smaller or less capable models for local counterfactual-inspired refinement, the precision of decisive error identification may be constrained.
Scalability to Long or Complex Traces.
Although the global-local combination in DCFA improves robustness to moderately long execution traces, extremely long or highly branching multi-agent interactions can still pose challenges. In such cases, both global causal-inspired attribution and local counterfactual-inspired reasoning may become less reliable, limiting applicability to systems with deeply nested or prolonged interaction patterns.
Attribution of Intrinsic LLM Reasoning Errors.
Compared with the decisive errors that reflect failures within the MAS workflow. Errors originating from inherent LLM hallucinations may be difficult to attribute accurately.
Ethical Considerations
This research is intended as a diagnostic tool to support failure analysis in LLM-based multi-agent systems, rather than an automated mechanism for judgment or accountability. Its attribution results depend on the reasoning behavior of underlying language models and may be imperfect, especially when failures stem from intrinsic model errors or ambiguous interactions. Over-reliance on such automated explanations could lead to misinterpretation of responsibility if used without human oversight. Therefore, we highlight that the proposed DCFA should be applied with caution in high-stakes settings and used to assist, not replace, human analysis of system failures.
Acknowledgment
This work was supported in part by the National Natural Science Foundation of China under Grants 62572346 and 62322208.
References
- Integrating large language model for improved causal discovery. IEEE Transactions on Artificial Intelligence. Cited by: §J.3, §2.3.
- Where did it all go wrong? a hierarchical look into multi-agent error attribution. In NeurIPS 2025 Workshop on Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling, Cited by: §J.2, §J.2, Appendix B, §1, §2.2, §5.1, §5.2.
- Measuring the robustness of multi-agent reinforcement learning systems under partial agent failure. In Proceedings of the Intelligent Robotics FAIR 2025, pp. 58–63. Cited by: §4.1.1.
- SALLMA: a software architecture for llm-based multi-agent systems. In 2025 IEEE/ACM International Workshop New Trends in Software Architecture (SATrends), pp. 5–8. Cited by: §1.
- Why do multi-agent llm systems fail?. Advances in Neural Information Processing Systems 38. Cited by: §J.1, §J.2, §J.2, §1, §1, §1, §2.1, §2.2, §2.2, §4.1.1.
- A survey of event causality identification: taxonomy, challenges, assessment, and prospects. ACM Computing Surveys 58 (3), pp. 1–37. Cited by: §J.3, §2.3.
- Large language models for causal hypothesis generation in science. Machine Learning: Science and Technology 6 (1), pp. 013001. Cited by: §J.3.
- Large language models are effective priors for causal graph discovery. arXiv preprint arXiv:2405.13551. Cited by: §J.3, §2.3.
- Trail: trace reasoning and agentic issue localization. arXiv preprint arXiv:2505.08638. Cited by: Appendix A, §J.2, §2.2.
- Magentic-one: a generalist multi-agent system for solving complex tasks. arXiv preprint arXiv:2411.04468. Cited by: Appendix A, §J.1, §J.2, §2.1, §2.2, §5.1.
- The book of why: the new science of cause and effect: by judea pearl and dana mackenzie, basic books (2018). isbn: 978-0465097609.. Taylor & Francis. Cited by: §4.1.2.
- Deepseek-r1: incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948. Cited by: §5.4.1.
- Multi-agent risks from advanced ai. arXiv preprint arXiv:2502.14143. Cited by: §1, §4.1.1.
- SentinelAgent: graph-based anomaly detection in multi-agent systems. arXiv preprint arXiv:2505.24201. Cited by: §1.
- Romas: a role-based multi-agent system for database monitoring and planning. arXiv preprint arXiv:2412.13520. Cited by: §1.
- Causal inference with generative artificial intelligence: application to texts as treatments. Journal of the American Statistical Association (just-accepted), pp. 1–27. Cited by: §J.3.
- [17] Efficient causal graph discovery using large language models. In ICLR 2024 Workshop: How Far Are We From AGI, Cited by: §J.3, §2.3.
- GPT-5 and open-weight large language models: advances in reasoning, transparency, and control. Information Systems, pp. 102620. Cited by: §5.4.1.
- A survey on llm-based multi-agent systems: workflow, infrastructure, and challenges. Vicinagearth 1 (1), pp. 9. Cited by: §1.
- Training data debugging for the fairness of machine learning software. In Proceedings of the 44th International Conference on Software Engineering, pp. 2215–2227. Cited by: Appendix I.
- Identifying while learning for document event causality identification. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 3815–3827. Cited by: §J.3.
- Agentbench: evaluating llms as agents. In International Conference on Learning Representations, Vol. 2024, pp. 52989–53046. Cited by: §J.1, §2.1.
- Large language models and causal inference in collaboration: a comprehensive survey. Findings of the Association for Computational Linguistics: NAACL 2025, pp. 7668–7684. Cited by: §J.3, §J.3.
- Open event causality extraction by the assistance of llm in task annotation, dataset, and method. In Proceedings of the Workshop: Bridging Neurons and Symbols for Natural Language Processing and Knowledge Graphs Reasoning (NeusymBridge)@ LREC-COLING-2024, pp. 33–44. Cited by: §J.3, §2.3.
- Towards adaptive software agents for debugging. In Proceedings of the 33rd ACM International Conference on the Foundations of Software Engineering, pp. 636–640. Cited by: §1.
- Gaia: a benchmark for general ai assistants. In The Twelfth International Conference on Learning Representations, Cited by: Appendix A, Appendix A, §5.1.
- Building multi-agent copilot towards autonomous agricultural data management and analysis. In 2024 IEEE International Conference on Big Data (BigData), pp. 4384–4393. Cited by: §1.
- Flow-of-action: sop enhanced llm-based multi-agent system for root cause analysis. In Companion Proceedings of the ACM on Web Conference 2025, pp. 422–431. Cited by: §J.1.
- Facilitating trustworthy human-agent collaboration in llm-based multi-agent system oriented software engineering. In Proceedings of the 33rd ACM International Conference on the Foundations of Software Engineering, pp. 1333–1337. Cited by: §1.
- Metacognitive self-correction for multi-agent system via prototype-guided next-execution reconstruction. In Findings of the Association for Computational Linguistics: ACL 2026, pp. 23320–23337. Cited by: §1.
- Navigating the path to ethical & responsible ai integration in health & life sciences with human and machine collaboration. Journal of the Society for Clinical Data Management 5 (1), pp. 1–12. Cited by: Appendix I.
- Enhancing event causality identification with llm knowledge and concept-level event relations. In Proceedings of the 31st International Conference on Computational Linguistics, pp. 7403–7414. Cited by: §J.3.
- LLM-based multi-agent decision-making: challenges and future directions. IEEE Robotics and Automation Letters. Cited by: §J.1, §2.1.
- Automating psychological hypothesis generation with ai: when large language models meet causal graph. Humanities and Social Sciences Communications 11 (1), pp. 896. Cited by: §J.3, §2.3.
- Measuring chain of thought faithfulness by unlearning reasoning steps. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pp. 9946–9971. Cited by: §J.3, §2.3.
- Causal ai scientist: facilitating causal data science with large language models. In NeurIPS 2025 AI for Science Workshop, Cited by: §J.3.
- Event causality identification with synthetic control. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp. 1725–1737. Cited by: §J.3.
- Abduct, act, predict: scaffolding causal inference for automated failure attribution in multi-agent systems. arXiv preprint arXiv:2509.10401. Cited by: §J.2, §J.2, Appendix B, §2.2, §5.1, §5.2.
- Autogen: enabling next-gen llm applications via multi-agent conversation. arXiv preprint arXiv:2308.08155. Cited by: Appendix A, §J.1, §J.2, §2.1, §5.1.
- Demystifying llm-based software engineering agents. Proceedings of the ACM on Software Engineering 2 (FSE), pp. 801–824. Cited by: §J.1.
- Beyond self-talk: a communication-centric survey of llm-based multi-agent systems. arXiv preprint arXiv:2502.14321. Cited by: §J.1.
- Agentnet: decentralized evolutionary coordination for llm-based multi-agent systems. Advances in Neural Information Processing Systems 38, pp. 107309–107336. Cited by: §J.1, §2.1.
- Assistantbench: can web agents solve realistic and time-consuming tasks?. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp. 8938–8968. Cited by: Appendix A, Appendix A, §5.1.
- Causaleval: towards better causal reasoning in language models. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pp. 12512–12540. Cited by: §J.3, §2.3.
- Zero-shot event causality identification via multisource evidence fuzzy aggregation with large language models. IEEE Transactions on Fuzzy Systems. Cited by: §4.1.2.
- Agentracer: who is inducing failure in the llm agentic systems?. In International Conference on Learning Representations, Vol. 2026, pp. 11377–11399. Cited by: §J.2, §1, §2.2.
- Which agent causes task failures and when? On automated failure attribution of LLM multi-agent systems. In Proceedings of the 42nd International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 267, pp. 76583–76599. Cited by: Appendix A, §J.1, §J.2, Appendix B, Appendix I, §1, §1, §2.1, §2.2, §4.2, §5.1, §5.2, §5.4.3, §5.4.5.
- SHIELDA: structured handling of exceptions in llm-driven agentic workflows. arXiv preprint arXiv:2508.07935. Cited by: §4.1.1.
- Where llm agents fail and how they can learn from failures. arXiv preprint arXiv:2509.25370. Cited by: Appendix A, §J.2, §2.2.
- Causal inference with latent variables: recent advances and future prospectives. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pp. 6677–6687. Cited by: §J.3.
Appendix A Dataset
We utilize the datasets from the Who&When benchmark Zhang et al. (2025), which comprise both Algorithm-generated and Hand-crafted multi-agent system traces, totaling 184 executions. Specifically, 126 traces are generated algorithmically using multi-agent systems built on CaptainAgent Wu et al. (2023), while 58 traces are manually curated from Hand-crafted systems such as Magnetic-One Fourney et al. (2024). The dataset encompasses a wide range of realistic multi-agent scenarios based on queries from GAIA Mialon et al. (2023) and AssistantBench Yoran et al. (2024).
The dataset covers a broad set of multi-agent tasks, including web-navigation-style challenges derived from AssistantBench Yoran et al. (2024) and general-assistant reasoning tasks inspired by GAIA Mialon et al. (2023). Each trace is annotated with fine-grained failure information, including the responsible agent, the true decisive error, and a natural language explanation of the failure.
The two categories exhibit distinct characteristics in terms of agent participation and trace length. Algorithm-generated traces involve less than 4 agents, with trace lengths ranging from 5 to 10 steps, a mean length of 8.7 steps, and a median length of 10 steps. In contrast, Hand-crafted traces involve between 1 and 5 agents, with trace lengths spanning 5 to 130 steps, a mean length of 51.6 steps, and a median length of 32 steps. The distributions of both datasets, illustrated in Fig. 5, highlight the greater variability present in the Hand-crafted traces.
Moreover, some benchmarks, such as AgentErrorBench with AgentDebug Zhu et al. (2025) and TRAIL Deshpande et al. (2025), may appear superficially similar to our task setting. However they differ fundamentally from our problem formulation and evaluation objective, making direct comparison inappropriate. AgentErrorBench focuses on human-agent interaction and single-agent trajectories, where the objective is to detect the earliest mistake that initiates cascading reasoning distortions leading to failure. TRAIL is designed for identifying and analyzing all errors in trajectories. In contrast, this work focuses on failure attribution in LLM-based multi-agent systems, where the decisive error is defined as the earliest mistake whose correction could reverse the overall system failure. For example, an upstream deviation (an agent returning incomplete search results) eventually leads to a downstream decisive error (another agent fabricating missing information). AgentErrorBench would typically detect the earlier deviation, which could potentially be mitigated through subsequent agent interactions such as verification or additional tool usage, whereas this work aims to identify the interaction whose counterfactual-inspired correction would directly rescue the final outcome.
Appendix B Baselines
To evaluate the effectiveness of DCFA, we compare it against five representative baseline strategies that differ in how the LLM interacts with the failure traces: All-at-Once, Step-by-Step, and Binary-Search from the Who&When benchmark Zhang et al. (2025), as well as A2P West et al. (2025) and ECHO Banerjee et al. (2025). Each baseline reflects a distinct reasoning paradigm in localizing the failure-responsible step and agent within multi-agent system traces.
- •
All-at-Once. The algorithm leverages the LLM to interpret the entire system trace in a single pass and directly infer the decisive error of the system failure.
- •
Step-by-Step. The algorithm partitions the complete system trace into individual interactions. At each step, the model determines whether an error has occurred in the current interaction. If an error is detected, the process terminates immediately and the LLM outputs the corresponding step index. Otherwise, the reasoning continues sequentially until the final interaction is reached.
- •
Binary-Search. The algorithm begins with the query and the full failure log. It first determines whether the fault occurred in the upper or lower half of the trace. The identified half is then provided for further inspection. This iterative procedure continues until a single interaction is isolated as the decisive error.
- •
A2P. The algorithm employs a causal-inspired scaffold that sequentially guides LLMs through abductive hypothesis generation, explicit action-level intervention, and short-horizon counterfactual-inspired simulation to assess whether correcting a given step can reverse a system failure.
- •
ECHO. The algorithm introduces a hierarchical error attribution approach for multi-agent systems, explicitly targeting both minor execution deviations and overt errors within long interaction traces. Through multi-level contextual abstraction and consensus-based analysis, ECHO can detect early, subtle deviations that precede downstream failures, enabling fine-grained and reliable error localization.
All the baselines are based on the LLMs to achieve the failure attribution. To comprehensively assess each algorithm, we test them with various LLMs, including both open-source models of different sizes and commercial models. The models tested include the Qwen3 series (Qwen3-Coder-30B and Qwen3-235B), the DeepSeek series (DeepSeek-R1-32B and DeepSeek-R1-671B), GPT-5, and Gemini-2.5-pro. All baselines are evaluated under identical inference conditions to ensure a fair and consistent comparison.
Appendix C Implementation Details
Our approach operates in a training-free manner. In the GCA stage, minor deviation identification, causal-inspired dependency graph construction, and hypothesis generation are performed by the selected commercial LLM, following the same configuration as the baseline setup. In the LCE stage, the bidirectional dependency neighborhood search and decisive error decision are executed using the locally deployed Qwen-Coder-30B model. The encoder used to transform the outcomes generated by the simulation process into a continuous embedding space is implemented using the text-embedding-3-large API provided by OpenAI.
The judgment function is implemented as a prompt-based evaluator operating over multi-agent system traces. Rather than re-executing the entire system, receives a system trace containing agent identities and message contents during in the interactions, and is prompted to assess whether a target interaction would plausibly occur under a counterfactual-inspired intervention
Commercial LLMs are accessed through their respective APIs, while for local experiments, we deploy Qwen-Coder-30B on a workstation equipped with two NVIDIA A800 GPUs. For both commercial and locally deployed LLMs, the temperature is fixed to 0 to ensure deterministic outputs.
All inference procedures of the baseline methods follow identical preprocessing steps and prompt configurations on the Who&When benchmark to ensure fair and reproducible evaluation across different models.
To evaluate attribution quality, we focus on Step-level Accuracy, which measures the proportion of cases in which the exact erroneous step is correctly identified. Notably, we do not focus on Agent-level Accuracy, as in systems with a small number of agents, Agent-level Accuracy can be inflated. This occurs because it overlooks the more precise Step-level localization of failures. Given our emphasis on accurate localization at both the agent and step levels, we exclusively use Step-level Accuracy, rejecting the misleading Agent-level baseline.
Appendix D Robustness Across Independent Runs
Although we set the decoding temperature to 0 for all experiments, minor randomness in LLM inference may still lead to small fluctuations in the outputs. To evaluate the robustness of our results, we repeated the experiments using the three best-performing LLMs and report the mean and standard deviation over three independent runs.
As shown in Table 4, the variance across runs is extremely small for all methods and models. This indicates that the observed performance differences are stable and not caused by stochastic variations in the LLM outputs. In particular, DCFA consistently achieves the highest performance across all settings with minimal variance, demonstrating the robustness of the proposed failure attribution framework.
| Method | Algorithm-Generated | Hand-Crafted | ||||
|---|---|---|---|---|---|---|
| DeepSeek-R1-671B | GPT-5 | Gemini-2.5-pro | DeepSeek-R1-671B | GPT-5 | Gemini-2.5-pro | |
| All-at-Once | 16.93 0.00 | 16.14 0.00 | 27.51 0.02 | 4.02 0.01 | 3.45 0.00 | 4.60 0.01 |
| Step-by-Step | 27.25 0.03 | 28.31 0.03 | 34.92 0.03 | 11.50 0.01 | 13.22 0.01 | 16.10 0.01 |
| Binary-Search | 33.86 0.01 | 31.75 0.03 | 26.46 0.02 | 5.75 0.01 | 10.92 0.01 | 5.75 0.01 |
| A2P | 16.40 0.04 | 16.40 0.05 | 31.75 0.05 | 6.32 0.01 | 12.64 0.01 | 9.77 0.01 |
| ECHO | 40.21 0.00 | 31.48 0.03 | 31.22 0.00 | 19.54 0.09 | 13.79 0.01 | 13.22 0.08 |
| DCFA | 53.44 0.01 | 47.09 0.04 | 46.69 0.01 | 21.38 0.04 | 22.41 0.01 | 21.84 0.03 |
Appendix E Detailed Computational Cost Analysis
E.1 Computational Cost Calculation
This section reports detailed empirical statistics of token usage and runtime for DCFA and baseline methods on the Who&When benchmark. For DCFA, the average prompt overhead is approximately 400 tokens. The average trace length is about 3k tokens for the Algorithm-Generated dataset and 10k tokens for the Hand-Crafted dataset. The GCA stage performs two full-context API calls, each taking roughly 16 seconds, resulting in an average runtime of about 32 seconds per instance. The LCE stage is executed locally using Qwen3-Coder-30B on two NVIDIA A800 GPUs. Each inference processes approximately 6k tokens and takes about 60 seconds. Since bidirectional dependency neighborhood search requires on average 4.5 such calls, the total LCE runtime is approximately 270 seconds. Consequently, the overall runtime of DCFA is around 300 seconds per trajectory.
Baseline methods differ primarily in how they process the interaction trace. All-at-Once analyzes the entire trace using a single LLM call of roughly 6k tokens. Search-by-Step processes one interaction at a time (about 50 tokens per step) and averages around 20 steps per trace, resulting in approximately 9k tokens. Binary-Search iteratively analyzes halves of the remaining trace and typically requires about five calls, leading to roughly 12k tokens. A2P uses a prompt of about 300 tokens and processes the full trace in a single call, while ECHO also uses a prompt of about 300 tokens but performs two full-context calls.
E.2 Trade-off Analysis
Although DCFA introduces additional computational overhead, it yields substantial performance improvements. In terms of computational complexity, the strongest baseline, ECHO, and the GCA component of DCFA both require two full-context LLM calls. Notably, the GCA module alone already achieves clear improvements over ECHO (+10.58% on the Algorithm-Generated dataset and +0.57% on the Hand-Crafted dataset), suggesting that the primary performance gains stem from enhanced reasoning over the trajectory. The LCE module, in contrast, relies on local LLMs and therefore incurs relatively additional cost. It further improves detection accuracy on longer trajectories (e.g., +3.45% on the Hand-Crafted dataset), indicating a favorable trade-off between attribution accuracy and computational efficiency.
E.3 Runtime Cost and Practicality
DCFA is designed as a post-hoc failure attribution framework, consistent with baselines such as ECHO and A2P, rather than as a real-time debugging module. Its primary goal is to analyze the complete system trace after a task failure and identify the decisive errors. Insights from this analysis can guide offline modifications to the MAS, including agent structures, context management strategies, or tool invocation policies, rather than adjusting the system during execution.
For online or real-time scenarios, a different design is required. Incremental reasoning can be applied at each interaction step, leveraging cached causal-inspired dependency subgraphs and restricting counterfactual-inspired evaluations to local neighborhoods. This avoids repeatedly processing the full trace and can substantially reduce runtime overhead, making failure attribution more practical for long-running or continuously operating multi-agent systems.
Appendix F Supplementary Analysis of GCA-Induced Causal-inspired Dependency Graphs
We analyze the structural statistics of the causal-inspired dependency graphs constructed during the GCA stage, with a particular focus on cross-model variability and its implications for decisive error localization. To facilitate a concrete comparison, we further present representative causal-inspired dependency graphs generated by GPT-5 and Gemini-2.5-pro on a selected sample (49.json) from the Hand-crafted dataset.
F.1 Model Capacity and Deviations in Causal-inspired Dependency Graph Edge Density
Table 5 reports the average number of edges per graph across different LLMs on both the Algorithm-generated and Hand-crafted datasets.
| Model | Algorithm-generated | Hand-crafted |
|---|---|---|
| Qwen3-Coder-30B | 7.89 | 68.90 |
| DeepSeek-R1-32B | 8.24 | 46.39 |
| Qwen3-235B | 9.41 | 61.52 |
| DeepSeek-R1-671B | 9.25 | 57.09 |
| GPT-5 | 10.12 | 55.40 |
| Gemini-2.5-pro | 8.76 | 56.38 |
| Average | 8.95 | 56.61 |
A clear stratification emerges when comparing models of different capability. Less capable models, such as Qwen3-Coder-30B and DeepSeek-R1-32B, exhibit pronounced deviations from the overall mean, especially on the Hand-crafted dataset, producing overly dense or sparse causal-inspired dependency graphs (e.g., 68.90 edges for Qwen3-Coder-30B and 46.39 edges for DeepSeek-R1-32B). Such extreme deviations, in either direction, indicate unstable causal induction and low quality of the causal-inspired dependency graph generation, likely due to context degradation in LLMs.
By contrast, higher-capability models, such as GPT-5, Gemini-2.5-pro, and DeepSeek-R1-671B, maintain edge counts closer to the global average, suggesting a stronger ability to suppress irrelevant dependencies and preserve salient relations.
F.2 Implications for Decisive Error Localization.
To qualitatively assess how these structural differences affect downstream reasoning, Figure 6 visualizes the causal-inspired dependency graphs constructed by Gemini-2.5-Pro and GPT-5 on the same hand-crafted example, 49.json. Although both models successfully capture the overall dependency structure, their local structures in the critical failure region differ substantially.
Specifically, the GCA stage using Gemini-2.5-Pro initially hypothesizes as the decisive error, reflecting a more entangled dependency subgraph in the interval –. In contrast, GCA with GPT-5 constructs a more compact and hierarchical causal-inspired dependency structure over the same region, enabling it to identify as the decisive error.
Importantly, this discrepancy does not prevent DCFA with Gemini-2.5-Pro from ultimately recovering the correct failure source. Through the LCE module in the subsequent DCFA stage, the model performs counterfactual-inspired reasoning over the extracted local dependency subgraph, progressively evaluating and eliminating less relevant upstream candidates and converging on the decisive error . This observation highlights that while reliable global causal-inspired dependency graphs facilitate earlier localization, the proposed framework remains robust to structural noise through localized counterfactual-inspired refinement.
Appendix G Edge Growth and Long-Trace Scalability
Analysis of the causal-inspired dependency graphs constructed by DCFA shows that edge growth is nearly linear with trajectory length. In realistic multi-agent traces, each interaction typically links only to temporally or semantically adjacent steps, rather than to all preceding interactions. For example, in the Algorithm-generated dataset, traces average 8.7 steps with 8.95 edges, while in the Hand-crafted dataset, traces average 51.6 steps with 56.61 edges. This correspondence indicates that edge growth scales linearly rather than quadratically, mitigating combinatorial complexity.
Maintaining consistent structured reasoning over long traces remains challenging for LLM-based approaches. LCE addresses this challenge by performing localized exploration within the dependency graph, reducing the reasoning scope and keeping the additional search overhead roughly linear in trace length. Alternative cost-efficient strategies could further alleviate scaling issues, including heuristic graph pruning (e.g., limiting search depth or filtering low-confidence edges), sliding-window reasoning without full graph traversal, or hierarchical trace compression prior to graph construction. These approaches offer options for scaling DCFA to longer or more complex trajectories while retaining the benefits of localized dependency reasoning.
Appendix H Details for Case Study on Causal-inspired Dependency Graph Search
To demonstrate the effectiveness of DCFA, we present a case study based on a system trace from the Hand-crafted multi-agent system (ID: 49.json). In this scenario, the user query triggers interactions among three agents: Orchestrator, WebSurfer, and Assistant. The Orchestrator decomposes the task and coordinates the reasoning workflow, the WebSurfer retrieves external information, and the Assistant integrates the gathered evidence to produce the final answer.
| Stage | Identified Decisive Error | Agent | Supporting Evidence |
|---|---|---|---|
| GCA (Global) | Orchestrator | Ignore missing operator details | |
| LCE (Local) | Assistant | Fabricate missing content |
The system trace captures each step of the multi-agent reasoning process, including information collection, context propagation, and intermediate reasoning actions. Specifically, the detail content of the user query is shown in the following box:
In response, the multi-agent system generates the system trace consisting of fifteen interactions . Each represents an interaction from one agent, as illustrated in Fig. 7. The system produces an incorrect answer “k”, whereas the correct missing character should be backtick “`”.
Within this trace, some steps involve minor deviations, such as incomplete retrieval or partial context propagation, which are generally correctable through downstream verification. In contrast, a decisive error occurs when the Assistant fabricates information based on incomplete or ambiguous input, producing content that is irreversible and misguides subsequent reasoning, ultimately leading to an incorrect outcome.
This overview sets the stage for a detailed case study analysis, where we examine how DCFA distinguishes contributory minor deviations from the decisive fabrication error and demonstrates its capability to pinpoint the root cause of failure in multi-agent reasoning.
Inference Phase of GCA
During global analysis, GCA first detects a potential minor deviation at , where the WebSurfer attempts to retrieve definitions of Unlambda operators but returns incomplete information.
Based on this, GCA builds a causal-inspired dependency graph centered on (Fig. 4) and ultimately identifies and attributes the system failure to , where the Orchestrator mistakenly assumes that the retrieved information is complete and proceeds without verification.
Refinement Phase of LCE.
To refine this hypothesis of GCA, the Local Counterfactual-inspired Enhancement (LCE) module focuses on the neighborhood around , forming a local subgraph (Fig. 4). It conducts counterfactual-inspired simulations to evaluate whether modifying prior steps would correct the outcome. Forward traversal from to shows no effect, while backward traversal identifies as the decisive cause.
LCE identifies as the decisive error, where the Assistant fabricates knowledge by falsely claiming that the WebSurfer’s incomplete results included operator definitions. This misinformation misleads subsequent reasoning, ultimately producing the wrong output in .
Decisive Error and DCFA Attribution.
Table 6 compares the decisive error attributions identified by the GCA and LCE modules. In this trace, corresponds to the WebSurfer attempting to retrieve essential information about Unlambda operators. Although the retrieved information is incomplete, this step constitutes a minor deviation: it is benign and can be effectively corrected by the system’s subsequent verification mechanisms, such as detecting missing fields and re-invoking tools to retrieve the omitted information. Therefore, does not constitute a decisive error.
represents the Orchestrator proceeding and propagating the partial context returned by WebSurfer. While the GCA module flags as a global-level misjudgment, this step primarily propagates the minor deviation from rather than introducing a fundamentally new error, making it contributory but not causative.
The decisive error occurs at , where the Assistant integrates the retrieved information and fabricates operator details, transforming a recoverable incompleteness into an irreversible factual error. This fabricated knowledge directly misguides subsequent reasoning and drives the final incorrect outcome. Unlike , the error at cannot be corrected through downstream verification, highlighting its role as the decisive upstream failure.
The DCFA framework identifies this decisive error through a two-stage reasoning process. First, GCA constructs a step-level causal-inspired dependency graph centered on minor deviations, capturing high-level dependencies and propagating influence to generate an initial hypothesis, here pointing to . Second, LCE performs a local neighborhood analysis around the hypothesis, employing counterfactual-inspired simulations to quantify the impact of each step on the final outcome. By evaluating the potential corrections in context, LCE isolates as the only step whose modification could prevent the error, thereby pinpointing the true decisive error.
This dual-view strategy, combining global reasoning over a causal-inspired dependency structure with local counterfactual-inspired validation, allows DCFA to transition from broad dependency mapping to precise error localization, effectively distinguishing contributory minor deviations from decisive errors. The approach not only enhances the interpretability of multi-agent reasoning traces but also identifies critical interactions whose correction is most likely to improve the final outcome, demonstrating clear advantages over purely global or local attribution methods.
Appendix I Bias Reduction
To further validate the robustness of our causal-inspired attribution, we examine DCFA’s capability to mitigate bias in decisive error localization compared with baselines. Here, bias refers to the systematic deviation of predicted steps from the ground-truth decisive error, which can obscure the actual failure source and mislead subsequent debugging Li et al. (2022); Shrotriya et al. (2025).
We compare DCFA with the representative baseline, ECHO, on both Algorithm-generated and Hand-crafted traces from the Who&When benchmark Zhang et al. (2025). For each trace, we quantify the deviation between the predicted and true causal-step indices:
| (17) |
where and are predicted indices, and denotes the ground-truth step. A smaller deviation implies lower bias and more reliable reasoning.
Figure 8 visualizes the deviation distributions for DCFA and ECHO. Across both DeepSeek-R1-671B and Gemini-2.5-pro backends, DCFA exhibits a notably sharper and more centralized distribution around zero. Overall, DCFA exhibits a smaller failure attribution offset than ECHO, indicating more accurate and less biased localization of the true decisive errors. This indicates that DCFA predictions consistently align closer to the ground-truth decisive error. Such improvement stems from the both global and local view causal-inspired enhancement of the DCFA, preventing error accumulation in step-level reasoning. Overall, the results demonstrate that DCFA reduces attribution bias, leading to more stable and interpretable failure diagnostics across synthetic and real-world traces.
Appendix J Related Works
J.1 LLM-based Multi-Agent Systems and System Failure
The recent progress of Large Language Model (LLM)-based agents has accelerated the development of Multi-Agent Systems (MAS), which have emerged as a promising paradigm for addressing complex, multi-step tasks that exceed the capabilities and token limitations of single, monolithic LLMs Yang et al. (2026); Sun et al. (2025); Pei et al. (2025); Xia et al. (2025). The principles underlying MAS, such as task decomposition, inter-agent communication, and emergent collective reasoning. Recent surveys have proposed systematic taxonomies of agent systems architectures Wu et al. (2023); Fourney et al. (2024), identifying core functional components required for effective coordination and execution. These components generally include explicit role assignment, hierarchical or iterative planning mechanisms, structured communication protocols, externalized memory modules, and evaluation and feedback mechanisms Yan et al. (2025).
Despite rapid progress, current agent systems remain limited in robustness, interpretability, and stability Zhang et al. (2025); Cemri et al. (2026). Agents often exhibit brittle coordination, unstable reasoning trajectories, and cascading errors, particularly in long-horizon or highly interactive environments Zhang et al. (2025); Cemri et al. (2026). Benchmarking efforts such as AgentBench Liu et al. (2024b) empirically demonstrate these weaknesses, showing that even state-of-the-art systems struggle to maintain coherence and consistency in collaborative, multi-step tasks. The recurrence of these failures exposes a critical research gap: understanding why and how LLM-based agent systems deviate from their intended execution pathways.
J.2 Failure Attribution in LLM-based Multi-agent Systems
Prior work has studied error analysis in agentic systems. AgentErrorBench with AgentDebug Zhu et al. (2025) locates mistakes that trigger cascading reasoning distortions in human-agent interactions in single-agent settings, whereas TRAIL Deshpande et al. (2025) analyzes errors in execution trajectories in multi-agent settings. These definitions, however, do not directly target the most actionable points for repairing task failures. Meanwhile, unlike single-agent systems with sequential execution and full information access, multi-agent systems involve heterogeneous roles, complex interaction topologies, partial observability, and non-trivial communication, producing intricate and interdependent dependency chains Cemri et al. (2026); Fourney et al. (2024); Wu et al. (2023). Thus, the focus shifts to identifying the earliest mistake whose correction could reverse overall system failure in multi-agent systems. The mistake, termed the decisive error Cemri et al. (2026); Fourney et al. (2024); Wu et al. (2023), represents the key point for restoring task success.
Recent research has begun to explore failure attribution in multi-agent systems using LLMs as diagnostic detectors Zhang et al. (2025); Zhang et al. (2026). Existing methods can be categorized into fine-tuning-based Zhang et al. (2026) and instruction-based Zhang et al. (2025); Cemri et al. (2026) paradigms. The fine-tuning-based approach, exemplified by AgenTracer Zhang et al. (2026), constructs specialized datasets to train dedicated models for failure diagnosis. While this approach offers strong controllability, it also incurs substantial costs for annotation, training, and maintenance. Instruction-based methods Zhang et al. (2025); Cemri et al. (2026), in contrast, rely on carefully designed prompts to guide LLMs in analyzing multi-agent system traces and identifying the decisive errors. For example, the Step-by-Step algorithm in the Who&When benchmark Zhang et al. (2025) formulates failure attribution as a sequential inspection process that progressively checks each interaction. Another algorithm in the benchmark, Binary-Search, adopts a divide-and-conquer strategy, recursively narrowing the trace scope until the decisive error is pinpointed. Additionally, Banerjee et al. Banerjee et al. (2025) propose hierarchical context representations and multi-perspective consensus to improve attribution reliability, while Weng et al. West et al. (2025) introduce A2P, a causal-guided attribution framework powered by the DeepScientist AI system.
Despite these advances, current methods West et al. (2025); Banerjee et al. (2025) rely primarily on the vanilla reasoning capabilities of LLMs, which often restricts their focus to superficial deviations and leads to performance degradation when processing long system traces. These limitations frequently result in inaccurate or incomplete attribution. To address these challenges, we propose a unified framework that incorporates structured reasoning and extracts local dependency chains to enable robust and reliable failure attribution in LLM-based multi-agent systems.
J.3 Causal Reasoning and Counterfactual Analysis via LLM
Recent advances have increasingly explored the use of large language models (LLMs) for causal reasoning and counterfactual-inspired analysis, including discovering, representing, and simulating causal relationships Su et al. (2025); Cheng et al. (2025); Luo et al. (2024). This line of research leverages the linguistic and reasoning capabilities of LLMs to extract and reason about causal structures from unstructured text. These studies show that LLMs can effectively extract events and propose candidate causal links with high recall, serving as powerful front-end modules for event causality analysis Liu et al. (2024a); Wang et al. (2024).
To improve reliability, several works Tong et al. (2024); Cohrs et al. (2025); Ban et al. (2025); Liu et al. (2025) integrate LLM-derived evidence with algorithmic estimators, aggregating multiple causal hypotheses or graph structures and taking their intersection to enhance robustness and filter out spurious relations. This hybrid approach demonstrates that LLMs can serve as high-level causal reasoning modules, while traditional estimators provide quantitative validation. Further research explores the use of LLMs as informative priors for causal graph discovery. Recent studies demonstrate that LLMs encode rich domain and commonsense knowledge that can guide structure learning and improve interpretability Darvariu et al. (2024); Jiralerspong et al. (). By injecting LLM-derived constraints or edge probabilities into graph search algorithms, these methods reduce sample complexity and produce more plausible causal graphs for downstream causal analysis.
Beyond causal structure discovery, causal inference and counterfactual simulation have been explored as mechanisms for testing causal hypotheses Liu et al. (2025); Zhu et al. (2024); Verma et al. (2025). Causal inference provides formal tools for reasoning about the effects of interventions Imai and Nakamura (2026), while counterfactual simulation assesses how altering an event or reasoning step might change subsequent outcomes. Recent works apply such techniques to LLM reasoning chains, using generated interventions to evaluate faithfulness and sensitivity to hypothetical changes in chain-of-thought reasoning Tutek et al. (2025); Yu et al. (2025). These findings suggest that LLMs can both represent causal structures and support the simulation of hypothetical interventions for evaluating outcome changes.
Inspired by these developments, our work adapts these ideas to failure attribution in LLM-based multi-agent systems. Rather than claiming formal causal inference, DCFA employs causal-inspired structured reasoning over system traces. Specifically, we construct a causal-inspired dependency graph to organize inter-agent dependencies and use counterfactual-inspired evaluation with LLM-mediated approximate interventions to assess how correcting candidate interactions may alter the final outcome. This dual-view design provides an interpretable basis for identifying the interaction that most plausibly constitutes the decisive error in long and unstructured multi-agent traces.
Appendix K Pseudocode Description of DCFA
This section briefly explains the pseudocode of the two modules in the DCFA framework.
Global Causal-inspired Attribution (GCA).
Algorithm 1 summarizes the overall global attribution procedure. The module first identifies candidate minor deviations in the system trace using the deviation identification process , which extracts interactions that exhibit abnormal behaviors together with brief textual explanations. Based on these candidates, the causal-inspired dependency graph construction procedure organizes the trace into a directed causal-inspired dependency graph that encodes plausible dependencies between interactions. Finally, the hypothesis generation process analyzes the trace, the detected deviations, and the constructed causal-inspired dependency graph to produce an initial hypothesis for the decisive error.
Local Counterfactual-inspired Enhancement (LCE).
Algorithm 2 refines the global hypothesis by performing a bidirectional greedy search on the causal-inspired dependency graph. The procedure first constructs a corrected trace using to provide a reference trajectory aligned with the ground-truth outcome. Starting from the hypothesis interaction , the algorithm iteratively explores downstream and upstream neighbors in two search loops. At each step, candidate interactions are evaluated using counterfactual-inspired interventions, and the interaction with the largest positive marginal improvement becomes the next pivot. The union of visited interactions forms a compact dependency neighborhood, which is then evaluated by the decision function to determine the final decisive error.
Appendix L Prompts for DCFA
This section presents the prompts used to implement the DCFA components described in Section 4. The prompts correspond to key reasoning steps in the Global Causal-inspired Attribution (GCA) and Local Counterfactual-inspired Enhancement (LCE) modules. For clarity and reproducibility, we provide representative prompt templates. Actual inputs (e.g., system traces, candidate events, and ground-truth answers) are dynamically inserted during execution.
Prompts for the GCA Module.
The GCA module performs global structured reasoning over the full system trace to generate an initial hypothesis for the decisive error.
First, the model identifies candidate minor deviations and constructs a causal-inspired dependency graph over the system trace. The prompt instructs the model to analyze the interaction events extracted from the system trace, detect potential mistake events, and identify dependency relationships among events based on the criteria of temporality, necessity, and sufficiency defined in Section 4.1. This step jointly performs minor deviation identification and causal-inspired dependency graph construction, producing both the set of candidate deviation events and the directed dependency relationships among them. The prompt used for this step is shown in Figure 9.
After the candidate deviations and causal-inspired dependency graph are obtained, the GCA module performs global reasoning to generate an initial hypothesis of the decisive error. Given the system trace, the set of candidate minor deviations, the explanations for each deviation, and the constructed causal-inspired dependency graph, the model evaluates how deviations propagate through the interaction chain and selects the event that most plausibly explains the final failure outcome. The corresponding prompt is illustrated in Figure 10.
Prompts for the LCE Module.
The LCE refines the hypothesis generated by GCA through localized counterfactual-inspired reasoning, comprising prompt-based steps for error selection.
First, in the counterfactual-inspired evaluation process, the model generates a corrected system trace. Given the original trace, the identified mistake events, and the ground-truth outcome, the prompt instructs the model to correct the erroneous events and causally affected downstream interactions while preserving the original event structure and formatting. This produces a fully corrected trace representing a counterfactual-inspired scenario in which the mistakes are fixed. The prompt is shown in Figure 11.
Next, for each counterfactual-inspired trace constructed during evaluation, the model simulates the corresponding system outcome based on the partially corrected interaction sequence. Starting from the corrected prefix and the remaining original interactions, the model reconstructs the downstream reasoning process and predicts the resulting final outcome. This prompt enables the model to estimate how correcting specific interactions affects the final result. The corresponding prompt is presented in Figure 12.
Finally, after the bidirectional dependency neighborhood search and counterfactual-inspired evaluation produce a set of candidate decisive error events, the model performs a final semantic decision and selection step. Given the candidate events, the causal-inspired dependency graph structure, and the ground-truth outcome, the model determines which interaction most plausibly constitutes the decisive error responsible for the observed system failure. The prompt is shown in Figure 13.