by
Dependency-Aware Chain-of-Thought Compression for Financial Reasoning
Abstract.
Chain of thought prompting improves complex reasoning, but its long intermediate traces create substantial inference cost and hinder practical deployment in financial settings. We present a Hierarchical Semantic Distillation Network, HSDN, for compressing reasoning chains while preserving answer accuracy and logical coherence. The framework combines semantic segmentation, dependency graph construction, dual encoder importance scoring, constrained segment selection, and local boundary rewriting. A frozen Qwen3 4B model is used only for feature extraction and final answer generation, while the compression process remains structured and interpretable. On the AFAC2025 benchmark, HSDN achieves 91.0% accuracy with 68.4% compression, outperforming strong compression baselines in overall score and reasoning coherence. The results show that graph guided compression is effective for high stakes financial reasoning tasks.
1. Introduction
Large language models have shown strong performance on multi step reasoning tasks, and chain of thought prompting has become a standard mechanism for improving intermediate deliberation quality (Wei et al., 2022). This capability is especially valuable in financial applications, where models must integrate textual evidence, numerical calculation, and compliance oriented interpretation under strict correctness requirements. At the same time, longer reasoning traces increase latency, memory usage, and serving cost, which limits their practicality in real world systems.Recent roofline-guided co-optimization work on Arm CPUs highlights how memory bandwidth and low-precision throughput constraints can be addressed with mixed-precision kernels and fused attention, which is relevant to our efficiency-oriented design (Zhou, 2026b).This efficiency challenge is closely related to serverless AI inference, where systems must balance cold start latency against the cost of retaining idle GPU resources, motivating adaptive lifecycle management strategies such as AdaScale (Zhou, 2026a). Zero shot reasoning studies further indicate that reasoning quality depends not only on model scale, but also on how intermediate steps are organized and exposed during inference (Kojima et al., 2022).Zero shot reasoning studies further indicate that reasoning quality depends not only on model scale, but also on how intermediate steps are organized and exposed during inference (Kojima et al., 2022). Recent multi-agent troubleshooting frameworks such as PRISM further suggest that specialized role decomposition and evidence verification can improve complex reasoning workflows (Yan et al., 2026). Existing approaches still face an important gap for high stakes financial reasoning. Direct generation is efficient but often omits key logical steps, while full chain of thought preserves detail at the expense of excessive verbosity. More advanced deliberate reasoning strategies improve search over reasoning paths, yet they do not explicitly address how to compress a completed chain while retaining dependency structure and numerical faithfulness (Yao et al., 2023). In financial scenarios, this omission is critical because dropping a seemingly minor step can invalidate later computations or break the justification trail needed for auditability. To address this problem, we propose a hierarchical semantic distillation framework for compressing reasoning chains in a structured manner. Our method first segments reasoning into semantic units, then builds a directed dependency graph, scores segment importance with question aware representations, and performs globally optimal selection under length and dependency constraints. A lightweight boundary rewriter restores fluency after segment removal, while a frozen large language model is used only for semantic features and final answer generation. This design yields an interpretable compression pipeline that reduces reasoning length without sacrificing the logical continuity required in financial decision support.
2. Related Work
Research on handling long inputs has largely focused on global mechanisms that reduce the cost of full sequence processing. Sparse attention architectures such as Longformer and BigBird improve scalability for long documents by restricting or restructuring attention patterns, offering a strong foundation for efficient long context modeling (Beltagy et al., 2020; Zaheer et al., 2020). However, these methods mainly optimize representation efficiency at the token level and do not directly determine which reasoning steps should be retained for answer faithful chain compression. A second line of work studies fine grained and dynamic reduction, where models selectively remove less important tokens during inference. TR BERT introduces dynamic token reduction conditioned on task relevance, and PoWER BERT progressively eliminates low impact word vectors to accelerate inference while maintaining prediction quality (Ye et al., 2021; Goyal et al., 2020). These methods demonstrate the value of adaptive compression, but they typically operate on local token salience and do not model explicit logical dependencies among reasoning segments, which are crucial in multi step financial inference.Recent work on dynamic retrieval-augmented generation further shows that selective tool use and sufficiency-aware routing can improve robustness when static context is insufficient (Liang et al., 2026). Our approach is also related to work on structured summarization and faithful generation.Recent hybrid deep learning work on supply chain delay prediction further underscores the value of jointly modeling temporal dynamics and graph structure in operational decision systems (Xue et al., 2026a).A recent risk-aware dynamic routing framework demonstrates how spatiotemporal graph neural networks can support resilient decision-making under congestion and fluctuating demand in large-scale logistics systems (Xue et al., 2026b).In particular, recent hybrid architectures that combine pyramid-style semantic encoding with graph attention and language-model-assisted rationale distillation further suggest the value of multi-granularity feature extraction for complex threat narratives (Xu, 2026b). HeterSumGraph shows that graph representations can capture document level relations for extractive summarization, while PRIMERA improves long document summarization through pyramid based pretraining (Wang et al., 2020; Xiao et al., 2022). For generation faithfulness, FactPEGASUS highlights the importance of preserving factual consistency during rewriting and compression (Wan and Bansal, 2022).Complementary to these generation-focused approaches, pairwise verification methods such as SENTINEL can detect subtle semantic inconsistencies between two candidate outputs for the same source, offering a useful perspective for assessing whether compression introduces manipulative distortions (Xu, 2026a).
3. Methodology
Chain-of-thought reasoning is essential for complex financial inference, yet verbose reasoning sequences impose substantial computational overhead. This paper presents a hierarchical semantic distillation framework that compresses lengthy reasoning chains while preserving logical integrity, addressing challenges unique to financial applications such as tabular data interpretation, multi-step numerical computation, and regulatory compliance verification. The framework introduces a structured pipeline comprising four key components: a graph-based dependency parser that constructs explicit reasoning topology to capture causal relationships and identify dispensable content; a dual-encoder architecture that scores segment importance through cross-modal attention between question semantics and reasoning content; a dynamic programming solver that guarantees globally optimal segment selection under length budgets while respecting dependency constraints; and a lightweight sequence-to-sequence rewriter that ensures coherence at segment boundaries without introducing factual inconsistencies. A frozen large language model serves solely as a semantic feature extractor and answer generator, keeping the core compression logic in interpretable algorithmic components. Evaluation on financial reasoning benchmarks demonstrates substantial length reduction while maintaining answer accuracy, with the graph-based formulation providing transparent compression rationale for human verification in high-stakes applications. The overall architecture is illustrated in Fig. 1.
4. Algorithm and Model
4.1. Semantic Segmentation
The first stage partitions the continuous reasoning text into discrete semantic units that serve as atomic elements for subsequent processing. Unlike sentence-level segmentation which often breaks logical units inappropriately, we train a specialized boundary detector that respects reasoning step boundaries.
4.1.1. Boundary Detection Network
We employ a bidirectional LSTM network augmented with conditional random fields to identify segment boundaries. Given the token sequence , the network first computes contextual representations:
| (1) |
| (2) |
The concatenated representation captures both preceding and following context, which proves essential for detecting boundaries that depend on what comes next. Early experiments with unidirectional models showed poor performance on boundaries preceding numerical computations, where the boundary significance only becomes apparent from the subsequent calculation.
The boundary probability at each position is computed through a feedforward layer:
| (3) |
where represents emission scores for boundary and non-boundary labels.
To ensure globally consistent segmentation, we apply a linear-chain CRF layer that models label transitions:
| (4) |
where is the transition matrix and is the partition function computed via the forward algorithm. The CRF layer prevents degenerate solutions such as consecutive boundaries or excessively long segments, which frequently occurred with independent classification.
4.1.2. Financial Domain Adaptations
Financial reasoning text presents unique segmentation challenges that required specific adaptations. Numerical expressions spanning multiple tokens, such as percentage changes or currency amounts, must remain intact within segments. We address this by incorporating a numerical span detector that identifies contiguous numerical expressions:
| (5) |
During CRF decoding, we modify transition scores to prohibit boundaries within detected numerical spans by setting
when and the span continues from position .
Additionally, financial text frequently contains table references and structured data mentions that should not be split. We found that simply expanding the context window was insufficient; instead, we pretrain the boundary detector on a auxiliary task of table cell boundary detection, which transfers effectively to reasoning text segmentation.
4.2. Dependency Graph Construction
With the reasoning chain segmented into units , we construct a directed acyclic graph that explicitly encodes logical dependencies between segments. This graph serves as the structural backbone for compression decisions, ensuring that removing a segment does not orphan its dependents. Fig. 2 illustrates the graph construction pipeline.
4.2.1. Node Representation
Each segment becomes a node in the graph. We compute node embeddings by mean-pooling token representations from a frozen language model encoder:
| (6) |
To capture segment-level semantics beyond token averaging, we apply a learned projection with residual connection:
| (7) |
This additional projection layer allows the model to learn task-specific representations while leveraging the pretrained encoder’s semantic knowledge. We experimented with fine-tuning the encoder but found that freezing it and adding the projection layer yielded comparable performance with substantially reduced memory requirements during training.
4.2.2. Edge Prediction
Edges represent logical dependencies where the source segment provides information required by the target segment. We predict edges using a biaffine attention mechanism that has proven effective for dependency parsing:
| (8) |
where and are head and dependent representations computed through separate multilayer perceptrons.
The edge probability is obtained via sigmoid activation:
| (9) |
A critical challenge we encountered was handling long-range dependencies that span many intermediate segments. The biaffine scorer tends to underestimate such dependencies due to the difficulty of directly relating distant segments. Our solution introduces a path-augmented scoring term:
| (10) |
where captures transitive dependency strength.
4.2.3. Acyclicity Constraint
The dependency graph must be acyclic to represent valid logical flow. We enforce this through a differentiable acyclicity regularization term based on the matrix exponential characterization:
| (11) |
where is the adjacency matrix with , and denotes element-wise multiplication. This regularizer equals zero if and only if the expected graph is acyclic.
During inference, we apply a greedy pruning procedure that removes the lowest-probability edge from any detected cycle, iterating until the graph is acyclic. In practice, the regularization during training ensures cycles are rare, typically requiring fewer than two pruning iterations.
4.3. Importance Scoring and Selection
Given the dependency graph, we must select which segments to retain in the compressed output. This stage comprises two components: a neural importance scorer that evaluates each segment’s contribution to answering the question, and a dynamic programming algorithm that finds the optimal selection respecting dependency constraints. The scoring and selection mechanism is depicted in Fig. 3.
4.3.1. Dual-Encoder Importance Scorer
The importance scorer employs a dual-encoder architecture that separately encodes the question and each reasoning segment, then computes relevance through cross-attention. This design allows efficient scoring of all segments with a single question encoding pass.
The question encoder applies self-attention layers to produce a contextualized representation:
| (12) |
where is the question length and is the number of layers.
For each segment, we compute cross-attention scores against the question:
| (13) |
| (14) |
The attended question representation captures which aspects of the question segment addresses. The importance score combines this relevance signal with segment-intrinsic features:
| (15) |
where is a feature vector containing segment length, position, numerical content indicators, and graph centrality measures.
An important trick we discovered is incorporating the segment’s graph centrality into the importance score. Segments with high out-degree, meaning many other segments depend on them, tend to contain foundational information that should be preserved. We compute PageRank scores on the dependency graph and include them in :
| (16) |
where is the damping factor.
4.3.2. Constrained Selection via Dynamic Programming
Selecting the optimal subset of segments is formulated as a constrained optimization problem:
| (17) |
subject to:
| (18) |
| (19) |
The first constraint enforces the length budget, while the second ensures that if a segment is selected, all its dependencies are also selected. This problem generalizes the knapsack problem with precedence constraints.
We solve it via dynamic programming on the topologically sorted graph. Let denote the maximum importance achievable considering segments with total length exactly . The recurrence is:
| (20) |
where the second case requires and .
The time complexity is , which is efficient for typical reasoning chain lengths. We implement the dependency check through bit manipulation, maintaining a bitmask of selected segments and verifying parent inclusion in constant time after preprocessing parent masks.
4.4. Boundary Rewriting
Directly concatenating retained segments often produces incoherent text with abrupt transitions. The rewriting module performs local edits at segment boundaries to restore fluency while preserving factual content.
4.4.1. Boundary Context Extraction
For each pair of consecutively selected segments where intermediate segments were removed, we extract boundary context:
| (21) |
where Suffix and Prefix extract tokens from segment ends and beginnings respectively.
4.4.2. Seq2Seq Rewriter
A lightweight transformer-based sequence-to-sequence model takes boundary context as input and generates a smoothed transition:
| (22) |
To prevent hallucination, the decoder employs a copy mechanism that biases generation toward input tokens:
| (23) |
where is a learned gate and are copy attention weights. The rewriter is trained on synthetic boundary pairs created by randomly removing segments from correct reasoning chains, avoiding error propagation from the full pipeline.
4.4.3. Consistency Verification
We verify factual consistency by computing embedding similarity between the original boundary region and the rewritten version:
| (24) |
If similarity falls below threshold , we fall back to simple concatenation with a generic connective phrase.
4.5. Training Procedure
The framework is trained in three stages to ensure stable optimization and effective knowledge transfer.
4.5.1. Stage 1: Component Pretraining
The segmentation network is pretrained on manually annotated reasoning chains using CRF negative log-likelihood:
| (25) |
The dependency graph predictor is pretrained on synthetic dependency data generated by prompting a large language model to annotate reasoning chain dependencies.
4.5.2. Stage 2: Joint Scorer Training
The importance scorer and graph predictor are trained jointly using a multi-task objective:
| (26) |
The scoring loss uses ground-truth labels derived from oracle compression identifying the minimal segment subset producing correct answers:
| (27) |
| (28) |
4.5.3. Stage 3: End-to-End Refinement
The complete pipeline is fine-tuned using reinforcement learning with a reward combining accuracy and compression:
| (29) |
We employ REINFORCE with baseline subtraction to reduce gradient variance:
| (30) |
where is an exponential moving average of recent rewards.
4.6. LLM Integration
The large language model component serves two specific roles in our framework, deliberately limited to leverage its strengths while avoiding the interpretability and efficiency drawbacks of end-to-end neural compression.
4.6.1. Feature Extraction
We use a frozen Qwen3-4B model as a semantic feature extractor, accessing its intermediate layer representations to initialize segment embeddings. Specifically, we extract features from layer 16 of 32, which empirically balances semantic abstraction with surface-level detail:
| (31) |
The frozen encoder provides rich pretrained representations without the computational cost of fine-tuning or the risk of catastrophic forgetting.
4.6.2. Answer Generation
After compression, the retained segments are concatenated with minimal rewriting and fed to the language model for final answer generation:
| (32) |
We apply standard decoding with temperature and nucleus sampling with to generate diverse answer candidates for the best-of-5 evaluation protocol.
4.7. Complexity Analysis
The computational complexity of each component scales tractably with input size. Segmentation requires time for the BiLSTM pass and for CRF decoding. Graph construction involves edge predictions where is the number of segments. The dynamic programming selection runs in time. Boundary rewriting processes at most boundaries with constant-length inputs each.
The overall complexity is dominated by the LLM feature extraction and answer generation, which are unavoidable for the task. Our framework adds minimal overhead to these fixed costs while providing structured, interpretable compression that pure neural approaches cannot match.
5. Evaluation Metrics
We adopt four metrics following the challenge protocol. Best-of-5 accuracy measures correctness when any of five samples matches the reference:
| (33) |
Compression ratio quantifies length reduction:
| (34) |
The competition score penalizes incorrect answers with maximum length :
| (35) |
Reasoning coherence score evaluates logical dependency preservation:
| (36) |
6. Experiment Results
Table 1 presents the main comparison and ablation study on the AFAC2025 benchmark. And the changes in model training indicators are shown in Fig4
.
| Method | Acc.(%) | CR(%) | RCS | Score |
|---|---|---|---|---|
| Qwen3-4B (Full CoT) | 94.0 | 0.0 | 1.000 | -184700 |
| Qwen3-4B (Direct) | 71.0 | 95.2 | – | -97890 |
| LLMLingua-2 | 85.0 | 58.3 | 0.724 | -83650 |
| CompAct | 86.0 | 61.7 | 0.756 | -78420 |
| RECOMP | 88.0 | 54.2 | 0.812 | -89760 |
| LongLLMLingua | 87.0 | 63.5 | 0.743 | -74380 |
| HSDN (Ours) | 91.0 | 68.4 | 0.867 | -61250 |
| w/o Dependency Graph | 87.0 | 71.2 | 0.712 | -72340 |
| w/o Boundary Rewriter | 90.0 | 68.4 | 0.791 | -63120 |
| w/o PageRank Features | 89.0 | 67.8 | 0.834 | -65780 |
| w/o RL Fine-tuning | 89.0 | 65.3 | 0.851 | -68450 |
As shown in Table 1, HSDN achieves the best trade-off between accuracy and compression. The dependency graph contributes most significantly, with its removal causing 4% accuracy drop. Compared to perplexity-based methods like LLMLingua-2, our graph-guided approach better preserves reasoning coherence.
7. Conclusion
In this work, we presented FinStack-Net, a hierarchical ensemble framework combining LightGBM, CatBoost, and a deep neural network with residual and attention mechanisms for fraud and gambling account detection. Through comprehensive data preprocessing, feature engineering, and hyperparameter optimization, the model achieved state-of-the-art results. Ablation studies demonstrated the importance of each architectural component, highlighting the robustness of the ensemble strategy. Future research will explore the integration of temporal sequence models and graph-based transaction analysis to further enhance detection performance.
References
- Longformer: the long-document transformer. arXiv preprint arXiv:2004.05150. Cited by: §2.
- Power-bert: accelerating bert inference via progressive word-vector elimination. In International conference on machine learning, pp. 3690–3699. Cited by: §2.
- Large language models are zero-shot reasoners. Advances in neural information processing systems 35, pp. 22199–22213. Cited by: §1.
- DynaRAG: bridging static and dynamic knowledge in retrieval-augmented generation. In 2026 9th International Symposium on Big Data and Applied Statistics (ISBDAS), pp. 442–445. Cited by: §2.
- FactPEGASUS: factuality-aware pre-training and fine-tuning for abstractive summarization. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 1010–1028. Cited by: §2.
- Heterogeneous graph neural networks for extractive document summarization. In Proceedings of the 58th annual meeting of the association for computational linguistics, pp. 6209–6219. Cited by: §2.
- Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems 35, pp. 24824–24837. Cited by: §1.
- PRIMERA: pyramid-based masked sentence pre-training for multi-document summarization. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 5245–5263. Cited by: §2.
- Detecting manipulated llm outputs via siamese transformers and semantic dependency graphs. In 2026 2nd International Conference on Artificial Intelligence and Computational Intelligence (AICI 2026), New York, NY, USA. External Links: ISBN 979-8-4007-2279-0/2026/02, Document Cited by: §2.
- Pyramid convolution and bidirectional graph attention for cyber threat detection from unstructured text. In Proceedings of the 2026 International Conference on Artificial Intelligence and Control, pp. 561–567. Cited by: §2.
- EAGLE: edge-aware graph learning for proactive delivery delay prediction in smart logistics networks. arXiv preprint arXiv:2604.05254. Cited by: §2.
- Resilient routing: risk-aware dynamic routing in smart logistics via spatiotemporal graph learning. arXiv preprint arXiv:2601.13632. Cited by: §2.
- PRISM: pipeline for root-cause investigation via specialized multi-agents. In 2026 International Conference on Generative Artificial Intelligence and Information Security (GAIIS), pp. 709–712. Cited by: §1.
- Tree of thoughts: deliberate problem solving with large language models. Advances in neural information processing systems 36, pp. 11809–11822. Cited by: §1.
- Tr-bert: dynamic token reduction for accelerating bert inference. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 5798–5809. Cited by: §2.
- Big bird: transformers for longer sequences. Advances in neural information processing systems 33, pp. 17283–17297. Cited by: §2.
- AdaScale: predictive and utility-aware autoscaling for serverless ai inference. In 2026 6th International Conference on Artificial Intelligence and Industrial Technology Applications (AIITA), pp. 544–550. Cited by: §1.
- Roofline-guided mixed quantization and kernel co-optimization for efficient large language model inference on arm cpus. In 2026 3rd International Conference on Digital Image Processing and Computer Applications (DIPCA), pp. 123–129. Cited by: §1.