arXiv is now an independent nonprofit! Learn more
License: CC BY 4.0
arXiv:2609.00413v1 [cs.AI] 10 Jul 2026
\setcctype

by

Dependency-Aware Chain-of-Thought Compression for Financial Reasoning

Wenjun Wu wenjun5@illinois.edu University of Illinois Urbana-ChampaignUrbanaUSA , Lei Fu fuleiac@gmail.com Independent ResearcherSan JoseUSA , Kejian Tong tongcs2021@gmail.com Independent ResearcherMukilteoUSA , Tao Ning ntgd1102@gmail.com Syracuse UniversitySan JoseUSA and Sichen Zhao zhao.siche@northeastern.edu Northeastern UniversityBostonUSA
(2026)
Abstract.

Chain of thought prompting improves complex reasoning, but its long intermediate traces create substantial inference cost and hinder practical deployment in financial settings. We present a Hierarchical Semantic Distillation Network, HSDN, for compressing reasoning chains while preserving answer accuracy and logical coherence. The framework combines semantic segmentation, dependency graph construction, dual encoder importance scoring, constrained segment selection, and local boundary rewriting. A frozen Qwen3 4B model is used only for feature extraction and final answer generation, while the compression process remains structured and interpretable. On the AFAC2025 benchmark, HSDN achieves 91.0% accuracy with 68.4% compression, outperforming strong compression baselines in overall score and reasoning coherence. The results show that graph guided compression is effective for high stakes financial reasoning tasks.

chain of thought compression, financial reasoning, dependency graph, prompt compression, structured distillation, long context
††journalyear: 2026††copyright: cc††conference: 2026 3rd International Conference on Machine Learning and Intelligent Computing; April 24–26, 2026; Zhengzhou, China††booktitle: 2026 3rd International Conference on Machine Learning and Intelligent Computing (MLIC 2026), April 24–26, 2026, Zhengzhou, China††doi: 10.1145/3829441.3829513††isbn: 979-8-4007-2465-7/2026/04††ccs: Computing methodologies Natural language processing††ccs: Computing methodologies Machine learning††ccs: Computing methodologies Knowledge representation and reasoning††ccs: Information systems Summarization††ccs: Information systems Decision support systems

1. Introduction

Large language models have shown strong performance on multi step reasoning tasks, and chain of thought prompting has become a standard mechanism for improving intermediate deliberation quality (Wei et al., 2022). This capability is especially valuable in financial applications, where models must integrate textual evidence, numerical calculation, and compliance oriented interpretation under strict correctness requirements. At the same time, longer reasoning traces increase latency, memory usage, and serving cost, which limits their practicality in real world systems.Recent roofline-guided co-optimization work on Arm CPUs highlights how memory bandwidth and low-precision throughput constraints can be addressed with mixed-precision kernels and fused attention, which is relevant to our efficiency-oriented design (Zhou, 2026b).This efficiency challenge is closely related to serverless AI inference, where systems must balance cold start latency against the cost of retaining idle GPU resources, motivating adaptive lifecycle management strategies such as AdaScale (Zhou, 2026a). Zero shot reasoning studies further indicate that reasoning quality depends not only on model scale, but also on how intermediate steps are organized and exposed during inference (Kojima et al., 2022).Zero shot reasoning studies further indicate that reasoning quality depends not only on model scale, but also on how intermediate steps are organized and exposed during inference (Kojima et al., 2022). Recent multi-agent troubleshooting frameworks such as PRISM further suggest that specialized role decomposition and evidence verification can improve complex reasoning workflows (Yan et al., 2026). Existing approaches still face an important gap for high stakes financial reasoning. Direct generation is efficient but often omits key logical steps, while full chain of thought preserves detail at the expense of excessive verbosity. More advanced deliberate reasoning strategies improve search over reasoning paths, yet they do not explicitly address how to compress a completed chain while retaining dependency structure and numerical faithfulness (Yao et al., 2023). In financial scenarios, this omission is critical because dropping a seemingly minor step can invalidate later computations or break the justification trail needed for auditability. To address this problem, we propose a hierarchical semantic distillation framework for compressing reasoning chains in a structured manner. Our method first segments reasoning into semantic units, then builds a directed dependency graph, scores segment importance with question aware representations, and performs globally optimal selection under length and dependency constraints. A lightweight boundary rewriter restores fluency after segment removal, while a frozen large language model is used only for semantic features and final answer generation. This design yields an interpretable compression pipeline that reduces reasoning length without sacrificing the logical continuity required in financial decision support.

2. Related Work

Research on handling long inputs has largely focused on global mechanisms that reduce the cost of full sequence processing. Sparse attention architectures such as Longformer and BigBird improve scalability for long documents by restricting or restructuring attention patterns, offering a strong foundation for efficient long context modeling (Beltagy et al., 2020; Zaheer et al., 2020). However, these methods mainly optimize representation efficiency at the token level and do not directly determine which reasoning steps should be retained for answer faithful chain compression. A second line of work studies fine grained and dynamic reduction, where models selectively remove less important tokens during inference. TR BERT introduces dynamic token reduction conditioned on task relevance, and PoWER BERT progressively eliminates low impact word vectors to accelerate inference while maintaining prediction quality (Ye et al., 2021; Goyal et al., 2020). These methods demonstrate the value of adaptive compression, but they typically operate on local token salience and do not model explicit logical dependencies among reasoning segments, which are crucial in multi step financial inference.Recent work on dynamic retrieval-augmented generation further shows that selective tool use and sufficiency-aware routing can improve robustness when static context is insufficient (Liang et al., 2026). Our approach is also related to work on structured summarization and faithful generation.Recent hybrid deep learning work on supply chain delay prediction further underscores the value of jointly modeling temporal dynamics and graph structure in operational decision systems (Xue et al., 2026a).A recent risk-aware dynamic routing framework demonstrates how spatiotemporal graph neural networks can support resilient decision-making under congestion and fluctuating demand in large-scale logistics systems (Xue et al., 2026b).In particular, recent hybrid architectures that combine pyramid-style semantic encoding with graph attention and language-model-assisted rationale distillation further suggest the value of multi-granularity feature extraction for complex threat narratives (Xu, 2026b). HeterSumGraph shows that graph representations can capture document level relations for extractive summarization, while PRIMERA improves long document summarization through pyramid based pretraining (Wang et al., 2020; Xiao et al., 2022). For generation faithfulness, FactPEGASUS highlights the importance of preserving factual consistency during rewriting and compression (Wan and Bansal, 2022).Complementary to these generation-focused approaches, pairwise verification methods such as SENTINEL can detect subtle semantic inconsistencies between two candidate outputs for the same source, offering a useful perspective for assessing whether compression introduces manipulative distortions (Xu, 2026a).

3. Methodology

Chain-of-thought reasoning is essential for complex financial inference, yet verbose reasoning sequences impose substantial computational overhead. This paper presents a hierarchical semantic distillation framework that compresses lengthy reasoning chains while preserving logical integrity, addressing challenges unique to financial applications such as tabular data interpretation, multi-step numerical computation, and regulatory compliance verification. The framework introduces a structured pipeline comprising four key components: a graph-based dependency parser that constructs explicit reasoning topology to capture causal relationships and identify dispensable content; a dual-encoder architecture that scores segment importance through cross-modal attention between question semantics and reasoning content; a dynamic programming solver that guarantees globally optimal segment selection under length budgets while respecting dependency constraints; and a lightweight sequence-to-sequence rewriter that ensures coherence at segment boundaries without introducing factual inconsistencies. A frozen large language model serves solely as a semantic feature extractor and answer generator, keeping the core compression logic in interpretable algorithmic components. Evaluation on financial reasoning benchmarks demonstrates substantial length reduction while maintaining answer accuracy, with the graph-based formulation providing transparent compression rationale for human verification in high-stakes applications. The overall architecture is illustrated in Fig. 1.

Refer to caption
Figure 1. Overview of the Hierarchical Semantic Distillation Network (HSDN). The pipeline comprises five stages: semantic segmentation via BiLSTM-CRF, dependency graph construction with biaffine attention, dual-encoder importance scoring, constrained segment selection via dynamic programming, and boundary rewriting with a copy-augmented seq2seq model. A frozen Qwen3-4B model provides semantic features and generates final answers.

4. Algorithm and Model

4.1. Semantic Segmentation

The first stage partitions the continuous reasoning text into discrete semantic units that serve as atomic elements for subsequent processing. Unlike sentence-level segmentation which often breaks logical units inappropriately, we train a specialized boundary detector that respects reasoning step boundaries.

4.1.1. Boundary Detection Network

We employ a bidirectional LSTM network augmented with conditional random fields to identify segment boundaries. Given the token sequence 𝐜=(c1,c2,…,cT)\mathbf{c}=(c_{1},c_{2},\ldots,c_{T}), the network first computes contextual representations:

(1) 𝐡→t=LSTM→​(ct,𝐡→t−1)\overrightarrow{\mathbf{h}}_{t}=\text{LSTM}_{\rightarrow}(c_{t},\overrightarrow{\mathbf{h}}_{t-1})
(2) 𝐡←t=LSTM←​(ct,𝐡←t+1)\overleftarrow{\mathbf{h}}_{t}=\text{LSTM}_{\leftarrow}(c_{t},\overleftarrow{\mathbf{h}}_{t+1})

The concatenated representation 𝐡t=[𝐡→t;𝐡←t]\mathbf{h}_{t}=[\overrightarrow{\mathbf{h}}_{t};\overleftarrow{\mathbf{h}}_{t}] captures both preceding and following context, which proves essential for detecting boundaries that depend on what comes next. Early experiments with unidirectional models showed poor performance on boundaries preceding numerical computations, where the boundary significance only becomes apparent from the subsequent calculation.

The boundary probability at each position is computed through a feedforward layer:

(3) 𝐞t=𝐖e⋅tanh⁡(𝐖h​𝐡t+𝐛h)+𝐛e\mathbf{e}_{t}=\mathbf{W}_{e}\cdot\tanh(\mathbf{W}_{h}\mathbf{h}_{t}+\mathbf{b}_{h})+\mathbf{b}_{e}

where 𝐞t∈ℝ2\mathbf{e}_{t}\in\mathbb{R}^{2} represents emission scores for boundary and non-boundary labels.

To ensure globally consistent segmentation, we apply a linear-chain CRF layer that models label transitions:

(4) P​(𝐲|𝐜)=1Z​(𝐜)​exp⁡(∑t=1T𝐞t​[yt]+∑t=1T−1𝐀​[yt,yt+1])P(\mathbf{y}|\mathbf{c})=\frac{1}{Z(\mathbf{c})}\exp\left(\sum_{t=1}^{T}\mathbf{e}_{t}[y_{t}]+\sum_{t=1}^{T-1}\mathbf{A}[y_{t},y_{t+1}]\right)

where 𝐀∈ℝ2×2\mathbf{A}\in\mathbb{R}^{2\times 2} is the transition matrix and Z​(𝐜)Z(\mathbf{c}) is the partition function computed via the forward algorithm. The CRF layer prevents degenerate solutions such as consecutive boundaries or excessively long segments, which frequently occurred with independent classification.

4.1.2. Financial Domain Adaptations

Financial reasoning text presents unique segmentation challenges that required specific adaptations. Numerical expressions spanning multiple tokens, such as percentage changes or currency amounts, must remain intact within segments. We address this by incorporating a numerical span detector that identifies contiguous numerical expressions:

(5) NumSpan(t)=𝟙[∃(i,j):i≤t≤j∧IsNumeric(ci,…,cj)]\text{NumSpan}(t)=\mathbb{1}\left[\exists(i,j):i\leq t\leq j\land\text{IsNumeric}(c_{i},\ldots,c_{j})\right]

During CRF decoding, we modify transition scores to prohibit boundaries within detected numerical spans by setting
𝐀​[yt−1,BOUNDARY]=−∞\mathbf{A}[y_{t-1},{\scriptstyle\text{BOUNDARY}}]=-\infty when NumSpan​(t)=1\text{NumSpan}(t)=1 and the span continues from position t−1t-1.

Additionally, financial text frequently contains table references and structured data mentions that should not be split. We found that simply expanding the context window was insufficient; instead, we pretrain the boundary detector on a auxiliary task of table cell boundary detection, which transfers effectively to reasoning text segmentation.

4.2. Dependency Graph Construction

With the reasoning chain segmented into units 𝐒=(s1,s2,…,sK)\mathbf{S}=(s_{1},s_{2},\ldots,s_{K}), we construct a directed acyclic graph 𝒢=(𝒱,ℰ)\mathcal{G}=(\mathcal{V},\mathcal{E}) that explicitly encodes logical dependencies between segments. This graph serves as the structural backbone for compression decisions, ensuring that removing a segment does not orphan its dependents. Fig. 2 illustrates the graph construction pipeline.

Refer to caption
Figure 2. Overview of the Hierarchical Semantic Distillation Network (HSDN). The pipeline comprises five stages: semantic segmentation via BiLSTM-CRF, dependency graph construction with biaffine attention, dual-encoder importance scoring, constrained segment selection via dynamic programming, and boundary rewriting with a copy-augmented seq2seq model. A frozen Qwen3-4B model provides semantic features and generates final answers.

4.2.1. Node Representation

Each segment sks_{k} becomes a node in the graph. We compute node embeddings by mean-pooling token representations from a frozen language model encoder:

(6) 𝐯k=1|sk|​∑t∈skLLMenc​(ct)\mathbf{v}_{k}=\frac{1}{|s_{k}|}\sum_{t\in s_{k}}\text{LLM}_{\text{enc}}(c_{t})

To capture segment-level semantics beyond token averaging, we apply a learned projection with residual connection:

(7) 𝐯~k=𝐯k+ReLU​(𝐖v​𝐯k+𝐛v)\tilde{\mathbf{v}}_{k}=\mathbf{v}_{k}+\text{ReLU}(\mathbf{W}_{v}\mathbf{v}_{k}+\mathbf{b}_{v})

This additional projection layer allows the model to learn task-specific representations while leveraging the pretrained encoder’s semantic knowledge. We experimented with fine-tuning the encoder but found that freezing it and adding the projection layer yielded comparable performance with substantially reduced memory requirements during training.

4.2.2. Edge Prediction

Edges represent logical dependencies where the source segment provides information required by the target segment. We predict edges using a biaffine attention mechanism that has proven effective for dependency parsing:

(8) score​(si→sj)=𝐡i⊤​𝐖biaff​𝐡j+𝐰i⊤​𝐡i+𝐰j⊤​𝐡j+b\text{score}(s_{i}\rightarrow s_{j})=\mathbf{h}_{i}^{\top}\mathbf{W}_{\text{biaff}}\mathbf{h}_{j}+\mathbf{w}_{i}^{\top}\mathbf{h}_{i}+\mathbf{w}_{j}^{\top}\mathbf{h}_{j}+b

where 𝐡i=MLPhead​(𝐯~i)\mathbf{h}_{i}=\text{MLP}_{\text{head}}(\tilde{\mathbf{v}}_{i}) and 𝐡j=MLPdep​(𝐯~j)\mathbf{h}_{j}=\text{MLP}_{\text{dep}}(\tilde{\mathbf{v}}_{j}) are head and dependent representations computed through separate multilayer perceptrons.

The edge probability is obtained via sigmoid activation:

(9) P​(ei​j=1)=σ​(score​(si→sj))P(e_{ij}=1)=\sigma(\text{score}(s_{i}\rightarrow s_{j}))

A critical challenge we encountered was handling long-range dependencies that span many intermediate segments. The biaffine scorer tends to underestimate such dependencies due to the difficulty of directly relating distant segments. Our solution introduces a path-augmented scoring term:

(10) scoreaug​(si→sj)=score​(si→sj)+γ⋅Φi​j\text{score}_{\text{aug}}(s_{i}\rightarrow s_{j})=\text{score}(s_{i}\rightarrow s_{j})+\gamma\cdot\Phi_{ij}

where Φi​j=maxp∈Paths​(i,j)​∏(a,b)∈pP​(ea​b=1)\Phi_{ij}=\max_{p\in\text{Paths}(i,j)}\prod_{(a,b)\in p}P(e_{ab}=1) captures transitive dependency strength.

4.2.3. Acyclicity Constraint

The dependency graph must be acyclic to represent valid logical flow. We enforce this through a differentiable acyclicity regularization term based on the matrix exponential characterization:

(11) ℒDAG=tr​(e𝐀⊙𝐀)−K\mathcal{L}_{\text{DAG}}=\text{tr}\left(e^{\mathbf{A}\odot\mathbf{A}}\right)-K

where 𝐀∈[0,1]K×K\mathbf{A}\in[0,1]^{K\times K} is the adjacency matrix with Ai​j=P​(ei​j=1)A_{ij}=P(e_{ij}=1), and ⊙\odot denotes element-wise multiplication. This regularizer equals zero if and only if the expected graph is acyclic.

During inference, we apply a greedy pruning procedure that removes the lowest-probability edge from any detected cycle, iterating until the graph is acyclic. In practice, the regularization during training ensures cycles are rare, typically requiring fewer than two pruning iterations.

4.3. Importance Scoring and Selection

Given the dependency graph, we must select which segments to retain in the compressed output. This stage comprises two components: a neural importance scorer that evaluates each segment’s contribution to answering the question, and a dynamic programming algorithm that finds the optimal selection respecting dependency constraints. The scoring and selection mechanism is depicted in Fig. 3.

4.3.1. Dual-Encoder Importance Scorer

The importance scorer employs a dual-encoder architecture that separately encodes the question and each reasoning segment, then computes relevance through cross-attention. This design allows efficient scoring of all segments with a single question encoding pass.

The question encoder applies self-attention layers to produce a contextualized representation:

(12) 𝐐=SelfAttn(L)​(Embed​(𝐱))∈ℝn×d\mathbf{Q}=\text{SelfAttn}^{(L)}(\text{Embed}(\mathbf{x}))\in\mathbb{R}^{n\times d}

where nn is the question length and LL is the number of layers.

For each segment, we compute cross-attention scores against the question:

(13) αk​t=exp⁡(𝐯~k⊤​𝐖q​𝐐t/d)∑t′=1nexp⁡(𝐯~k⊤​𝐖q​𝐐t′/d)\alpha_{kt}=\frac{\exp(\tilde{\mathbf{v}}_{k}^{\top}\mathbf{W}_{q}\mathbf{Q}_{t}/\sqrt{d})}{\sum_{t^{\prime}=1}^{n}\exp(\tilde{\mathbf{v}}_{k}^{\top}\mathbf{W}_{q}\mathbf{Q}_{t^{\prime}}/\sqrt{d})}
(14) 𝐪k=∑t=1nαk​t​𝐐t\mathbf{q}_{k}=\sum_{t=1}^{n}\alpha_{kt}\mathbf{Q}_{t}

The attended question representation 𝐪k\mathbf{q}_{k} captures which aspects of the question segment sks_{k} addresses. The importance score combines this relevance signal with segment-intrinsic features:

(15) Ik=σ​(𝐰⊤​[𝐯~k;𝐪k;𝐯~k⊙𝐪k;fk]+b)I_{k}=\sigma\left(\mathbf{w}^{\top}[\tilde{\mathbf{v}}_{k};\mathbf{q}_{k};\tilde{\mathbf{v}}_{k}\odot\mathbf{q}_{k};f_{k}]+b\right)

where fkf_{k} is a feature vector containing segment length, position, numerical content indicators, and graph centrality measures.

An important trick we discovered is incorporating the segment’s graph centrality into the importance score. Segments with high out-degree, meaning many other segments depend on them, tend to contain foundational information that should be preserved. We compute PageRank scores on the dependency graph and include them in fkf_{k}:

(16) PR​(sk)=1−dK+d​∑sj∈Parents​(sk)PR​(sj)|Children​(sj)|\text{PR}(s_{k})=\frac{1-d}{K}+d\sum_{s_{j}\in\text{Parents}(s_{k})}\frac{\text{PR}(s_{j})}{|\text{Children}(s_{j})|}

where d=0.85d=0.85 is the damping factor.

4.3.2. Constrained Selection via Dynamic Programming

Selecting the optimal subset of segments is formulated as a constrained optimization problem:

(17) max𝐳∈{0,1}K​∑k=1Kzk⋅Ik\max_{\mathbf{z}\in\{0,1\}^{K}}\sum_{k=1}^{K}z_{k}\cdot I_{k}

subject to:

(18) ∑k=1Kzk⋅|sk|≤Lbudget\sum_{k=1}^{K}z_{k}\cdot|s_{k}|\leq L_{\text{budget}}
(19) zj=1⟹zi=1∀(si,sj)∈ℰz_{j}=1\implies z_{i}=1\quad\forall(s_{i},s_{j})\in\mathcal{E}

The first constraint enforces the length budget, while the second ensures that if a segment is selected, all its dependencies are also selected. This problem generalizes the knapsack problem with precedence constraints.

We solve it via dynamic programming on the topologically sorted graph. Let dp​[k]​[l]\text{dp}[k][l] denote the maximum importance achievable considering segments 1,…,k1,\ldots,k with total length exactly ll. The recurrence is:

(20) dp​[k]​[l]=max⁡{dp​[k−1]​[l]dp​[k−1]​[l−|sk|]+Ik\text{dp}[k][l]=\max\begin{cases}\text{dp}[k-1][l]\\ \text{dp}[k-1][l-|s_{k}|]+I_{k}\end{cases}

where the second case requires l≥|sk|l\geq|s_{k}| and DepsSelected​(k)=true\text{DepsSelected}(k)=\text{true}.

The time complexity is O​(K⋅Lbudget)O(K\cdot L_{\text{budget}}), which is efficient for typical reasoning chain lengths. We implement the dependency check through bit manipulation, maintaining a bitmask of selected segments and verifying parent inclusion in constant time after preprocessing parent masks.

Refer to caption
Figure 3. Dual-encoder importance scoring and constrained selection. The question and segment streams are fused via cross-attention to produce importance scores IkI_{k}, which are then optimized through dynamic programming under length budget and dependency constraints.

4.4. Boundary Rewriting

Directly concatenating retained segments often produces incoherent text with abrupt transitions. The rewriting module performs local edits at segment boundaries to restore fluency while preserving factual content.

4.4.1. Boundary Context Extraction

For each pair of consecutively selected segments (si,sj)(s_{i},s_{j}) where intermediate segments were removed, we extract boundary context:

(21) 𝐛i​j=[Suffix​(si,w);Prefix​(sj,w)]\mathbf{b}_{ij}=[\text{Suffix}(s_{i},w);\text{Prefix}(s_{j},w)]

where Suffix and Prefix extract w=15w=15 tokens from segment ends and beginnings respectively.

4.4.2. Seq2Seq Rewriter

A lightweight transformer-based sequence-to-sequence model takes boundary context as input and generates a smoothed transition:

(22) P​(𝐫|𝐛i​j)=∏t=1|𝐫|P​(rt|r<t,𝐛i​j)P(\mathbf{r}|\mathbf{b}_{ij})=\prod_{t=1}^{|\mathbf{r}|}P(r_{t}|r_{<t},\mathbf{b}_{ij})

To prevent hallucination, the decoder employs a copy mechanism that biases generation toward input tokens:

(23) P​(rt=w)=λt​Pgen​(w)+(1−λt)​∑i:bi=wαt​iP(r_{t}=w)=\lambda_{t}P_{\text{gen}}(w)+(1-\lambda_{t})\sum_{i:b_{i}=w}\alpha_{ti}

where λt\lambda_{t} is a learned gate and αt​i\alpha_{ti} are copy attention weights. The rewriter is trained on synthetic boundary pairs created by randomly removing segments from correct reasoning chains, avoiding error propagation from the full pipeline.

4.4.3. Consistency Verification

We verify factual consistency by computing embedding similarity between the original boundary region and the rewritten version:

(24) sim​(𝐛i​j,𝐫)=Embed​(𝐛i​j)⊤​Embed​(𝐫)‖Embed​(𝐛i​j)‖⋅‖Embed​(𝐫)‖\text{sim}(\mathbf{b}_{ij},\mathbf{r})=\frac{\text{Embed}(\mathbf{b}_{ij})^{\top}\text{Embed}(\mathbf{r})}{\|\text{Embed}(\mathbf{b}_{ij})\|\cdot\|\text{Embed}(\mathbf{r})\|}

If similarity falls below threshold τ\tau, we fall back to simple concatenation with a generic connective phrase.

4.5. Training Procedure

The framework is trained in three stages to ensure stable optimization and effective knowledge transfer.

4.5.1. Stage 1: Component Pretraining

The segmentation network is pretrained on manually annotated reasoning chains using CRF negative log-likelihood:

(25) ℒseg=−log⁡P​(𝐲∗|𝐜)\mathcal{L}_{\text{seg}}=-\log P(\mathbf{y}^{*}|\mathbf{c})

The dependency graph predictor is pretrained on synthetic dependency data generated by prompting a large language model to annotate reasoning chain dependencies.

4.5.2. Stage 2: Joint Scorer Training

The importance scorer and graph predictor are trained jointly using a multi-task objective:

(26) ℒjoint=ℒscore+λ1​ℒedge+λ2​ℒDAG\mathcal{L}_{\text{joint}}=\mathcal{L}_{\text{score}}+\lambda_{1}\mathcal{L}_{\text{edge}}+\lambda_{2}\mathcal{L}_{\text{DAG}}

The scoring loss uses ground-truth labels derived from oracle compression identifying the minimal segment subset producing correct answers:

(27) ℒscore=−∑k=1K[yk​log⁡Ik+(1−yk)​log⁡(1−Ik)]\mathcal{L}_{\text{score}}=-\sum_{k=1}^{K}\left[y_{k}\log I_{k}+(1-y_{k})\log(1-I_{k})\right]
(28) ℒedge=−∑i,j[ei​j∗​log⁡P​(ei​j)+(1−ei​j∗)​log⁡(1−P​(ei​j))]\mathcal{L}_{\text{edge}}=-\sum_{i,j}\left[e_{ij}^{*}\log P(e_{ij})+(1-e_{ij}^{*})\log(1-P(e_{ij}))\right]

4.5.3. Stage 3: End-to-End Refinement

The complete pipeline is fine-tuned using reinforcement learning with a reward combining accuracy and compression:

(29) R=𝟙​[correct]⋅(1+β⋅|𝐜|−|𝐜∗||𝐜|)−𝟙​[¬correct]⋅ρR=\mathbb{1}[\text{correct}]\cdot\left(1+\beta\cdot\frac{|\mathbf{c}|-|\mathbf{c}^{*}|}{|\mathbf{c}|}\right)-\mathbb{1}[\neg\text{correct}]\cdot\rho

We employ REINFORCE with baseline subtraction to reduce gradient variance:

(30) ∇ℒRL=−𝔼​[(R−b)​∇log⁡πθ​(𝐳|𝐱,𝐜)]\nabla\mathcal{L}_{\text{RL}}=-\mathbb{E}\left[(R-b)\nabla\log\pi_{\theta}(\mathbf{z}|\mathbf{x},\mathbf{c})\right]

where bb is an exponential moving average of recent rewards.

4.6. LLM Integration

The large language model component serves two specific roles in our framework, deliberately limited to leverage its strengths while avoiding the interpretability and efficiency drawbacks of end-to-end neural compression.

4.6.1. Feature Extraction

We use a frozen Qwen3-4B model as a semantic feature extractor, accessing its intermediate layer representations to initialize segment embeddings. Specifically, we extract features from layer 16 of 32, which empirically balances semantic abstraction with surface-level detail:

(31) LLMenc​(ct)=𝐇t(16)\text{LLM}_{\text{enc}}(c_{t})=\mathbf{H}^{(16)}_{t}

The frozen encoder provides rich pretrained representations without the computational cost of fine-tuning or the risk of catastrophic forgetting.

4.6.2. Answer Generation

After compression, the retained segments are concatenated with minimal rewriting and fed to the language model for final answer generation:

(32) P​(𝐚|𝐱,𝐜∗)=∏t=1|𝐚|P​(at|a<t,𝐱,𝐜∗)P(\mathbf{a}|\mathbf{x},\mathbf{c}^{*})=\prod_{t=1}^{|\mathbf{a}|}P(a_{t}|a_{<t},\mathbf{x},\mathbf{c}^{*})

We apply standard decoding with temperature τ=0.7\tau=0.7 and nucleus sampling with p=0.9p=0.9 to generate diverse answer candidates for the best-of-5 evaluation protocol.

4.7. Complexity Analysis

The computational complexity of each component scales tractably with input size. Segmentation requires O​(T)O(T) time for the BiLSTM pass and O​(T)O(T) for CRF decoding. Graph construction involves O​(K2)O(K^{2}) edge predictions where K≪TK\ll T is the number of segments. The dynamic programming selection runs in O​(K⋅Lbudget)O(K\cdot L_{\text{budget}}) time. Boundary rewriting processes at most K−1K-1 boundaries with constant-length inputs each.

The overall complexity is dominated by the LLM feature extraction and answer generation, which are unavoidable for the task. Our framework adds minimal overhead to these fixed costs while providing structured, interpretable compression that pure neural approaches cannot match.

5. Evaluation Metrics

We adopt four metrics following the challenge protocol. Best-of-5 accuracy measures correctness when any of five samples matches the reference:

(33) Acc=1N​∑i=1N𝟙​[⋁j=15Match​(oi(j),ℛi)]\text{Acc}=\frac{1}{N}\sum_{i=1}^{N}\mathbb{1}\left[\bigvee_{j=1}^{5}\text{Match}(o_{i}^{(j)},\mathcal{R}_{i})\right]

Compression ratio quantifies length reduction:

(34) CR=1−∑i=1N|𝐜i∗|∑i=1N|𝐜i|\text{CR}=1-\frac{\sum_{i=1}^{N}|\mathbf{c}_{i}^{*}|}{\sum_{i=1}^{N}|\mathbf{c}_{i}|}

The competition score penalizes incorrect answers with maximum length LmaxL_{\max}:

(35) Score=−∑i=1NLeff​(qi),Leff​(qi)={minj:correct⁡|oi(j)|if correctLmaxotherwise\text{Score}=-\sum_{i=1}^{N}L_{\text{eff}}(q_{i}),\quad L_{\text{eff}}(q_{i})=\begin{cases}\min\limits_{j:\text{correct}}|o_{i}^{(j)}|&\text{if correct}\\ L_{\max}&\text{otherwise}\end{cases}

Reasoning coherence score evaluates logical dependency preservation:

(36) RCS=1N​∑i=1N|ℰ​(𝒢i)∩ℰ​(𝒢i∗)||ℰ​(𝒢i∗)|\text{RCS}=\frac{1}{N}\sum_{i=1}^{N}\frac{|\mathcal{E}(\mathcal{G}_{i})\cap\mathcal{E}(\mathcal{G}_{i}^{*})|}{|\mathcal{E}(\mathcal{G}_{i}^{*})|}

6. Experiment Results

Table 1 presents the main comparison and ablation study on the AFAC2025 benchmark. And the changes in model training indicators are shown in Fig4

Refer to caption
Figure 4. Model indicator change chart.

.

Table 1. Performance comparison and ablation study
Method Acc.(%) CR(%) RCS Score
Qwen3-4B (Full CoT) 94.0 0.0 1.000 -184700
Qwen3-4B (Direct) 71.0 95.2 – -97890
LLMLingua-2 85.0 58.3 0.724 -83650
CompAct 86.0 61.7 0.756 -78420
RECOMP 88.0 54.2 0.812 -89760
LongLLMLingua 87.0 63.5 0.743 -74380
HSDN (Ours) 91.0 68.4 0.867 -61250
   w/o Dependency Graph 87.0 71.2 0.712 -72340
   w/o Boundary Rewriter 90.0 68.4 0.791 -63120
   w/o PageRank Features 89.0 67.8 0.834 -65780
   w/o RL Fine-tuning 89.0 65.3 0.851 -68450

As shown in Table 1, HSDN achieves the best trade-off between accuracy and compression. The dependency graph contributes most significantly, with its removal causing 4% accuracy drop. Compared to perplexity-based methods like LLMLingua-2, our graph-guided approach better preserves reasoning coherence.

7. Conclusion

In this work, we presented FinStack-Net, a hierarchical ensemble framework combining LightGBM, CatBoost, and a deep neural network with residual and attention mechanisms for fraud and gambling account detection. Through comprehensive data preprocessing, feature engineering, and hyperparameter optimization, the model achieved state-of-the-art results. Ablation studies demonstrated the importance of each architectural component, highlighting the robustness of the ensemble strategy. Future research will explore the integration of temporal sequence models and graph-based transaction analysis to further enhance detection performance.

References

  • I. Beltagy, M. E. Peters, and A. Cohan (2020) Longformer: the long-document transformer. arXiv preprint arXiv:2004.05150. Cited by: §2.
  • S. Goyal, A. R. Choudhury, S. Raje, V. Chakaravarthy, Y. Sabharwal, and A. Verma (2020) Power-bert: accelerating bert inference via progressive word-vector elimination. In International conference on machine learning, pp. 3690–3699. Cited by: §2.
  • T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa (2022) Large language models are zero-shot reasoners. Advances in neural information processing systems 35, pp. 22199–22213. Cited by: §1.
  • P. Liang, M. Yuan, J. Liu, J. Yang, X. Li, W. Yan, and Y. Wu (2026) DynaRAG: bridging static and dynamic knowledge in retrieval-augmented generation. In 2026 9th International Symposium on Big Data and Applied Statistics (ISBDAS), pp. 442–445. Cited by: §2.
  • D. Wan and M. Bansal (2022) FactPEGASUS: factuality-aware pre-training and fine-tuning for abstractive summarization. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 1010–1028. Cited by: §2.
  • D. Wang, P. Liu, Y. Zheng, X. Qiu, and X. Huang (2020) Heterogeneous graph neural networks for extractive document summarization. In Proceedings of the 58th annual meeting of the association for computational linguistics, pp. 6209–6219. Cited by: §2.
  • J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al. (2022) Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems 35, pp. 24824–24837. Cited by: §1.
  • W. Xiao, I. Beltagy, G. Carenini, and A. Cohan (2022) PRIMERA: pyramid-based masked sentence pre-training for multi-document summarization. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 5245–5263. Cited by: §2.
  • Y. Xu (2026a) Detecting manipulated llm outputs via siamese transformers and semantic dependency graphs. In 2026 2nd International Conference on Artificial Intelligence and Computational Intelligence (AICI 2026), New York, NY, USA. External Links: ISBN 979-8-4007-2279-0/2026/02, Document Cited by: §2.
  • Y. Xu (2026b) Pyramid convolution and bidirectional graph attention for cyber threat detection from unstructured text. In Proceedings of the 2026 International Conference on Artificial Intelligence and Control, pp. 561–567. Cited by: §2.
  • Z. Xue, M. Huo, and Y. Wang (2026a) EAGLE: edge-aware graph learning for proactive delivery delay prediction in smart logistics networks. arXiv preprint arXiv:2604.05254. Cited by: §2.
  • Z. Xue, S. Zhao, Y. Qi, X. Zeng, and Z. Yu (2026b) Resilient routing: risk-aware dynamic routing in smart logistics via spatiotemporal graph learning. arXiv preprint arXiv:2601.13632. Cited by: §2.
  • W. Yan, Y. Wu, P. Liang, M. Yuan, J. Liu, J. Yang, and X. Li (2026) PRISM: pipeline for root-cause investigation via specialized multi-agents. In 2026 International Conference on Generative Artificial Intelligence and Information Security (GAIIS), pp. 709–712. Cited by: §1.
  • S. Yao, D. Yu, J. Zhao, I. Shafran, T. Griffiths, Y. Cao, and K. Narasimhan (2023) Tree of thoughts: deliberate problem solving with large language models. Advances in neural information processing systems 36, pp. 11809–11822. Cited by: §1.
  • D. Ye, Y. Lin, Y. Huang, and M. Sun (2021) Tr-bert: dynamic token reduction for accelerating bert inference. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 5798–5809. Cited by: §2.
  • M. Zaheer, G. Guruganesh, K. A. Dubey, J. Ainslie, C. Alberti, S. Ontanon, P. Pham, A. Ravula, Q. Wang, L. Yang, et al. (2020) Big bird: transformers for longer sequences. Advances in neural information processing systems 33, pp. 17283–17297. Cited by: §2.
  • Q. Zhou (2026a) AdaScale: predictive and utility-aware autoscaling for serverless ai inference. In 2026 6th International Conference on Artificial Intelligence and Industrial Technology Applications (AIITA), pp. 544–550. Cited by: §1.
  • Q. Zhou (2026b) Roofline-guided mixed quantization and kernel co-optimization for efficient large language model inference on arm cpus. In 2026 3rd International Conference on Digital Image Processing and Computer Applications (DIPCA), pp. 123–129. Cited by: §1.