arXiv is now an independent nonprofit! Learn more
License: CC BY 4.0
arXiv:2609.02336v1 [cs.AI] 02 Sep 2026

SALA: Semantic-Aware Logical Alignment for Complex Reasoning in In-Context Learning

Zhao Ji Affiliation: School of Software Engineering, Sun Yat-sen University, Zhuhai, China Affiliation: Zhuhai Key Laboratory of Trusted Large Language Models, Zhuhai, China    Wenqing Chen ††thanks: Corresponding author. Affiliation: School of Software Engineering, Sun Yat-sen University, Zhuhai, China Affiliation: Zhuhai Key Laboratory of Trusted Large Language Models, Zhuhai, China    Zhixuan Chu Affiliation: The State Key Laboratory of Blockchain and Data SecurityZhejiang University, Hangzhou, China    Jianxing Yu Affiliation: School of Artificial Intelligence, Sun Yat-sen University, Zhuhai, China Affiliation: Key Laboratory of Sustainable Tourism Smart Assessment TechnologyMinistry of Culture and Tourism, Zhuhai, China    Jingping Liu Affiliation: School of Software Engineering, Sun Yat-sen University, Zhuhai, China Affiliation: Zhuhai Key Laboratory of Trusted Large Language Models, Zhuhai, China    Shanhe Zhao Affiliation: Merchants Union Consumer Finance Company Limited, Shenzhen, China    Zibin Zheng Affiliation: School of Software Engineering, Sun Yat-sen University, Zhuhai, China Affiliation: Zhuhai Key Laboratory of Trusted Large Language Models, Zhuhai, China
Abstract

Effective in-context learning (ICL) for complex reasoning relies on selecting the right demonstrations. Traditional retrieval methods based on surface similarity fail to capture the underlying problem-solving logic. Recent logic-based methods address this by matching predefined reasoning steps, but the rigid rules and exact-match criteria is improper to handle flexible or diverse reasoning processes. To address the problem, we propose SALA, a Semantic-Aware Logical Alignment framework. Instead of relying on a fixed inventory, SALA automatically learns task-specific reasoning operations. It then embeds these operations into a continuous semantic space and uses dynamic time warping (DTW) to align the reasoning sequences. This approach allows for soft, flexible matching of reasoning logic while remaining highly interpretable. Experiments across four reasoning benchmarks and three LLMs demonstrate that SALA outperforms existing demonstration selection methods. Further analysis confirms the roles of the operation induction and the logical semantic alignment.

Refer to caption
Figure 1: An example of the induction process of task-specific operations.

1 Introduction

Large Language Models (LLMs) Ouyang et al. (2022); Achiam et al. (2023); Touvron et al. (2023); Bai et al. (2023); Ye et al. (2023b); Bahrini et al. (2023) have shown strong in-context learning (ICL) ability Dong et al. (2024); Wies et al. (2023); Luo et al. (2024); Bhattamishra et al. (2024), enabling them to solve complex reasoning tasks from a few labeled demonstrations. ICL performance largely depends on demonstration selection, since useful examples can improve accuracy Ye et al. (2023a), whereas mismatched or noisy examples may introduce reasoning biases and degrade performance Zhang et al. (2025c). The importance of selecting useful examples is further enhanced in recent agentic LLM systems, where reasoning experiences are stored as memories and retrieved to guide future tasks Ouyang et al. (2025).

Existing studies have improved demonstration selection in general domain Rubin et al. (2022); Li et al. (2023); Ye et al. (2023a); Yang et al. (2023); Qin et al. (2024). However, most retrieval models rely on embedding similarity and the reasoning logic can be obscured by surface semantics.

The key question is how to represent demonstrations for reasoning tasks. Recent reasoning-aware methods construct intermediate representations before retrieval, including generated reasoning paths Qin et al. (2024), reasoning patterns Zhang et al. (2025b), latent reasoning skills Xu et al. (2024), and reasoning graphs Lin et al. (2025). These methods suggest that raw question embeddings are often insufficient, since reasoning-relevant signals can be diluted by topics, entities, and surface semantics. Problem-solving logic (PSL) takes a more symbolic route by representing each problem as a sequence of predefined reasoning operations and matching demonstrations at the operation level Ma et al. (2025). Such operation-based representations reduce the influence of surface wording and provide a more explicit view of the problem-solving process.

Despite these advances, existing representations still struggle to make reasoning logic both comparable and adaptable. Free-form rationales or reasoning paths can express diverse reasoning processes, but their natural-language form often makes the representation noisy and unstable. Symbolic operation sequences, as explored by PSL, make reasoning steps easier to compare, but fixed operation spaces and exact matching rules limit their ability to generalize across task-specific and semantically equivalent reasoning patterns. This motivates a logical alignment framework that represents reasoning explicitly while aligning it semantically.

To this end, we propose SALA, a Semantic-Aware Logical Alignment framework for reasoning-oriented demonstration selection. SALA constructs a task-adaptive operation space by inducing reasoning operations from downstream data, enabling demonstrations to be represented with problem-solving units beyond the predefined operation set. Figure 1 illustrates how SALA induces additional task-specific operations when the predefined operation set cannot fully express the reasoning logic of a downstream question. Then we embed operation descriptions into a continuous semantic space and applies dynamic time warping (DTW) to align reasoning-operation sequences Sakoe and Chiba (1978). In this way, we keep the reasoning representation explicit while enabling soft semantic alignment beyond exact symbolic correspondence.

We conduct experiments on four reasoning benchmarks across three LLMs. SALA achieves better average performance against similarity-based, learning-based, and reasoning-aware demonstration selection baselines. Ablation studies further show that both task-adaptive operation construction and semantic sequence alignment contribute to the improvement. The key contributions of this work are as follows:

  • •

    We propose SALA, a reasoning-oriented demonstration selection framework that represents problem-solving logic with explicit operation sequences and aligns them in semantic space.

  • •

    We introduce a task-adaptive operation space that extends predefined operations with reusable reasoning units induced from downstream data, enabling more expressive problem-solving representations without manual operation design.

  • •

    We design a semantic DTW-based alignment strategy to measure logical similarity across reasoning-operation sequences with different lengths and decomposition granularities.

  • •

    We validate SALA on four reasoning benchmarks and three LLMs, showing consistent average gains over recent state-of-the-art baselines and confirming the roles of operation construction and semantic alignment through ablations.

Refer to caption
Figure 2: Overview of the SALA framework. SALA constructs a task-adaptive operation set, parses queries and demonstrations into reasoning-operation sequences, and applies DTW-based semantic alignment to retrieve logically similar demonstrations for ICL prompting.

2 Related Work

Existing work on ICL demonstration selection Dong et al. (2024) has studied this problem from lexical, semantic, learned, diversity-aware, and reasoning-aware perspectives. We review the most relevant lines of work in general and reasoning domains.

2.1 General Demonstration Selection

Early demonstration selection methods retrieve examples according to lexical overlap or sentence-level semantic similarity. Sparse retrieval methods such as BM25 Robertson and Zaragoza (2009) estimate query-example word overlap, while embedding-based methods use pretrained encoders such as BERT Devlin et al. (2019) to retrieve semantically similar demonstrations. Later studies improve this basic retrieval paradigm by considering supportiveness, diversity, and compositionality. For example, support-example selection chooses demonstrations according to their usefulness for the target query Li and Qiu (2023). DPP-based selection balances relevance and diversity Yang et al. (2023), and compositional exemplar selection models interactions among demonstrations beyond nearest-neighbor retrieval Ye et al. (2023a). Iterative selection further shows that demonstration retrieval can be treated as a multi-step process rather than a one-shot nearest-neighbor search Qin et al. (2024).

Another line of work formulates demonstration selection as a learnable retrieval or optimization problem. EPR trains a retriever from language-model feedback Rubin et al. (2022), while UDR learns a unified retriever for cross-task demonstration selection Li et al. (2023). More recent methods use model preference, gradient matching, reinforcement learning, or many-shot optimization to improve selection quality Zhang et al. (2025c); Zhang et al. (2025a); Wang et al. (2025); Purohit et al. (2025). These methods improve the retrieval of useful demonstrations, but their selection signals are still mostly derived from input similarity or model-level feedback rather than the reasoning process itself. For complex reasoning tasks, a helpful demonstration should not only be topically or semantically related to the query, but also follow a compatible problem-solving logic.

2.2 Reasoning Demonstration Selection

Recent work has started to construct reasoning-oriented representations. One direction uses natural-language reasoning descriptions. Skill-KNN rewrites inputs into skill-based descriptions before applying embedding-based retrieval An et al. (2023). Luo et al. Luo et al. (2023) extend retrieval-based ICL to chain-of-thought (CoT) prompting, and iterative demonstration selection uses generated reasoning paths to guide example retrieval Qin et al. (2024). These methods keep the representation flexible, but the reasoning signal is still expressed in free-form text, which can make the retrieval noisy or unstable.

A second direction introduces more structured reasoning representations. Reasoning-pattern methods select demonstrations according to task-specific reasoning patterns Zhang et al. (2025b), LaRS learns latent reasoning skills from CoT rationales Xu et al. (2024), and RGER represents intermediate reasoning steps as reasoning graphs for exemplar retrieval Lin et al. (2025). For mathematical reasoning, LMS3 further shows that useful demonstrations should balance semantic similarity and inference stability Liu et al. (2025). These methods provide stronger reasoning-oriented signals than raw question embeddings, but latent skills are less directly interpretable, and graph-based representations require more complex structure matching across examples.

PSL-guided ICL is the most closely related symbolic approach Ma et al. (2025). It represents each problem as a sequence of predefined QDMR-style reasoning operations Wolfson et al. (2020) and selects demonstrations through operation-level exact matching. While this shows the value of explicit operation sequences for representing problem-solving logic, it still relies on a fixed operation space and rigid symbolic matching. SALA moves beyond symbolic matching by combining operation-level reasoning representations with semantic alignment. It constructs a task-adaptive operation space from downstream data and aligns operation sequences in semantic space with DTW, allowing explicit reasoning representations to be compared flexibly across diverse reasoning processes.

3 Methodology

As shown in Figure 2, SALA follows a four-stage pipeline for reasoning-oriented demonstration selection, with adaptive operation construction and DTW-based semantic alignment serving as the two key mechanisms.

3.1 Problem Formulation

Let 𝒟={(qi,yi)}i=1N\mathcal{D}=\{(q_{i},y_{i})\}_{i=1}^{N} denote the demonstration pool derived from the training set of a downstream task, where qiq_{i} and yiy_{i} are the question and its answer, respectively. Given a test query q∗q^{\ast}, the goal is to select kk demonstrations from 𝒟\mathcal{D} that support the target LLM in solving q∗q^{\ast}.

For reasoning tasks, we model the problem-solving logic of a question as an ordered sequence of reasoning operations. Given an operation set 𝒪\mathcal{O}, a parser maps a question qq into

π⁡(q,𝒪)=[o1,…,om],oj∈𝒪.\pi(q;\mathcal{O})=[o_{1},\ldots,o_{m}],\quad o_{j}\in\mathcal{O}. (1)

After constructing the task-adaptive operation set 𝒪t​a​s​k\mathcal{O}_{task}, we denote the query sequence and the candidate sequence as Q∗=π⁡(q∗,𝒪t​a​s​k)Q^{\ast}=\pi(q^{\ast};\mathcal{O}_{task}) and Ei=π⁡(qi,𝒪t​a​s​k)E_{i}=\pi(q_{i};\mathcal{O}_{task}), respectively. Demonstration selection is then formulated as:

𝒟k​(q∗)=TopK(qi,yi)∈𝒟⁡S​i​m​(Q∗,Ei),\mathcal{D}_{k}(q^{\ast})=\operatorname{TopK}_{(q_{i},y_{i})\in\mathcal{D}}Sim(Q^{\ast},E_{i}), (2)

where S​i​m​(Q∗,Ei)Sim(Q^{\ast},E_{i}) measures the alignment between the problem-solving logic of the query and that of the candidate demonstration. SALA focuses on constructing the task-adaptive operation set 𝒪t​a​s​k\mathcal{O}_{task} and defining S​i​m​(⋅,⋅)Sim(\cdot,\cdot) through semantic DTW alignment.

3.2 Adaptive Construction of Reasoning Operation Sets

To address the limited transferability of fixed predefined reasoning operation sets, we construct a task-adaptive reasoning operation set 𝒪t​a​s​k\mathcal{O}_{task} by extending the 13 predefined QDMR reasoning operations 𝒪p​r​e\mathcal{O}_{pre} Wolfson et al. (2020) (see Appendix B for the full list) with operations induced from downstream training data. The construction process consists of candidate operation induction followed by two-stage deduplication.

3.2.1 Candidate Operation Induction

For each training question qiq_{i} in the candidate demonstration pool 𝒟\mathcal{D}, we condition an LLM on the predefined QDMR operation set 𝒪p​r​e\mathcal{O}_{pre} and prompt it to identify additional operations only when the existing set is insufficient to cover the problem-solving logic of qiq_{i}. If no extension is needed, the model returns an empty result; otherwise, it outputs one or more candidate operation descriptions in the predefined format. Collecting the non-empty outputs over all questions yields a candidate operation pool 𝒪c​a​n​d={o1,o2,…,oM}\mathcal{O}_{cand}=\{o_{1},o_{2},\ldots,o_{M}\}, where MM is the number of candidate new operations. The full prompt used for this step is provided in Appendix C.1.

3.2.2 Two-Stage Deduplication of New Operations

To avoid redundancy and ensure that newly induced operations truly complement the predefined ones, we apply a two-stage deduplication procedure.

Stage 1: Heuristic Name-Based Deduplication. We first perform a lightweight heuristic deduplication over the names of candidate operations. For each candidate operation ok∈𝒪c​a​n​do_{k}\in\mathcal{O}_{cand}, we extract its operation name, denoted as nkn_{k}. We then normalize the name by lowercasing it and removing spaces and underscore characters, yielding a normalized form n~k\tilde{n}_{k}. Candidate entries whose extracted names contain obvious formatting artifacts or malformed markers are discarded at this stage. Let 𝒪r​e​p\mathcal{O}_{rep} denote the retained representative operation set, initialized as an empty set. We process the candidates sequentially and compare the normalized name of the current candidate with the normalized operation names already stored in 𝒪r​e​p\mathcal{O}_{rep}.

For a current candidate oko_{k} with normalized name n~k\tilde{n}_{k}, and an existing representative operation o¯i∈𝒪r​e​p\bar{o}_{i}\in\mathcal{O}_{rep} with normalized name n~o¯i\tilde{n}_{\bar{o}_{i}}, we apply the following deterministic rules:

  • •

    If n~k⊆n~o¯i\tilde{n}_{k}\subseteq\tilde{n}_{\bar{o}_{i}}, we replace o¯i\bar{o}_{i} with oko_{k}.

  • •

    If n~o¯i⊆n~k\tilde{n}_{\bar{o}_{i}}\subseteq\tilde{n}_{k}, we discard oko_{k}.

  • •

    Otherwise, we keep both entries unchanged.

If no containment relation is found between n~k\tilde{n}_{k} and any existing representative in 𝒪r​e​p\mathcal{O}_{rep}, we append oko_{k} to 𝒪r​e​p\mathcal{O}_{rep}. After all candidates are processed, 𝒪r​e​p={o¯1,o¯2,…,o¯P}\mathcal{O}_{rep}=\{\bar{o}_{1},\bar{o}_{2},\ldots,\bar{o}_{P}\} becomes the representative operation set produced by Stage 1. This heuristic removes near-duplicate operation names caused by minor formatting differences or partially overlapping name variants, while preserving the original full operation descriptions for subsequent processing. Because the procedure depends only on normalized names, containment checks, and candidate order, it is reproducible given the same candidate operation pool.

Stage 2: Function Duplication Detection Based on LLM Judgment. When 𝒪r​e​p={o¯1,o¯2,…,o¯P}\mathcal{O}_{rep}=\{\bar{o}_{1},\bar{o}_{2},\ldots,\bar{o}_{P}\} is obtained after Stage 1, we further remove functional redundancy through a sequential judgment process. We initialize the task-adaptive operation set as 𝒪t​a​s​k(0)=𝒪p​r​e\mathcal{O}_{task}^{(0)}=\mathcal{O}_{pre}. For each representative operation o¯i∈𝒪r​e​p\bar{o}_{i}\in\mathcal{O}_{rep}, we construct a function comparison prompt that includes the core function description of o¯i\bar{o}_{i} and the core function descriptions of all operations in the current operation set 𝒪t​a​s​k(i−1)\mathcal{O}_{task}^{(i-1)}. The LLM outputs a binary judgment result J⁡(o¯i,𝒪t​a​s​k(i−1))∈{0,1}J(\bar{o}_{i},\mathcal{O}_{task}^{(i-1)})\in\{0,1\}, where J⁡(o¯i,𝒪t​a​s​k(i−1))=1J(\bar{o}_{i},\mathcal{O}_{task}^{(i-1)})=1 means that o¯i\bar{o}_{i} functionally overlaps with at least one operation in 𝒪t​a​s​k(i−1)\mathcal{O}_{task}^{(i-1)}, and J⁡(o¯i,𝒪t​a​s​k(i−1))=0J(\bar{o}_{i},\mathcal{O}_{task}^{(i-1)})=0 means no overlap.

The operation set is then updated recursively:

𝒪t​a​s​k(i)={𝒪t​a​s​k(i−1)∪{o¯i},if ​J​(o¯i,𝒪t​a​s​k(i−1))=0,𝒪t​a​s​k(i−1),otherwise.\mathcal{O}_{task}^{(i)}=\begin{cases}\mathcal{O}_{task}^{(i-1)}\cup\{\bar{o}_{i}\},&\text{if }J(\bar{o}_{i},\mathcal{O}_{task}^{(i-1)})=0,\\ \mathcal{O}_{task}^{(i-1)},&\text{otherwise}.\end{cases} (3)

After all representative operations are examined, the final task-adaptive reasoning operation set is 𝒪t​a​s​k=𝒪t​a​s​k(P)\mathcal{O}_{task}=\mathcal{O}_{task}^{(P)}.

This sequential process ensures that each accepted operation complements both the predefined operations and the previously accepted induced operations. The prompts for candidate induction (Section C.1) and deduplication (Section C.3) are fully documented in Appendix C.

3.3 Construction of the Reasoning Operation Embedding Library

To capture the functional semantic associations of operations and support semantic alignment, we convert discrete operations into continuous semantic embeddings and construct a reasoning operation embedding library, providing the foundation for subsequent similarity calculations.

First, we generate semantic description texts for each operation o∈𝒪t​a​s​ko\in\mathcal{O}_{task}, denoted as d​e​s​c​(o)desc(o). In our implementation, each d​e​s​c​(o)desc(o) describes the core function of the corresponding operation, and the description texts are used as inputs to a pretrained embedding model.

Next, we use a pretrained text embedding model as the semantic encoder, denoted as fe​m​b​(⋅):s​t​r→ℝdf_{emb}(\cdot):str\rightarrow\mathbb{R}^{d}, where dd is the embedding dimension. Specifically, the implementation uses a Hugging Face embedding model to encode the operation description text and obtains the initial semantic embedding vector through mean pooling over the hidden states:

uoi​n​i​t=fe​m​b​(d​e​s​c​(o))∈ℝdu_{o}^{init}=f_{emb}(desc(o))\in\mathbb{R}^{d} (4)

We apply ℓ2\ell_{2} normalization to the initial vector before DTW matching:

vo=uoi​n​i​t‖uoi​n​i​t‖2v_{o}=\frac{u_{o}^{init}}{\|u_{o}^{init}\|_{2}} (5)

where vo∈ℝdv_{o}\in\mathbb{R}^{d} and ‖vo‖2=1\|v_{o}\|_{2}=1 is the final semantic embedding vector of operation oo.

Finally, the mapping relationship, {(o,vo)∣o∈𝒪t​a​s​k}\{(o,v_{o})\mid o\in\mathcal{O}_{task}\}, is used to form the reasoning operation embedding library Le​m​bL_{emb}. In practice, the embeddings are precomputed once for all operations in 𝒪t​a​s​k\mathcal{O}_{task} and then reused during retrieval.

3.4 DTW-Based Semantic Alignment

In this section, we propose a semantic alignment strategy that maps reasoning operation sequences to semantic embedding sequences and uses DTW to calculate sequence-level similarity Sakoe and Chiba (1978), enabling soft logical alignment across sequences of different lengths.

3.4.1 Reasoning Operation Sequence Parsing

Based on the constructed task-adaptive reasoning operation set 𝒪t​a​s​k\mathcal{O}_{task}, we use an LLM to map each query and candidate demonstration into an ordered reasoning operation sequence. Each sequence represents the problem-solving process as a series of operations drawn from 𝒪t​a​s​k\mathcal{O}_{task}. The parsing prompt is provided in Appendix C.2.

3.4.2 Semantic Sequence Generation

We retrieve the semantic embedding of each reasoning operation in the parsed reasoning operation sequence from the embedding library, converting discrete reasoning operation sequences into continuous semantic embedding sequences. For the query sequence Q∗=[o1∗,o2∗,…,om∗]Q^{\ast}=[o^{\ast}_{1},o^{\ast}_{2},\ldots,o^{\ast}_{m}], its operation embedding sequence is 𝐐∗=[vo1∗,vo2∗,…,vom∗]\mathbf{Q}^{\ast}=[v_{o^{\ast}_{1}},v_{o^{\ast}_{2}},\ldots,v_{o^{\ast}_{m}}], where voi∗v_{o^{\ast}_{i}} is the semantic embedding of the ii-th query operation oi∗o^{\ast}_{i}. For a candidate demonstration sequence EiE_{i}, we simplify it to E=[o1E,o2E,…,onE]E=[o^{E}_{1},o^{E}_{2},\ldots,o^{E}_{n}] when computing DTW, and its semantic sequence is 𝐄=[vo1E,vo2E,…,vonE]\mathbf{E}=[v_{o^{E}_{1}},v_{o^{E}_{2}},\ldots,v_{o^{E}_{n}}], where vojEv_{o^{E}_{j}} is the semantic embedding of the jj-th candidate operation ojEo^{E}_{j}.

3.4.3 DTW-based Similarity Calculation

The DTW algorithm calculates the semantic similarity between the candidate semantic sequence 𝐄\mathbf{E} and the query semantic sequence 𝐐∗\mathbf{Q}^{\ast} through three steps. First, we construct a pairwise distance matrix D∈ℝm×nD\in\mathbb{R}^{m\times n}, where the element D⁡(i,j)D(i,j) represents the semantic distance between the ii-th vector voi∗v_{o^{\ast}_{i}} in 𝐐∗\mathbf{Q}^{\ast} and the jj-th vector vojEv_{o^{E}_{j}} in 𝐄\mathbf{E}. Consistent with the implementation, we use Euclidean distance to measure this cost:

D⁡(i,j)=‖voi∗−vojE‖2.D(i,j)=\|v_{o^{\ast}_{i}}-v_{o^{E}_{j}}\|_{2}. (6)

where ∥⋅∥2\|\cdot\|_{2} denotes the Euclidean norm. Since the operation embeddings are precomputed in a fixed dd-dimensional space, this distance directly measures the semantic discrepancy between two reasoning operations.

Next, we search for the optimal warping path. The warping path W=[w1,w2,…,wk]W=[w_{1},w_{2},\ldots,w_{k}] is a sequence of coordinate pairs (i,j)(i,j) that satisfies three constraints to ensure a valid alignment: the boundary constraint (starting at (1,1)(1,1) and ending at (m,n)(m,n) to cover the entire problem-solving logic), the monotonicity constraint (avoiding reverse matching to maintain consistency with the reasoning process), and the continuity constraint (ensuring continuous paths without jumps to avoid missing key reasoning steps). The optimal warping path is the one with the minimum cumulative cost among all feasible paths, where the cumulative cost γ⁡(i,j)\gamma(i,j) from the starting point (1,1)(1,1) to the point (i,j)(i,j) is calculated recursively as:

γ⁡(i,j)=D⁡(i,j)+min⁡{γ⁡(i−1,j),γ⁡(i,j−1),γ⁡(i−1,j−1)}.\gamma(i,j)=D(i,j)+\\ \min\{\gamma(i-1,j),\gamma(i,j-1),\gamma(i-1,j-1)\}. (7)

with the initial condition γ⁡(1,1)=D⁡(1,1)\gamma(1,1)=D(1,1).

Finally, we calculate the sequence similarity. To reduce the impact of sequence length differences, we divide the cumulative DTW distance by the length of the optimal alignment path to obtain the average alignment distance, and then transform it into a bounded similarity score:

S​i​m​(Q∗,Ei)=11+min⁡(γ⁡(m,n)k,1).Sim(Q^{\ast},E_{i})=\frac{1}{1+\min\left(\frac{\gamma(m,n)}{k},1\right)}. (8)

where kk denotes the length of the optimal warping path. Here, γ⁡(m,n)/k\gamma(m,n)/k is the average DTW alignment distance per step. Since operation embeddings are ℓ2\ell_{2}-normalized unit vectors, their pairwise Euclidean distance lies in [0,2][0,2], and the min⁡(⋅,1)\min(\cdot,1) clipping bounds the effective distance to [0,1][0,1]. Therefore, S​i​m​(Q∗,Ei)∈[0.5,1]Sim(Q^{\ast},E_{i})\in[0.5,1], and a value closer to 1 indicates higher semantic similarity and greater consistency in problem-solving logic between the demonstration and the query.

We rank all demonstrations by DTW semantic similarity in descending order and select the top-kk most similar ones. Then, following an easy-to-hard curriculum Ma et al. (2025), we sort the selected demonstrations in ascending order of reasoning operation sequence length.

4 Experiments

4.1 Experimental Setup

4.1.1 Datasets

We evaluate SALA on four reasoning benchmarks covering diverse task types and difficulty levels. SVAMP Patel et al. (2021) contains 1,000 arithmetic word problems. GSM8K Cobbe et al. (2021) consists of 1,319 high-quality grade-school math problems that typically require multi-step reasoning. CommonsenseQA Talmor et al. (2019) is a multiple-choice commonsense reasoning benchmark with 10,881 questions, each associated with five answer options. StrategyQA Geva et al. (2021) is an open-domain question-answering benchmark with 2,290 questions. Detailed dataset statistics and split information are provided in Appendix A.

4.1.2 Baselines

We compare SALA with seven representative ICL baselines, including Random sampling, EPR with a contrastive retriever Rubin et al. (2022), BM25 retrieval Robertson and Zaragoza (2009), TopK-BERT retrieval Devlin et al. (2019), and DPP-BERT, which combines BERT-based retrieval with a Determinantal Point Process to balance relevance and diversity Chen et al. (2018); Yang et al. (2023). We also include two reasoning-aware baselines: PSL, which selects demonstrations using problem-solving logic guidance Ma et al. (2025), and LMS3, which selects demonstrations based on semantic similarity and inference stability Liu et al. (2025).

4.1.3 Implementation Details

We evaluate SALA on three LLMs: Llama3-8B-Instruct Grattafiori et al. (2024), Qwen2.5-7B-Instruct Yang et al. (2024), and DeepSeek-V4-Pro DeepSeek-AI (2026) (accessed via API). For each benchmark, demonstrations are retrieved from the training set. To induce task-specific reasoning operations, we use Seed-OSS-36B-Instruct ByteDance Seed Team (2025) on the training set of each benchmark. Using a separate LLM keeps operation construction independent of the evaluated target LLMs and ensures a fair comparison under the same operation space. For semantic matching, we use bert-base-uncased Devlin et al. (2019) to encode reasoning operations. All experimental results are averaged over five runs. The number of selected examples for all baselines is based on the settings in PSL Ma et al. (2025) for different benchmarks. Detailed configurations can be found in Appendix D.

Method GSM8K SVAMP CMSQA StrQA Avg
Llama3-8B-Instruct
Random 80.91 85.40 72.91 82.27 80.37
EPR 79.45 84.33 69.94 83.99 79.43
BM25 79.88 85.53 69.26 83.49 79.54
TopK-BERT 80.97 85.00 71.25 84.13 80.34
DPP-BERT 79.65 84.20 70.81 82.91 79.39
PSL 82.03 84.67 72.89 84.28 80.97
LMS3 82.65 85.40 73.89 84.89 81.71
SALA 82.59 87.13 74.07 87.25 82.76
Qwen2.5-7B-Instruct
Random 89.25 82.73 71.84 83.29 81.78
EPR 87.95 88.00 74.12 81.95 83.01
BM25 87.08 86.53 73.84 81.74 82.28
TopK-BERT 87.26 85.33 75.18 82.97 82.69
DPP-BERT 88.07 83.86 73.59 83.05 82.14
PSL 89.76 89.33 73.05 80.20 83.09
LMS3 91.52 83.40 73.23 81.98 82.53
SALA 91.16 91.13 74.73 83.90 85.23
DeepSeek-V4-Pro
Random 94.54 90.33 74.45 90.25 87.39
BM25 93.87 92.73 76.59 91.59 88.70
TopK-BERT 96.94 92.20 77.61 92.37 89.78
DPP-BERT 95.45 93.00 75.67 92.58 89.18
PSL 96.48 91.86 76.71 93.68 89.68
SALA 97.12 94.67 79.28 94.61 91.42
Table 1: Main results across three LLMs. EPR and LMS3 are unavailable on DeepSeek-V4-Pro (API-only access). CMSQA denotes CommonsenseQA and StrQA denotes StrategyQA. Bold indicates the best result per model–benchmark.
Variant GSM8K SVAMP CMSQA StrQA Avg
Llama3-8B-Instruct
w/o DTW+OPs 81.65 85.00 72.40 83.41 80.62
w/o OPs 82.20 87.06 73.73 86.20 82.30
w/o DTW 82.61 86.87 74.01 85.09 82.15
Full 82.59 87.13 74.07 87.25 82.76
Qwen2.5-7B-Instruct
w/o DTW+OPs 88.63 86.67 73.55 80.49 82.34
w/o OPs 90.54 87.73 74.22 82.27 83.69
w/o DTW 89.20 90.20 73.02 81.16 83.40
Full 91.16 91.13 74.73 83.90 85.23
Table 2: Ablation study results. “w/o DTW” uses exact prefix matching similar to PSL; “w/o OPs” uses only the 13 predefined QDMR operations. CMSQA denotes CommonsenseQA and StrQA denotes StrategyQA.

4.2 Main Results

Table 1 summarizes the performance of SALA and the baseline methods. On Llama3-8B-Instruct, SALA achieves the best overall average and the top results on SVAMP, CommonsenseQA, and StrategyQA, while remaining competitive on GSM8K. On Qwen2.5-7B-Instruct, SALA again yields the best average, with the strongest results on SVAMP and StrategyQA. On DeepSeek-V4-Pro DeepSeek-AI (2026), EPR and LMS3 are omitted as they require model training or internal hidden states inaccessible through API calls. SALA achieves the best results across all four benchmarks with an average of 91.42%. These results show that SALA consistently improves ICL performance across models of different scales and architectures.

4.3 Ablation Study

To further demonstrate the effectiveness of the SALA framework, we conduct an ablation study on its key modules.

Figure 3: The percentage of samples using task-specific reasoning operations in the training and test sets.
Figure 4: Average reasoning operation sequence length with and without task-specific operations across four benchmarks. (a) Average reasoning sequence length of training set samples; (b) Average reasoning sequence length of test set samples.

As shown in Table 2, removing either component leads to a consistent performance drop on both models, indicating that both task-adaptive operation construction and DTW-based semantic alignment are necessary for SALA. When the induced operations are removed, SALA falls back to the 13 predefined QDMR operations. This reduces the average accuracy by 0.5 points on Llama3-8B-Instruct and 1.5 points on Qwen2.5-7B-Instruct, showing that the additional operations help represent task-specific reasoning patterns that are not sufficiently covered by the predefined operation set. When DTW is removed and exact prefix matching is used instead, the average accuracy drops by 0.6 points on Llama3-8B-Instruct and 1.8 points on Qwen2.5-7B-Instruct. This suggests that semantic sequence alignment is important for handling semantically related operations and different decomposition granularities. Removing both components causes the largest degradation, with average drops of 2.1 and 2.9 points on the two models, respectively.

4.4 SALA Analysis

4.4.1 Necessity of Inducing Task-Specific Operations

To examine the role of task-specific reasoning operations, we measure the proportion of samples whose reasoning sequences require such operations in both the training and test sets. As shown in Figure 3, this proportion exceeds 68% across all datasets. In GSM8K, it is above 98% for both the training and test sets, whereas for the remaining datasets it generally falls between 70% and 80%. These findings suggest that task-specific reasoning operations are widely involved in reasoning decomposition across diverse tasks. Some representative induced operations are listed in Appendix E.

4.4.2 Qualitative Case Analysis

To further understand how each module contributes to demonstration selection, we present four case studies. Cases 1–2 isolate the effect of task-specific operations (Module 1), while Cases 3–4 isolate DTW semantic alignment (Module 2).

Cases 1–2: Task-Specific Operations Condense Reasoning Sequences

Figure 4 shows that adding task-specific operations consistently reduces average reasoning sequence length across all benchmarks. This compaction is critical for accurate demonstration matching. In Case 1 (Table 7 in Appendix F.1), with Modulo, the query sequence shrinks from 9 to 7 operations, and the matched demonstration is another remainder problem; without Modulo, the inflated division-multiplication-subtraction sequence attracts a proportion problem with an incompatible reasoning pattern. In Case 2 (Table 8 in Appendix F.1), with Algebra, the sequence shrinks from 13 to 7 operations; the compact Define–Algebra pattern matches another equation-solving demonstration, while the 13-step QDMR-only decomposition attracts a multi-item subtraction problem.

Cases 3–4: DTW Overcomes Limitations of Prefix Matching

Cases 3–4 use only the 13 QDMR operations to isolate Module 2. Case 3 (Table 9 in Appendix F.2) shows that DTW selects a percentage-markup demonstration that matches the query’s reasoning pattern despite different third operations, while the longest valid prefix demonstration shares 6 of 7 operations but encodes a subtract-discount pattern incompatible with the query’s add-insurance requirement. Case 4 (Table 10 in Appendix F.2) shows that the longest valid prefix demonstration covers 5 of 7 operations but lacks any percentage step, causing the model to miss the tax computation entirely; DTW instead selects a semantically aligned percentage demonstration. Full case details with query examples and demonstration sequences are provided in Appendix F.

5 Conclusion and Future Work

We propose SALA, a reasoning-oriented demonstration selection framework that combines task-adaptive operation construction with semantic DTW-based alignment. SALA represents problem-solving logic as explicit reasoning-operation sequences, while allowing flexible matching in semantic space. Experiments on four reasoning benchmarks and three LLMs show that SALA consistently outperforms strong demonstration selection baselines. In future work, we plan to extend SALA to broader reasoning domains and explore richer reasoning structures beyond linear operation sequences.

Limitations

SALA uses LLMs to induce task-adaptive operations and parse questions into reasoning-operation sequences. While this reduces the need for manual operation engineering, the induced operation set and parsed sequences may vary with the induction model and prompting strategy. In this work, we use a separate LLM and fixed prompts to keep the construction process consistent across experiments.

SALA currently represents problem-solving logic as a linear sequence of operations. This representation works well for the reasoning benchmarks studied in this paper, but more complex tasks may benefit from richer structures, such as hierarchical or graph-based reasoning representations. We leave the extension of SALA to these settings for future work.

Acknowledgements

This work was supported by the National Natural Science Foundation of China (62306344, 62276279), the Guangdong Basic and Applied Basic Research Foundation (2024A1515010253, 2026A1515011800, 2024B1515020032), the Open Research Fund of the State Key Laboratory of Blockchain and Data Security, Zhejiang University (Grant No. A2537), and the Guangdong S&T Programme Key-Area Research and Development Program of Guangdong Province (2026B0101100004).

Generative AI tools were used during the preparation and revision of this manuscript to polish the language, revise selected passages, check the consistency of terminology and reported numerical values across sections, and assist with formatting. All AI-assisted changes incorporated into the manuscript were reviewed and verified by the authors, who take full responsibility for the paper’s content.

References

  • Achiam et al. (2023) J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774. Cited by: §1.
  • An et al. (2023) S. An, B. Zhou, Z. Lin, Q. Fu, B. Chen, N. Zheng, W. Chen, and J. Lou Skill-based few-shot selection for in-context learning. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Singapore, pp. 13472–13492. External Links: Document, Link Cited by: §2.2.
  • Bahrini et al. (2023) A. Bahrini, M. Khamoshifar, H. Abbasimehr, R. J. Riggs, M. Esmaeili, R. M. Majdabadkohne, and M. Pasehvar ChatGPT: applications, opportunities, and threats. In 2023 Systems and Information Engineering Design Symposium (SIEDS), Vol. , pp. 274–279. External Links: Document Cited by: §1.
  • Bai et al. (2023) J. Bai, S. Bai, Y. Chu, Z. Cui, K. Dang, X. Deng, Y. Fan, W. Ge, Y. Han, F. Huang, et al. Qwen technical report. arXiv preprint arXiv:2309.16609. Cited by: §1.
  • Bhattamishra et al. (2024) S. Bhattamishra, A. Patel, P. Blunsom, and V. Kanade Understanding in-context learning in transformers and LLMs by learning to learn discrete functions. In The Twelfth International Conference on Learning Representations, External Links: Link Cited by: §1.
  • ByteDance Seed Team (2025) ByteDance Seed Team Seed-oss open-source models. Note: https://github.com/ByteDance-Seed/seed-oss Cited by: §4.1.3.
  • Chen et al. (2018) L. Chen, G. Zhang, and E. Zhou Fast greedy map inference for determinantal point process to improve recommendation diversity. In Advances in Neural Information Processing Systems, S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett (Eds.), Vol. 31, pp. . External Links: Link Cited by: §4.1.2.
  • Cobbe et al. (2021) K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168. Cited by: §4.1.1.
  • DeepSeek-AI (2026) DeepSeek-AI DeepSeek-v4: towards highly efficient million-token context intelligence. Cited by: §4.1.3, §4.2.
  • Devlin et al. (2019) J. Devlin, M. Chang, K. Lee, and K. Toutanova BERT: pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), J. Burstein, C. Doran, and T. Solorio (Eds.), Minneapolis, Minnesota, pp. 4171–4186. External Links: Link, Document Cited by: §2.1, §4.1.2, §4.1.3.
  • Dong et al. (2024) Q. Dong, L. Li, D. Dai, C. Zheng, J. Ma, R. Li, H. Xia, J. Xu, Z. Wu, B. Chang, X. Sun, L. Li, and Z. Sui A survey on in-context learning. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Miami, Florida, USA, pp. 1107–1128. External Links: Link, Document Cited by: §1, §2.
  • Geva et al. (2021) M. Geva, D. Khashabi, E. Segal, T. Khot, D. Roth, and J. Berant Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies. Transactions of the Association for Computational Linguistics 9, pp. 346–361. External Links: Link, Document Cited by: §4.1.1.
  • Grattafiori et al. (2024) A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783. Cited by: §4.1.3.
  • Li et al. (2023) X. Li, K. Lv, H. Yan, T. Lin, W. Zhu, Y. Ni, G. Xie, X. Wang, and X. Qiu Unified demonstration retriever for in-context learning. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), A. Rogers, J. Boyd-Graber, and N. Okazaki (Eds.), Toronto, Canada, pp. 4644–4668. External Links: Link, Document Cited by: §1, §2.1.
  • Li and Qiu (2023) X. Li and X. Qiu Finding support examples for in-context learning. In Findings of the Association for Computational Linguistics: EMNLP 2023, H. Bouamor, J. Pino, and K. Bali (Eds.), Singapore, pp. 6219–6235. External Links: Link, Document Cited by: §2.1.
  • Lin et al. (2025) Y. Lin, B. Zhong, S. Jiang, J. Siebert, and Q. Chen Reasoning graph enhanced exemplars retrieval for in-context learning. In Proceedings of the 31st International Conference on Computational Linguistics, Abu Dhabi, UAE, pp. 9737–9759. External Links: Link Cited by: §1, §2.2.
  • Liu et al. (2025) J. Liu, Z. Huang, C. Wang, X. Huang, C. Zhai, and E. Chen What makes in-context learning effective for mathematical reasoning. In Forty-second International Conference on Machine Learning, External Links: Link Cited by: §2.2, §4.1.2.
  • Luo et al. (2023) M. Luo, X. Xu, Z. Dai, P. Pasupat, M. Kazemi, C. Baral, V. Imbrasaite, and V. Y. Zhao Dr.ICL: Demonstration-Retrieved In-Context Learning. In NeurIPS 2023 Workshops: R0-FoMo, External Links: Link Cited by: §2.2.
  • Luo et al. (2024) M. Luo, X. Xu, Y. Liu, P. Pasupat, and M. Kazemi In-context learning with retrieved demonstrations for language models: a survey. Transactions on Machine Learning Research. Note: Survey Certification External Links: ISSN 2835-8856, Link Cited by: §1.
  • Ma et al. (2025) X. Ma, W. Jiang, and H. Huang Problem-solving logic guided curriculum in-context learning for LLMs complex reasoning. In Findings of the Association for Computational Linguistics: ACL 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, pp. 8394–8412. External Links: Link, Document, ISBN 979-8-89176-256-5 Cited by: §1, §2.2, §3.4.3, §4.1.2, §4.1.3.
  • Ouyang et al. (2022) L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. Christiano, J. Leike, and R. Lowe Training language models to follow instructions with human feedback. In Proceedings of the 36th International Conference on Neural Information Processing Systems, NIPS ’22, Red Hook, NY, USA. External Links: ISBN 9781713871088 Cited by: §1.
  • Ouyang et al. (2025) S. Ouyang, J. Yan, I. Hsu, Y. Chen, K. Jiang, Z. Wang, R. Han, L. T. Le, S. Daruki, X. Tang, V. Tirumalashetty, G. Lee, M. Rofouei, H. Lin, J. Han, C. Lee, and T. Pfister ReasoningBank: scaling agent self-evolving with reasoning memory. CoRR abs/2509.25140. External Links: Link, Document, 2509.25140 Cited by: §1.
  • Patel et al. (2021) A. Patel, S. Bhattamishra, and N. Goyal Are NLP models really able to solve simple math word problems?. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, K. Toutanova, A. Rumshisky, L. Zettlemoyer, D. Hakkani-Tur, I. Beltagy, S. Bethard, R. Cotterell, T. Chakraborty, and Y. Zhou (Eds.), Online, pp. 2080–2094. External Links: Link, Document Cited by: §4.1.1.
  • Purohit et al. (2025) K. Purohit, V. V, S. Bhattacharya, and A. Anand Sample efficient demonstration selection for in-context learning. In Proceedings of the 42nd International Conference on Machine Learning, A. Singh, M. Fazel, D. Hsu, S. Lacoste-Julien, F. Berkenkamp, T. Maharaj, K. Wagstaff, and J. Zhu (Eds.), Proceedings of Machine Learning Research, Vol. 267, pp. 49959–49982. External Links: Link Cited by: §2.1.
  • Qin et al. (2024) C. Qin, A. Zhang, C. Chen, A. Dagar, and W. Ye In-context learning with iterative demonstration selection. In Findings of the Association for Computational Linguistics: EMNLP 2024, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Miami, Florida, USA, pp. 7441–7455. External Links: Link, Document Cited by: §1, §1, §2.1, §2.2.
  • Robertson and Zaragoza (2009) S. Robertson and H. Zaragoza The probabilistic relevance framework: bm25 and beyond. Found. Trends Inf. Retr. 3 (4), pp. 333–389. External Links: ISSN 1554-0669, Link, Document Cited by: §2.1, §4.1.2.
  • Rubin et al. (2022) O. Rubin, J. Herzig, and J. Berant Learning to retrieve prompts for in-context learning. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, M. Carpuat, M. de Marneffe, and I. V. Meza Ruiz (Eds.), Seattle, United States, pp. 2655–2671. External Links: Link, Document Cited by: §1, §2.1, §4.1.2.
  • Sakoe and Chiba (1978) H. Sakoe and S. Chiba Dynamic programming algorithm optimization for spoken word recognition. IEEE Transactions on Acoustics, Speech, and Signal Processing 26 (1), pp. 43–49. External Links: Document Cited by: §1, §3.4.
  • Talmor et al. (2019) A. Talmor, J. Herzig, N. Lourie, and J. Berant CommonsenseQA: a question answering challenge targeting commonsense knowledge. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), J. Burstein, C. Doran, and T. Solorio (Eds.), Minneapolis, Minnesota, pp. 4149–4158. External Links: Link, Document Cited by: §4.1.1.
  • Touvron et al. (2023) H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al. Llama 2: open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288. Cited by: §1.
  • Wang et al. (2025) X. Wang, J. Wu, Y. Yuan, D. Cai, M. Li, and W. Jia Demonstration selection for in-context learning via reinforcement learning. In Proceedings of the 42nd International Conference on Machine Learning, ICML’25. Cited by: §2.1.
  • Wies et al. (2023) N. Wies, Y. Levine, and A. Shashua The learnability of in-context learning. In Proceedings of the 37th International Conference on Neural Information Processing Systems, NIPS ’23, Red Hook, NY, USA. Cited by: §1.
  • Wolfson et al. (2020) T. Wolfson, M. Geva, A. Gupta, M. Gardner, Y. Goldberg, D. Deutch, and J. Berant Break it down: a question understanding benchmark. Transactions of the Association for Computational Linguistics 8, pp. 183–198. External Links: Link, Document Cited by: Appendix B, §2.2, §3.2.
  • Xu et al. (2024) Z. Xu, H. Wang, D. Bespalov, X. Wu, P. Stone, and Y. Qi LaRS: latent reasoning skills for chain-of-thought reasoning. In Findings of the Association for Computational Linguistics: EMNLP 2024, Miami, Florida, USA, pp. 3624–3643. External Links: Document, Link Cited by: §1, §2.2.
  • Yang et al. (2024) A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, H. Lin, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Lin, K. Dang, K. Lu, K. Bao, K. Yang, L. Yu, M. Li, M. Xue, P. Zhang, Q. Zhu, R. Men, R. Lin, T. Li, T. Xia, X. Ren, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Wan, Y. Liu, Z. Cui, Z. Zhang, and Z. Qiu Qwen2.5 technical report. arXiv preprint arXiv:2412.15115. Cited by: §4.1.3.
  • Yang et al. (2023) Z. Yang, Y. Zhang, D. Sui, C. Liu, J. Zhao, and K. Liu Representative demonstration selection for in-context learning with two-stage determinantal point process. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, H. Bouamor, J. Pino, and K. Bali (Eds.), Singapore, pp. 5443–5456. External Links: Link, Document Cited by: §1, §2.1, §4.1.2.
  • Ye et al. (2023a) J. Ye, Z. Wu, J. Feng, T. Yu, and L. Kong Compositional exemplars for in-context learning. In Proceedings of the 40th International Conference on Machine Learning, A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, and J. Scarlett (Eds.), Proceedings of Machine Learning Research, Vol. 202, pp. 39818–39833. External Links: Link Cited by: §1, §1, §2.1.
  • Ye et al. (2023b) J. Ye, X. Chen, N. Xu, C. Zu, Z. Shao, S. Liu, Y. Cui, Z. Zhou, C. Gong, Y. Shen, J. Zhou, S. Chen, T. Gui, Q. Zhang, and X. Huang A comprehensive capability analysis of gpt-3 and gpt-3.5 series models. External Links: 2303.10420, Link Cited by: §1.
  • Zhang et al. (2025a) J. Zhang, B. Li, J. Bai, R. Li, Y. Wang, C. Lin, and W. Rong Selecting demonstrations for many-shot in-context learning via gradient matching. In Findings of the Association for Computational Linguistics: ACL 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, pp. 11686–11704. External Links: Link, Document, ISBN 979-8-89176-256-5 Cited by: §2.1.
  • Zhang et al. (2025b) Y. Zhang, X. Wang, L. Wu, and J. Wang Enhancing chain of thought prompting in large language models via reasoning patterns. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39, pp. 25985–25993. External Links: Document Cited by: §1, §2.2.
  • Zhang et al. (2025c) Z. Zhang, S. Lan, L. Song, J. Bian, Y. Li, and K. Ren Learning to select in-context demonstration preferred by large language model. In Findings of the Association for Computational Linguistics: ACL 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, pp. 11345–11360. External Links: Link, Document, ISBN 979-8-89176-256-5 Cited by: §1, §2.1.

Appendix A Dataset Statistics

Table 3 summarizes the training and test splits of the four benchmark datasets used in this paper.

Dataset Train Test
GSM8K 7,473 1,319
SVAMP 700 300
CommonsenseQA 9,741 1,140
StrategyQA 1,603 687
Table 3: Training and test split statistics for the datasets used in our experiments.

Appendix B Predefined QDMR Reasoning Operations

SALA extends the 13 predefined QDMR reasoning operations Wolfson et al. (2020) with task-specific operations. Table 4 lists all 13 predefined operations with their core functions, example questions, and example sequences.

Operation Core Function Example Question Example Sequence
Select Selects a specific entity or set How many touchdowns were scored overall? SELECT[’touchdowns’]; AGGREGATE[’number’, ’#1’]
Filter Selects a subset matching conditions I would like a flight from Toronto to San Diego please. SELECT[’flights’]; FILTER[’#1’, ’from Toronto’]; FILTER[’#2’, ’to San Diego’]
Arithmetic Basic arithmetic (+, −- , ×\times, ÷\div) on numeric attributes How many more red objects are there than blue objects? SELECT[’red objects’]; SELECT[’blue objects’]; PROJECT[’number’, ’#1’]; PROJECT[’number’, ’#2’]; ARITHMETIC[’difference’, ’#3’, ’#4’]
Comparative Selects subset greater/less than a threshold Who are the authors with more than 500 papers? SELECT[’authors’]; PROJECT[’papers’, ’#1’]; GROUP[’number’, ’#2’, ’#1’]; COMPARATIVE[’#1’, ’#3’, ’more than 500’]
Superlative Selects the entity with the extreme value What is the keyword contained by the most papers? SELECT[’papers’]; PROJECT[’keywords’, ’#1’]; GROUP[’number’, ’#1’, ’#2’]; SUPERLATIVE[’#2’, ’#3’, ’highest’]
Aggregate Computes mathematical properties (count, avg) of a set How many states border Colorado? SELECT[’Colorado’]; PROJECT[’border states’, ’#1’]; AGGREGATE[’number’, ’#2’]
Union Merges two sets Tell me the president and vice-president. SELECT[’president’]; SELECT[’vice-president’]; UNION[’#1’, ’#2’]
Intersection Takes the intersection of two sets Show parties with representatives in both New York and Pennsylvania. SELECT[’representatives’]; FILTER[’#1’, ’in New York’]; FILTER[’#1’, ’in Pennsylvania’]; INTERSECTION[’parties’, ’#2’, ’#3’]
Project Obtains a specific attribute of an entity Who is the head coach of the Los Angeles Lakers? SELECT[’Los Angeles Lakers’]; PROJECT[’head coach’, ’#1’]
Sort Arranges elements by a specified rule Find student addresses sorted by monthly rental. SELECT[’students’]; PROJECT[’addresses’, ’#1’]; PROJECT[’monthly rental’, ’#2’]; SORT[’#2’, ’#3’]
Group Computes properties per group element How many female students per club? SELECT[’clubs’]; FILTER[’#1’, ’female students’]; GROUP[’number’, ’#2’, ’#1’]
Discard Selects subset not satisfying a condition Find professors not playing Canoeing. SELECT[’professors’]; FILTER[’#1’, ’playing Canoeing’]; DISCARD[’#1’, ’#2’]
Boolean Determines if entity satisfies a condition Were Scott Derrickson and Ed Wood of the same nationality? SELECT[’Scott Derrickson’]; SELECT[’Ed Wood’]; PROJECT[’nationality’, ’#1’]; PROJECT[’nationality’, ’#2’]; BOOLEAN[’#3’, ’the same as’, ’#4’]
Table 4: The 13 predefined QDMR reasoning operations used as the base inventory in SALA.

Appendix C Prompt Templates

We provide the full prompt templates used in each step of the SALA pipeline. All prompts were used in English during our experiments.

C.1 Operation Induction Prompt

This prompt instructs the LLM to induce reasoning operations not covered by the predefined 13.

Task: Given an input question and the 13 existing reasoning operations (listed below), induce any new reasoning operations needed to solve the problem.

Existing reasoning operations: {operations}

Requirements: Supplement with new operations. Identify any new reasoning operations from the input question that are not covered by the existing ones. Describe the core functionality, provide an example question, example decomposition, and corresponding reasoning operation sequence in the same format as above. Do not output any extraneous content. If you determine that the existing operations are sufficient to solve the input question, simply output “No new operation needed.”

Input question: {problem}

C.2 Sequence Analysis Prompt

This prompt decomposes a question into an operation sequence using the extended operation set.

Decompose the input question into a reasoning operation sequence using the given reasoning operations.

Given reasoning operations: {operations}

Example Question: what flights are available tomorrow from denver to philadelphia? Reasoning operation sequence: [’SELECT[’flights’]’, ’FILTER[’#1’, ’from denver’]’, ’FILTER[’#2’, ’to philadelphia’]’, ’FILTER[’#3’, ’if available’]’] Operation name list: [’select’, ’filter’, ’filter’, ’filter’]

Follow the example format strictly and provide the reasoning operation sequence and corresponding operation name list for the input question below.

Input question: {question}

C.3 LLM Deduplication Prompt

This prompt checks whether a newly induced operation overlaps with existing ones (Stage 2 of deduplication).

Task: Compare whether the functionality of the new reasoning operation overlaps with any existing reasoning operation, and determine whether the new operation is a composite operation.

Existing reasoning operations: {old_batch}

Requirements: If the new operation’s functionality overlaps with any existing operation, or if the new operation can be decomposed into existing operations (i.e., it is composite), output ‘<1>’. Otherwise, output ‘<0>’.

New reasoning operation: {new}

C.4 ICL Prompt Template

This template formats the retrieved demonstrations and query for the final ICL inference.

Before answering my question, there are some examples you can use for reference:

<example>{examples} </example>

Here is my question: {question}

Please answer it:

C.5 Evaluation Prompts

We use three dataset-specific evaluation prompts to compare candidate answers against ground-truth answers.

Numeric Evaluation (used for GSM8K, SVAMP):

Your task is to compare the Candidate Answer with the Correct Answer and determine whether they are consistent.

Question: {question}

Candidate Answer: {candidate}

Correct Answer: {correct}

Criteria: - Numerical values are consistent if they represent the same quantity despite different formats (e.g., 0.88 is consistent with 88%). Return <1>. - Numerical values are also considered consistent if rounding leads to the same result (e.g., if the correct answer is 8 and the candidate is 7.96, return <1>).

Output: Compare whether the answers are consistent. If consistent, output <1>; otherwise, output <0>.

Text Evaluation (used for CommonsenseQA):

Your task is to compare the Candidate Answer with the Correct Answer and determine whether they are consistent.

Question: {question}

Candidate Answer: {candidate}

Correct Answer: {correct}

Criteria: - Evaluate whether the Candidate Answer and Correct Answer share the same meaning in the context of the given question. If they use different words but convey the same meaning, consider them consistent.

Output: If consistent, output <1>; otherwise, output <0>.

Boolean Evaluation (used for StrategyQA):

Your task is to compare the Candidate Answer with the Correct Answer and determine whether they are consistent.

Question: {question}

Candidate Answer: {candidate}

Correct Answer: {correct}

Criteria: - The Candidate Answer may be verbose while the Correct Answer is very concise. If the final conclusion of the Candidate Answer matches the boolean value (True or False) of the Correct Answer, consider them consistent.

Output: If consistent, output <1>; otherwise, output <0>.

Appendix D Hyperparameter Settings

Table 5 lists the key hyperparameters used in our experiments.

Parameter Value
Operation Induction
LLM for induction & deduplication Seed-OSS-36B-Instruct
Temperature 0.01
Max tokens 4,096
Semantic Embedding
Embedding model bert-base-uncased
Embedding dimension 768
Max input length 1,024 tokens
DTW Retrieval
Retrieval top-kk 8
Inference & Evaluation
Target LLMs Llama3-8B-Instruct, Qwen2.5-7B-Instruct, DeepSeek-V4-Pro
Temperature 0.01
Max tokens 4,096
Evaluation LLM GPT-4
Table 5: Key hyperparameter settings for SALA.

Appendix E Induced Reasoning Operation Examples across Downstream Tasks

Table 6 presents representative reasoning operations induced by SALA on each downstream benchmark through the adaptive operation set construction process.

Operation Core Function Example Question Example Sequence
Define Declares an unknown variable representing a target quantity to be solved Randy spent $10 on lunch. He then spent a quarter of his remaining money on an ice cream cone costing $5. What was his initial money? DEFINE[’M’, ’initial money’]; ARITHMETIC[’subtract’, ’#1’, ’10’]; ARITHMETIC[’quarter’, ’#2’]; ALGEBRA[’#3’, ’=’, ’5’]
Algebra Defines unknown variables, builds relations with known values, and solves via equations Yvonne brings a box of chocolates to school. Half have nuts and half do not. The students eat 80% of the ones with nuts and half of the ones without nuts. If there are 28 left, how many were originally in the box? DEFINE[’total chocolates’, ’#x’]; ARITHMETIC[’division’, ’#x’, 2]; ARITHMETIC[’division’, ’#x’, 2]; ARITHMETIC[’multiply’, ’#2’, ’0.2’]; ARITHMETIC[’multiply’, ’#3’, ’0.5’]; ARITHMETIC[’addition’, ’#4’, ’#5’]; ALGEBRA[’#x’, ’#6’, ’=’, ’28’]
Modulo Computes the remainder after dividing one number by another Emma’s bank account has $100. Each day of the week she spends $8. At the end of the week, she withdraws as many $5 bills as possible. How many dollars remain? SELECT[’initial amount’]; SELECT[’daily spending’]; SELECT[’days in a week’]; ARITHMETIC[’product’, ’#2’, ’#3’]; ARITHMETIC[’difference’, ’#1’, ’#4’]; MODULO[’#5’, ’5’]
SVAMP
Constant Obtains a fixed numeric value given directly or implicitly in the problem Mary is baking a cake. The recipe calls for 5 cups of sugar and 14 cups of flour. She already put in 11 cups of flour. How many more cups of sugar than cups of flour does she need to add now? SELECT[’recipe’]; PROJECT[’sugar’, ’#1’]; PROJECT[’flour’, ’#1’]; SELECT[’flour’]; PROJECT[’already added’, ’#1’]; CONSTANT[’0’]; ARITHMETIC[’subtract’, ’#2’, ’#6’]; ARITHMETIC[’subtract’, ’#3’, ’#5’]; ARITHMETIC[’subtract’, ’#7’, ’#8’]
Literal Retrieves a specific numeric or attribute value directly stated in the problem If 479 students suggested adding mashed potatoes while 489 suggested adding bacon, how many more students suggested bacon than mashed potatoes? LITERAL[’mashed potato’, ’479’]; LITERAL[’bacon’, ’489’]; ARITHMETIC[’difference’, ’#1’, ’#2’]
Retain Preserves the original attribute value of an entity when no change to that attribute is described Ed had 12 more marbles than Doug. Ed lost 20 marbles. If Ed now has 17, how many marbles does Doug have now? SELECT[’Edś current marbles’]; SELECT[’Edś lost marbles’]; ARITHMETIC[’addition’, ’#1’, ’#2’]; SELECT[’Ed-Doug gap’]; ARITHMETIC[’subtract’, ’#3’, ’#4’]; RETAIN[’Dougś marbles’, ’#5’]
StrategyQA
EntityLinking Resolves a referring expression to its corresponding real-world entity Was historical Dracula from a town in Bucharest? ENTITYLINKING["historical Dracula"]; PROJECT["birthplace", "#1"]; BOOLEAN["#2", "in", "Bucharest"]
Relate Obtains the value of a specific relationship between two or more entities What is the distance between Dusseldorf and Stonehenge? SELECT[’Dusseldorf’]; SELECT[’Stonehenge’]; RELATE[’distance’, ’#1’, ’#2’]
Deductive Derives a conclusion through logical reasoning from multiple factual premises Aristotle died in 322 BC. The Model Parliament was held in 1295. The House of Lords grew out of the Model Parliament. Was Aristotle a member of the House of Lords? SELECT[’Aristotle’]; PROJECT[’death year’, ’#1’]; SELECT[’Model Parliament’]; PROJECT[’held year’, ’#3’]; SELECT[’House of Lords’]; PROJECT[’origin’, ’#5’]; BOOLEAN[’#2’, ’earlier than’, ’#4’]; DEDUCTIVE[’#1 is a member of #5’, ’#6’, ’#7’]
CommonsenseQA
Contrast Obtains the opposite or contrasting concept of a given concept or attribute The troublemaker had been hoping for a soft punishment, but the ruling handed down was quite what? SELECT[’soft punishment’]; CONTRAST[’opposite’, ’#1’]; SELECT[’the ruling’]; PROJECT[’#2’, ’#3’]
Similar Finds entities or collections that are similar in type or features to a specified entity A hurricane is similar to what other wind event? SELECT[’hurricane’]; SIMILAR[’wind events’, ’#1’]
Source Locates where or how a specific attribute of an entity can be found or accessed If I wanted to find out the hours that Marmot was open, I might look where? SELECT[’Marmot’]; SOURCE[’open hours’, ’#1’]
Table 6: Representative reasoning operations induced by SALA on four downstream benchmarks.

Appendix F Case Study — Detailed Tables

This appendix provides the full detailed tables for the four case studies summarized in Section 4.4 (Table 11). Cases 1–2 isolate Module 1 (task-specific operations) and Cases 3–4 isolate Module 2 (DTW semantic alignment). All experiments use Llama3-8B-Instruct.

F.1 Effect of Task-Specific Operations on Reasoning Sequence Compactness

Module 1 induces task-specific operations that condense reasoning sequences which would otherwise require long chains of QDMR operations. This compactness directly affects which demonstrations are retrieved and whether the model produces the correct answer.

F.1.1 Case 1: Modulo Condenses Remainder Computation

Query
Emma has $100 in her bank account. She spends $8 each day for a full week. At the end of the week, she withdraws as many $5 bills as possible. How many dollars remain in her account?
Ground Truth $4
w/ new OPs $4 ✓  (with Modulo)
w/o new OPs $2 ×\times  (QDMR only)
With New Operations (Modulo)
Query sequence (len=7) [SELECT, PROJECT, SELECT, PROJECT, ARITHMETIC, ARITHMETIC, MODULO]
Matched demo “A factory packs 125 toys per day. Each box holds 8 toys. After filling full boxes, how many toys are left?”
Demo sequence (len=4) [SELECT, PROJECT, ARITHMETIC, MODULO]
Analysis This is another remainder problem; the model correctly performs the modulo computation.
Without New Operations (QDMR Only)
Query sequence (len=9) [SELECT, PROJECT, SELECT, PROJECT, ARITHMETIC, ARITHMETIC, ARITHMETIC, ARITHMETIC, ARITHMETIC]
To simulate modulo, the sequence appends division →\rightarrow multiplication →\rightarrow subtraction.
Matched demo “A recipe needs 3 cups of flour per 5 servings. How many cups for 20 servings?”
Demo sequence (len=7) [SELECT, PROJECT, SELECT, PROJECT, ARITHMETIC, ARITHMETIC, ARITHMETIC]
Analysis This is a proportional reasoning problem, not a remainder problem. The model is misled by the similar tail of consecutive ARITHMETIC steps (5 in the query, 3 in the demo).
Table 7: Case 1: With Modulo, the compact sequence attracts another remainder problem. Without it, the inflated division-multiplication-subtraction chain attracts a proportion problem that shares the same surface pattern but has a different reasoning structure.

F.1.2 Case 2: Algebra Condenses Variable Solving

Query
Yvonne brings a box of chocolates to school. Half have nuts and half do not. The students eat 80% of the ones with nuts and half of the ones without nuts. If there are 28 chocolates left, how many chocolates were originally in the box?
Ground Truth 80
w/ new OPs 80 ✓  (with Algebra)
w/o new OPs 90 ×\times  (QDMR only)
With New Operations (Algebra)
Query sequence (len=7) [DEFINE, ARITHMETIC, ARITHMETIC, ARITHMETIC, ARITHMETIC, ARITHMETIC, ALGEBRA]
Matched demo “Randy spent $10 on lunch and a quarter of the remaining money on an ice cream cone costing $5. What was his initial money?”
Demo sequence (len=5) [DEFINE, ARITHMETIC, ARITHMETIC, ARITHMETIC, ALGEBRA]
Analysis Both queries define an unknown variable and solve for it via an equation. The model recovers the algebraic structure and correctly outputs 80.
Without New Operations (QDMR Only)
Query sequence (len=13) [SELECT, SELECT, PROJECT, PROJECT, ARITHMETIC, ARITHMETIC, PROJECT, PROJECT, ARITHMETIC, ARITHMETIC, ARITHMETIC, ARITHMETIC, ARITHMETIC]
Without DEFINE and ALGEBRA, each chocolate category must be separately selected and projected; remaining amounts are computed stepwise.
Matched demo “Danny has 94 guppies, 76 angelfish, 89 tiger sharks, and 58 Oscar fish. If he sells 30, 48, 17, and 24 respectively, how many fish remain?”
Demo sequence [SELECT, PROJECT, SELECT, PROJECT, SELECT, PROJECT, SELECT, PROJECT, ARITHMETIC, …]
Analysis This is a multi-item subtraction problem whose SELECT–PROJECT repetition is structurally similar to the inflated query, but its reasoning logic (per-item subtraction then aggregation) is unrelated to variable solving.
Table 8: Case 2: With Algebra, the compact DEFINE–…–ALGEBRA pattern matches equation-solving demonstrations. QDMR-only decomposition produces a long chain of SELECT–PROJECT pairs that attracts structurally similar but logically unrelated multi-item problems.

F.2 DTW Semantic Alignment vs. Prefix Subsequence Matching

To isolate Module 2, both methods in this subsection use only the 13 QDMR operations.

F.2.1 Case 3: Granularity Mismatch — Best Structural Match Excluded by Prefix Constraint

Query
Janet buys a brooch for her daughter. She pays $500 for the material and another $800 for the jeweler to construct it. After that, she pays 10% of that to get it insured. How much did she pay?
Query sequence (len=7) [SELECT, PROJECT, SELECT, PROJECT, ARITHMETIC, ARITHMETIC, ARITHMETIC]
Ground Truth $1,430
DTW $1,430 ✓
Prefix $1,170 ×\times
DTW-Selected Demonstration
Demo “John commissions a drawing. A black and white costs $160. Color is 50% more. How much for both?”
Demo sequence (len=5) [SELECT, PROJECT, ARITHMETIC, ARITHMETIC, ARITHMETIC]
Reasoning Extract base cost →\rightarrow apply percentage markup →\rightarrow sum.
Prefix match Not a valid prefix. c3c_{3} = ARITHMETIC ≠\neq q3q_{3} = SELECT, so the demo’s entire sequence does not match the query’s prefix. Excluded.
Prefix-Selected Demonstration (longest valid prefix = 6)
Demo “A book costs $20 and a notebook costs $5. With a 10% total discount, final price?”
Demo sequence (len=6) [SELECT, PROJECT, SELECT, PROJECT, ARITHMETIC, ARITHMETIC]
Reasoning Sum two items →\rightarrow apply percentage →\rightarrow subtract (discount).
Prefix match Entire demo sequence c1:6c_{1:6} = q1:6q_{1:6} ✓. Valid prefix.
Analysis
The DTW-selected demo shares the query’s “base cost →\rightarrow percentage” reasoning but is excluded by prefix matching because c3≠q3c_{3}\neq q_{3}. The prefix-selected demo matches the first 6 operations exactly, but its final reasoning step is “subtract discount” while the query requires “add insurance.” The model, misled by the subtraction pattern, outputs $1,170 ($1,300 −- $130) instead of $1,430 ($1,300 + $130).
Table 9: Case 3: DTW selects a structurally aligned demo (single-item percentage markup) despite being excluded by prefix matching due to a granularity mismatch at position 3. Prefix matching instead selects a two-item discount demo whose entire sequence matches the query’s prefix, but whose final reasoning step (subtract discount) contradicts the query’s requirement (add insurance).

F.2.2 Case 4: Incomplete Reasoning — Longest Valid Prefix Omits Percentage Computation

Query
Alice bought a $50 dress and a $30 bag. The store applies 8% sales tax. What is the total?
Query sequence (len=7) [SELECT, PROJECT, SELECT, PROJECT, ARITHMETIC, ARITHMETIC, ARITHMETIC]
Ground Truth $86.40
DTW $86.40 ✓
Prefix $80 ×\times
DTW-Selected Demonstration
Demo “John commissions a drawing. A black and white costs $160. Color is 50% more. How much for both?”
Demo sequence (len=5) [SELECT, PROJECT, ARITHMETIC, ARITHMETIC, ARITHMETIC]
Reasoning Extract base cost →\rightarrow apply percentage →\rightarrow sum.
Prefix match Not a valid prefix. c3c_{3} = ARITHMETIC ≠\neq q3q_{3} = SELECT. Excluded.
Prefix-Selected Demonstration (longest valid prefix = 5)
Demo “A Toyota costs $20,000 and a Honda costs $25,000. What is the total cost?”
Demo sequence (len=5) [SELECT, PROJECT, SELECT, PROJECT, ARITHMETIC]
Reasoning Extract two item costs →\rightarrow sum them.
Prefix match Entire demo sequence c1:5c_{1:5} = q1:5q_{1:5} ✓. Valid prefix.
Analysis
The DTW-selected demo teaches the model to apply a percentage after obtaining a base value, which matches the query’s “add tax” logic. It is excluded by prefix matching because c3≠q3c_{3}\neq q_{3}. The prefix-selected demo is a valid prefix covering 5 of 7 operations, but its reasoning stops at “sum two costs” with no percentage step. The model follows this incomplete pattern, outputs $80 (the pre-tax subtotal), and misses the 8% tax entirely.
Table 10: Case 4: DTW selects a demonstration whose reasoning pattern (base value →\rightarrow percentage →\rightarrow total) matches the query, despite not being a valid prefix. Prefix matching selects a structurally simpler demo that is a valid prefix, but whose reasoning stops at “sum two costs” without any percentage computation, causing the model to omit the tax step.

Table 11 summarizes the four cases.

Case What is compared Key insight
1 13 QDMR vs. + Modulo Without Modulo, remainder is simulated via 3 extra ARITHMETIC steps, attracting proportion problems instead of remainder problems.
2 13 QDMR vs. + Algebra Without Algebra, variable solving expands to 13-step SELECT–PROJECT chains, attracting multi-item subtraction problems.
3 Prefix vs. DTW Prefix matching excludes the best structural match because c3≠q3c_{3}\neq q_{3}; DTW spans the granularity gap. The longest valid prefix demo shares 6 of 7 operations but ends with “subtract” instead of “add,” misleading the model.
4 Prefix vs. DTW The longest valid prefix demo matches 5 of the query’s 7 operations but has no percentage step, causing the model to miss the tax computation. DTW selects a semantically aligned demo (c3≠q3c_{3}\neq q_{3}) that includes percentage reasoning.
Table 11: Summary of the four case studies and their key insights.