LoomSum: Weaving Quantitative and Narrative Evidence for Faithful Long Text–Table Summarization
Abstract
Long documents often distribute important information across extensive narrative passages and multiple tables, making faithful summarization particularly challenging. Existing methods may generate individually supported quantitative facts and analytical statements yet associate them incorrectly, producing quantitatively plausible yet analytically unfaithful summaries. In this work, we propose LoomSum, a training-free framework that extracts source-grounded atomic evidence, explicitly links table-derived facts with supporting narrative analyses, and plans the discourse structure before generation. We also introduce Table-Grounded Faithfulness (TGF), a claim-level metric that separately evaluates Numeric Grounding, Analysis Support, and Relation Consistency. Experiments on the text–table summarization benchmarks FINDSum and USTT show that LoomSum improves analytical faithfulness while maintaining strong summarization quality. Human evaluation finds positive component-level associations with the corresponding human judgments. Our Relation Consistency metric further shows stronger agreement with human relation judgments than generic factuality metrics, indicating that explicit cross-modal linking helps reduce errors in which supported quantities are paired with incorrect narrative interpretations. Together, these findings show that faithful long text–table summarization requires not only grounding individual facts, but also preserving the relations between them.
1 Introduction
Faithful summarization of long, heterogeneous documents remains challenging despite rapid progress in long-context large language models (LLMs). Access to longer inputs reduces the need for aggressive truncation, but does not guarantee that all source regions are used effectively (Ravaut et al., 2024). Long-context Faithful summarization of long, heterogeneous documents remains challenging despite rapid progress in long-context large language models (LLMs). Although longer context windows reduce the need for aggressive truncation, they do not ensure that information from all source regions is used effectively (Ravaut et al., 2024). Long-context models exhibit systematic positional biases (Liu et al., 2024; Cao et al., 2024), unevenly represent evidence from different parts of the input (Ravaut et al., 2024), and become more prone to factual errors when supporting evidence appears in the middle of a document (Wan et al., 2025). Faithful long-document summarization therefore requires not only selecting salient information, but also integrating related evidence distributed across distant and heterogeneous source regions. This challenge is particularly pronounced in financial reports, where quantitative results are often presented in tables, while their causes, implications, and uncertainties are discussed in narrative passages elsewhere (Zhang et al., 2026). Prior work studies this setting through long text–multi-table summarization in FINDSum (Liu et al., 2022) and joint table–text financial summarization in USTT (Wang et al., 2023). However, access to both modalities does not guarantee that evidence from them is correctly associated. Indeed, Cao et al. (2024) identify context mismatch as a prominent form of numerical hallucination in financial-report summaries: a number may be supported by the source but used in an incorrect semantic context. We focus on a specific cross-modal form of this problem: a model may correctly reproduce a quantitative fact and a source-supported analytical statement, yet incorrectly associate the two. For example, a reported revenue increase may be attributed to an explanation that actually concerns a different metric, entity, or reporting period. We refer to this failure as quantitative–narrative misalignment: the quantitative and narrative components of a generated claim may each be supported, while the relation asserted between them is not. Figure 1 illustrates this failure and summarizes our main empirical finding: explicitly modeling cross-modal links reduces omitted and incorrect quantitative–narrative associations, as well as unsupported analyses, while improving the TGF score.
Related Works. Existing approaches only partially address this failure. Hierarchical, extract-then-abstract, and retrieval-based methods reduce long inputs by selecting salient sentences, table records, or relational facts before generation (Chen and Bansal, 2018; Gehrmann et al., 2018; Zhu et al., 2021; Cho et al., 2024; Kirstein et al., 2025). For financial summarization, FINDSum selects textual segments and table tuples separately (Liu et al., 2022), while Wang et al. (2023) model interactions between sentence and table representations. These methods improve content selection, but do not maintain the quantitative–narrative relation as an explicit, traceable unit throughout generation. Many also require task-specific fine-tuning, which introduces adaptation costs when transferring to new datasets or domains. Existing evaluation metrics expose a related limitation (Sun et al., 2026). Reference-based and general factuality metrics measure summary similarity or overall source–claim consistency (Fabbri et al., 2022; Zha et al., 2023; Laban et al., 2022), while attribution frameworks identify supporting evidence for generated claims (Song et al., 2024; Wu et al., 2024; Chu et al., 2025). Recent work further evaluates structured factuality and coverage: Jeong et al. (2025) uses agent-based evaluation for long narratives, Elaraby and Litman (2026) measures the preservation of salient argument roles, and Samarinas et al. (2025) combines atomic-claim verification with aspect-level coverage. Numerical and table-aware metrics additionally assess consistency with structured records or the selection of quantitative content (Dhingra et al., 2019; Wang et al., 2020; Chen et al., 2020; Liu et al., 2022). However, these methods do not specifically evaluate whether a table-derived quantity is associated with its correct narrative interpretation. Consequently, a claim may receive a favorable evaluation when both components are independently grounded, even though their combination is unsupported.
To address these limitations, we propose LoomSum, a training-free claim-to-discourse framework for long text–table summarization. LoomSum first represents narrative passages and table records as source-grounded atomic claims. It then aligns quantitative claims with narrative evidence that explains, compares, qualifies, or contextualizes them, producing linked cross-modal claims with explicit provenance. Related evidence is subsequently consolidated and organized into a discourse plan (Puduppully and Lapata, 2021), from which an off-the-shelf LLM generates the final summary. This representation is designed to preserve quantitative–narrative relations from evidence selection through surface realization without task-specific parameter updates. We further introduce Table-Grounded Faithfulness (TGF), a claim-level evaluation framework that separately assesses Numeric Grounding, Analysis Support, and Relation Consistency. This decomposition distinguishes errors in quantitative facts, narrative interpretations, and the relations between them.
Experiments on FINDSum and USTT show that LoomSum improves analytical faithfulness while maintaining competitive summary quality and content coverage. A claim-level human evaluation further shows that TGF correlates with human judgments, particularly for Numeric Grounding and Relation Consistency.
In summary, our contributions are as follows:
- •
We identify and formalize quantitative–narrative misalignment, and empirically characterize how existing long text–table summarizers omit or mismatch relations between quantitative facts and narrative interpretations.
- •
We propose LoomSum, a training-free claim-to-discourse framework that constructs source-grounded atomic claims, explicitly aligns quantitative and narrative evidence, and organizes the claims through discourse planning for better summarization.
- •
We introduce and human-validate TGF, a reference-summary-free, source-grounded metric that separately evaluates Numeric Grounding, Analysis Support, and Relation Consistency, capturing relation-level errors overlooked by existing factuality and numerical metrics.
- •
We conduct extensive experiments on FINDSum and USTT datasets, demonstrating improvements in analytical faithfulness and summary quality through automatic evaluation, ablation studies, and qualitative analyses.
2 Method
2.1 Overview
In this section, we present LoomSum, a training-free framework for long text–table summarization that explicitly links and organizes textual and tabular evidence before generation. Given a source document , the proposed method constructs source-grounded text and table claims, aligns quantitative facts with their narrative interpretations, consolidates redundant claims, organizes the resulting evidence into a discourse plan, and generates the final summary from this structured representation. The following subsections describe each stage in detail.
2.2 Claim Construction and Cross-Modal Alignment
Claim Construction. Inspired by claim-level decomposition used in FineSurE (Song et al., 2024), LoomSum converts the source into short, independently verifiable claims before generation: . For text claims, the LLM extracts atomic propositions while preserving entities, quantities, units, reporting periods, and stated relations. Each claim retains its supporting source span and document position, enabling subsequent alignment and provenance tracking without relying solely on the model-generated reformulation. For table claims, the model constructs atomic quantitative propositions from serialized table records that preserve the table title, row and column headers, value, unit, and reporting period. Each claim retains a pointer to its original table record, allowing the quantitative content to remain traceable to the structured source.
cross-modal Alignment. Quantitative facts and their narrative interpretations may appear in distant regions of a long document. LoomSum therefore aligns table claims with text claims that explain, compare, qualify, attribute, or contextualize them: . Each cross-modal analytic claim retains the quantitative component, the narrative component, their asserted relation, and the provenance of both. The alignment stage is constrained to link existing source-derived claims rather than introduce new factual content, making the resulting quantitative–narrative relations explicit and traceable.
2.3 Claim Consolidation and Planning
Long documents may repeat the same information across tables, narrative passages, and sections. To reduce redundancy, LoomSum conservatively consolidates claims: . The LLM is instructed to merge claims only when they express the same proposition and agree on the associated entity, metric, reporting period, direction, and numerical value. Claims with conflicting values, temporal scopes, or trends are retained separately. When several equivalent representations exist, a cross-modal analytic claim is preferred because it preserves both the quantitative fact and its narrative interpretation. Provenance from all consolidated instances is retained so that their original supporting evidence remains traceable. Finally, following prior work on content planning Puduppully and Lapata (2021), we organize the consolidated claims into an ordered discourse plan: . The planner specifies which claims should be realized, their discourse roles, their supporting evidence, and the quantitative–narrative relations to preserve. Separating planning from surface realization promotes coherent information ordering and reduces the risk of omitting or separating linked evidence during generation.
2.4 Summary Generation
The summary generator is guided by a small set of demonstrations, , retrieved from the training split through in-context learning. Demonstrations are selected according to structural properties such as table density, cross-modal relation patterns, and source-summary length. Each demonstration contains source-derived claims, their discourse organization, and the corresponding reference summary. These examples illustrate how evidence should be compressed, ordered, and combined while preserving quantitative–narrative relations. Given the discourse plan and the retrieved demonstrations, the generator produces the final summary: . Each generated sentence is also accompanied by the identifiers of its supporting claims, providing lightweight sentence-level provenance and facilitating subsequent inspection.
3 Table-Grounded Faithfulness Score
Existing metrics assess reference similarity, numerical content, or general claim–source consistency, but do not explicitly determine whether a quantitative fact is associated with the correct narrative interpretation. We therefore introduce Table-Grounded Faithfulness (TGF), a reference-summary-free, source-grounded metric with three dimensions: Numeric Grounding(NG), Analysis Support(AS), and Relation Consistency(RC) to evaluate whether generated quantitative analyses are supported by the input document. Following the atomic-fact evaluation paradigm of FActScore (Min et al., 2023), TGF decomposes an LLM-generated summary into a set of independently verifiable claims, , and retrieves an evidence packet from the corresponding source document for each claim. Claims with multiple independent propositions are split, except when a quantitative premise and its interpretation must remain together to evaluate their relation.
LLM-based Scoring. For each dimension , an LLM judge evaluates against using a five-point rubric. A score of denotes unsupported or contradicted content, whereas denotes full support; intermediate scores indicate increasing degrees of partial support. Because the judge does not provide calibrated probabilities over the five ratings, we repeat each judgment times. Let denote the rating assigned to claim for dimension in the -th judgment. We compute the normalized dimension score as
| (6) |
This maps each rating to a scale. We further set , , and .
Evaluation dimensions. Numeric Grounding evaluates whether the quantitative content of is supported by together with its semantic context. The judge considers the value, metric, entity, unit, and direction of change. Numerical surface forms are normalized only to resolve equivalent formats. A correct value associated with the wrong metric, entity, unit, or reporting period therefore does not receive full credit. Analysis Support evaluates whether trends, comparisons, interpretations, and explanations are licensed by the evidence . Plausible but unstated interpretations receive lower scores than analyses explicitly supported by the source. Relation Consistency applies to claims combining quantitative and narrative components and evaluates whether the source supports their asserted association. Relation Consistency is designed to capture this quantitative–narrative misalignment.
Aggregation. Each claim contributes once to TGF. We set for numeric-only claims, for analysis-only claims, and for cross-modal analytic claims. For a cross-modal analytical claim to be faithful, its quantitative fact, analytical explanation, and the relation connecting them must all be supported. We therefore adopt a non-compensatory aggregation in which the claim score is determined by its least-supported required component (the min() operator). This prevents strongly supported individual facts from compensating for an unsupported analytical link. Then, the summary-level score can be written as , and the overall score across the test set is the document macro-average .
4 Experiments
We organize our experiments around three research questions we aim to answer. RQ1: Overall effectiveness. Does LoomSum improve summary quality, numeric content coverage, and table-grounded faithfulness over existing summarization methods? RQ2: Component contributions. How do cross-modal alignment and planning contribute to the performance of LoomSum? RQ3: Metric validity. Does TGF align with human judgments and capture relation-level errors that are overlooked by generic factuality metrics?
4.1 Settings
In this work, we evaluate LoomSum on two complementary table–text summarization benchmarks that differ in language and input scale. FINDSum is an English long-document, multi-table financial summarization dataset (Liu et al., 2022). It contains two subsets: ROO, which focuses on changes in revenue, expenses, and profitability, and Liquidity, which covers liquidity, cash flows, debt, and capital resources. We additionally evaluate on USTT, a Chinese financial table–text summarization dataset (Wang et al., 2023), in which each instance contains a table, an associated narrative passage, and a human-written summary grounded in both modalities. We use the official test split for all datasets. Detailed dataset statistics are provided in Appendix A.
Baselines. We compare LoomSum with methods that use different input modalities. For the text-only setting, TextRank (Mihalcea and Tarau, 2004) is an unsupervised extractive baseline, while BART (Lewis et al., 2020) and BigBird-Pegasus (Zaheer et al., 2020) are neural abstractive summarizers operating only on the narrative text. For the table-only setting, BART and BigBird-Pegasus verbalize selected table tuples, while Table-RAG retrieves relevant table rows before generation. For the joint text–table setting, GCG verbalizes table tuples and combines them with the narrative input for final generation, following Liu et al. (2022). Direct-LC provides the complete narrative and serialized tables directly to a long-context LLM, whereas RAG-Sum retrieves relevant narrative blocks and table rows using fixed summary-oriented queries. Unless otherwise stated, all LLM-based methods use GPT-5-nano (Singh et al., 2025) as the base model and the same decoding configuration as LoomSum.
Evaluation Metrics. We evaluate summary quality using reference-based, numerical-content, and source-grounded faithfulness metrics. For reference-based evaluation, we report ROUGE-1, ROUGE-2, and ROUGE-L scores (Lin, 2004), which measure unigram, bigram, and longest-common-subsequence overlap with the reference summary, respectively. We also report BERTScore (Zhang et al., 2019) to assess semantic similarity, following prior works (Liu et al., 2022; Wang et al., 2023). For numerical-content evaluation, we adopt Numerical Precision (NP), Numerical Coverage (NC), and their harmonic mean, Numerical Selection (NS), from Liu et al. (2022). To evaluate source-grounded factual consistency, we include AlignScore (Zha et al., 2023) and our proposed TGF score, described in Section 3. Because TGF evaluates the faithfulness of expressed claims rather than omitted content, we additionally report Grounded Claim Coverage (G-Cov). We decompose the reference and generated summaries into atomic claims and measure the proportion of reference claims covered by at least one generated claim. G-Cov therefore complements the precision-oriented TGF by identifying methods that achieve high faithfulness through overly conservative generation. In the human validation study, we additionally include SummaC (Laban et al., 2022). AlignScore and SummaC serve as generic factuality baselines, allowing us to assess whether TGF’s dimension-specific design provides additional agreement with human judgments.
5 Results
We evaluate the generated summaries using automatic metrics and human judgments and present the results in this section. We further conduct ablation studies to examine the contribution of individual components and provide a case study.
5.1 Results on FINDSum (RQ1)
Table 1 reports the main results on the ROO and Liquidity subsets of FINDSum. Due to space constraints, we report ROUGE-1 (R-1), ROUGE-2 (R-2), ROUGE-L (R-L), BERTScore (BS), Numeric Selection (NS), and Table-Grounded Faithfulness (TGF) in the main text; complete results and detailed analysis are provided in Appendix C. Several patterns emerge from the results. First, LLM-based abstractive methods generally outperform unsupervised baselines across both reference-based and faithfulness-oriented metrics under all three input configurations. Text-only methods also generally outperform their table-only counterparts, suggesting that narrative passages provide a more directly usable representation of the content emphasized in the reference summaries. Among all methods, LoomSum achieves the strongest performance on most reported metrics across both subsets, including the highest NS and TGF scores. The NS gains indicate improved retention of salient numerical content, while the TGF improvements show that quantities, analytical statements, and their asserted relations are more consistently supported by the source. These results support the central design of LoomSum: cross-modal alignment preserves quantitative–narrative relations, while discourse planning integrates the aligned evidence into a coherent summary.
Component-wise TGF analysis. Figure 3 shows three component-level scores of TGF: Numeric Grounding (NG), Analysis Support (AS), and Relation Consistency (RC). On FINDSum-ROO, LoomSum achieves scores of 74 in NG, 72 in AS, and 69 in RC, respectively, outperforming the strongest baseline. The gains are larger on FINDSum-Liquidity, where LoomSum reaches 82 in NG, 80 in AS, and 77 in RC. The progressively larger gains from NG to RC indicate that LoomSum’s primary advantage lies not only in reproducing grounded quantities, but also in preserving the analytical relations in which those quantities participate.
| Input | Method | FINDSum-Liquidity | FINDSum-ROO | ||||||||||
| R-1 | R-2 | R-L | BS | NS | TGF | R-1 | R-2 | R-L | BS | NS | TGF | ||
| Text- only | TextRank | 33.06 | 6.98 | 13.54 | 74.33 | 20.05 | 65.86 | 35.04 | 7.34 | 14.94 | 76.81 | 24.80 | 63.19 |
| BART | 21.93 | 6.22 | 10.78 | 72.92 | 20.44 | 63.70 | 23.19 | 5.83 | 10.39 | 67.81 | 19.80 | 54.33 | |
| BigBird-Pegasus | 25.84 | 7.86 | 12.59 | 73.80 | 21.40 | 64.28 | 24.10 | 7.01 | 13.49 | 68.24 | 22.80 | 64.37 | |
| Table- only | BART | 17.71 | 5.98 | 8.47 | 66.53 | 7.50 | 48.15 | 18.37 | 4.62 | 8.96 | 67.21 | 6.50 | 46.69 |
| BigBird-Pegasus | 23.16 | 6.11 | 11.50 | 68.50 | 13.34 | 50.92 | 19.25 | 6.90 | 12.91 | 69.70 | 12.47 | 52.22 | |
| Table-RAG† | 30.05 | 5.88 | 12.70 | 65.59 | 14.68 | 48.92 | 31.80 | 7.32 | 14.91 | 69.72 | 12.27 | 51.93 | |
| Text + Table | GCG† | 33.96 | 7.79 | 14.14 | 76.54 | 24.59 | 66.71 | 29.25 | 8.89 | 16.58 | 80.53 | 23.75 | 64.47 |
| Direct-LC† | 32.01 | 7.49 | 13.74 | 75.38 | 25.11 | 69.29 | 30.19 | 8.83 | 16.52 | 80.55 | 24.10 | 68.21 | |
| RAG-Sum† | 33.51 | 7.72 | 14.30 | 76.31 | 23.48 | 76.34 | 27.82 | 8.04 | 15.80 | 80.70 | 27.08 | 69.47 | |
| LoomSum (Ours)† | 32.74 | 9.51 | 16.35 | 76.78 | 25.98 | 79.43 | 35.45 | 10.77 | 18.55 | 81.34 | 27.03 | 71.38 | |
5.2 Results on USTT (RQ1)
Table 3 reports the results on USTT. We observe similar findings as in FindSum. First, Table-only methods perform substantially worse than text-based and multimodal counterparts, suggests that tabular records alone provide insufficient context for recovering the narrative structure and analytical content emphasized in the reference summaries. LoomSum achieves the highest R-1, R-L, and NS scores, indicating stronger overall content coverage and substantially better retention of reference-relevant numerical information.
An interesting result is that TextRank obtains the highest TGF score despite considerably lower ROUGE and NS performance. This reflects the distinction between faithfulness and coverage: an extractive method can achieve high claim-level faithfulness by selecting a small set of directly supported sentences while omitting other salient information. The component scores in the right panel of Figure 3 clarify this result. TextRank achieves the highest NG score, whereas LoomSum obtains the highest AS and RC scores. Thus, TextRank is effective at preserving isolated source quantities, while LoomSum provides stronger analytical support and more reliable quantitative–narrative integration.
Across all three evaluation sets, RC is generally the most challenging dimension. Nevertheless, LoomSum consistently achieves the highest RC score, providing direct evidence that explicit cross-modal alignment mitigates the quantitative–narrative misalignment identified in Section 1. More detailed analysis can be found in Appendix.
| Input | Method | R-1 | R-2 | R-L | BS | NS | TGF |
| Text- only | TextRank | 24.77 | 10.20 | 13.83 | 60.24 | 21.44 | 66.41 |
| BART | 22.84 | 8.90 | 14.25 | 59.80 | 19.12 | 54.26 | |
| BigBird-Pegasus | 24.18 | 7.69 | 14.17 | 60.12 | 20.35 | 56.91 | |
| Table- only | BART | 20.46 | 5.21 | 13.72 | 60.01 | 13.32 | 52.53 |
| BigBird-Pegasus | 20.71 | 5.54 | 12.01 | 59.25 | 13.98 | 55.79 | |
| Table-RAG | 19.26 | 5.66 | 10.77 | 61.50 | 12.18 | 56.13 | |
| Text + Table | GCG | 24.83 | 8.93 | 12.92 | 63.84 | 18.76 | 59.71 |
| Direct-LC | 25.14 | 9.37 | 13.26 | 63.71 | 18.63 | 58.51 | |
| RAG-Sum | 24.56 | 9.91 | 12.93 | 64.47 | 15.89 | 60.64 | |
| LoomSum (Ours) | 26.10 | 9.18 | 15.16 | 64.16 | 27.36 | 63.71 |
| Variant | Align | Plan | ICL | R-L | NS | TGF |
| FINDSum–ROO | ||||||
| Flat Evidence | ✗ | ✗ | ✗ | 13.47 | 21.46 | 55.68 |
| + Cross-Modal Alignment | ✓ | ✗ | ✗ | 15.89 | 22.30 | 62.93 |
| + Planning | ✓ | ✓ | ✗ | 16.04 | 25.71 | 67.34 |
| LoomSum (Ours) | ✓ | ✓ | ✓ | 18.55 | 27.03 | 71.38 |
| FINDSum–Liquidity | ||||||
| Flat Evidence | ✗ | ✗ | ✗ | 13.12 | 20.21 | 58.94 |
| + Cross-Modal Alignment | ✓ | ✗ | ✗ | 14.47 | 21.47 | 65.69 |
| + Planning | ✓ | ✓ | ✗ | 16.44 | 24.16 | 74.48 |
| LoomSum (Ours) | ✓ | ✓ | ✓ | 16.35 | 25.98 | 79.43 |
5.3 Results on Human Evaluation (RQ3)
We conduct a human evaluation to examine whether the three components of TGF—Numeric Grounding (NG), Analysis Support (AS), and Relation Consistency (RC)—agree with human judgments of the corresponding properties. Rather than eliciting a single holistic faithfulness score, we evaluate each dimension independently, consistent with the multidimensional design of TGF.
Annotation setup. We randomly select 100 samples(documents) and their generated summaries from the FINDSum test set. The study is done by two annotators. NG is evaluated for claims containing quantitative information, AS for claims containing analytical statements, and RC for claims that associate a quantitative fact with a narrative interpretation. Each annotator independently assesses the applicable dimensions using the corresponding source evidence, including relevant table records and narrative passages. Annotators are blinded to the generation methods and all automatic metric scores.
How well does TGF align with human judgments? We assess the validity of each TGF component by computing Spearman’s and Kendall’s between the LLM-derived scores defined in Section 3 and the corresponding human scores. As shown in Figure 3, TGF-NG exhibits a moderate positive correlation with human Numeric Grounding judgments (, , ). TGF-RC achieves the strongest correlation (, , ), indicating that it reliably reflects human assessments of whether quantitative facts are associated with the correct narrative interpretations. TGF-AS also correlates positively with human Analysis Support judgments (, ), although the association is not statistically significant. Overall, the results provide human validation for the NG and RC components, with particularly strong evidence for Relation Consistency. The weaker AS correlation suggests that determining whether an interpretation is sufficiently supported remains more subjective and challenging for both human and automatic evaluation.
Does TGF capture human judgments beyond generic factuality metrics? We additionally compare TGF with AlignScore and SummaC, two generic factuality metrics that each produce a single overall consistency score. For a fair comparison, we correlate each metric with Human NG, AS, and RC on the same annotated samples. Figure 4 shows that TGF consistently achieves stronger agreement with human judgments than the generic baselines across both Spearman’s and Kendall’s . The advantage is especially pronounced for Relation Consistency, where TGF substantially outperforms AlignScore and SummaC. This result supports our central claim that generic source–claim consistency metrics are not sufficient to detect quantitative–narrative relation errors. TGF-RC explicitly targets this failure mode, which explains its markedly stronger correspondence with human RC judgments. We also observe gains for Numeric Grounding and Analysis Support, suggesting that the component-specific design of TGF provides more fine-grained alignment with human assessment than a single undifferentiated factuality score.
5.4 Ablations (RQ2)
Table 3 reports a cumulative ablation of the three main components of LoomSum. Flat Evidence uses only independently extracted text and table claims, without explicit cross-modal alignment, planning, or in-context demonstrations. We then progressively add each component until reaching the full model. cross-modal alignment consistently improves all metrics on both FINDSum subsets, showing that incorporating table claims and explicitly connecting them to narrative evidence improves quantitative–narrative faithfulness beyond text-only evidence. Adding planning yields further gains, particularly in NS and TGF. These improvements suggest that organizing related claims before generation helps retain salient numerical content and preserve their associated interpretations. Finally, in-context demonstrations produce the full LoomSum model and further improve NS and TGF on both subsets. Although R-L score slightly decreases from 16.44 to 16.35 in the Liquidity subset, the NS and TGF continue to improve, indicating that in-context guidance primarily benefits faithful evidence realization rather than lexical overlap.
Sensitivity Analysis. We further examine the robustness of the proposed method to the choice of base model and observe broadly consistent performance across different tasks; full results are reported in Appendix B.
5.5 Case Study
Finally, to complement the main results, we examine an example from the FINDSum-ROO testset in Table 4. The source establishes a quantitative–narrative chain in which “$51.4 million in net interest income and a $10.8 million credit-loss provision benefit yield $62.2 million in post-provision net interest income”. The accompanying narrative further identifies the mortgage-loan portfolio and investments in beneficial interests as the sources of the provision benefit.
| Source evidence. Table: Net interest income: 51.4; provision benefit: +10.8; post-provision net interest income: 62.2. Text: “… credit loss provisions of $10.8 million on our mortgage loan portfolio and investments in beneficial interests during the year ended December 31, 2020.” | |
| Direct-LC | “… initial credit loss provisions of $5.1 million in Q1 2020, the company recorded a net recovery of $10.8 million in credit loss provisions for the full year 2020 due to better-than-expected loan performance. …” |
| GCG | “… credit loss provisions of $10.8 million on mortgage loans and beneficial interests for the year ended December 31, 2020; net $0.9 million of unrealized losses …” |
| RAG-Sum | “… macroeconomic conditions and expected lifetime credit losses on our mortgage loan and beneficial interest portfolios are subject to meaningful uncertainty …” |
| LoomSum | “… net interest income after provision for credit benefit/(losses) for 2020 was $62.2 million, compared with the original $51.4 million, reflecting the $10.8 million provision benefit recorded on our mortgage loan portfolio and investments in beneficial interests …” |
| Reference | “… our net interest income after the recovery of provision for credit losses increased to $62.2 million […] primarily as a result of $10.8 million in net recovery of provision for credit losses …” |
The baselines preserve only fragments of the source relation. Direct-LC reports the $10.8 million recovery but omits its effect on the $62.2 million post-provision income and introduces an unsupported explanation. GCG similarly retains the provision benefit as an isolated fact, while RAG-Sum shifts toward general risk discussion and omits the central quantities altogether. In contrast, LoomSum connects the provision benefit to the resulting income and retains its portfolio-level attribution, consistent with the reference. This example illustrates how explicit cross-modal alignment preserves a coherent driver–outcome relation rather than disconnected facts, matching the aggregate gains in TGF Relation Consistency. We provide more examples for qualitative analysis in Appendix D.
6 Conclusion
In this work, we formalize quantitative–narrative misalignment, where individually supported quantitative facts and analytical statements are combined into unsupported relations. We then propose LoomSum, a training-free framework that explicitly links table-derived facts with narrative interpretations and organizes them through discourse planning. We also introduce Table-Grounded Faithfulness (TGF) to evaluate Numeric Grounding, Analysis Support, and Relation Consistency. Experiments on FINDSum and USTT show that LoomSum improves analytical faithfulness without compromising reference-based quality. Human evaluation supports TGF, with the Relation Consistency component aligning more closely with human judgments than generic factuality metrics. Overall, our results show that faithful text–table summarization requires preserving not only individual facts, but also the relations between them. Future work may explore richer relation-aware models and broader domains.
Limitations
This work provides an initial investigation of quantitative–narrative misalignment in long text–table summarization. LoomSum is intentionally designed as a lightweight, training-free framework to isolate the value of explicitly linking quantitative facts with their narrative interpretations. It should therefore be viewed as one possible instantiation of relation-aware summarization rather than an exhaustive solution to the problem. Its modular pipeline may also propagate errors across evidence extraction, cross-modal alignment, discourse planning, and generation. Future work could explore jointly optimized or learned relation models, richer evidence graphs, and generation objectives that directly preserve cross-modal relations.
Our empirical evaluation is conducted on FINDSum and USTT, both of which contain financial text–table inputs. Although the underlying failure mode may arise in other domains that combine structured records with narrative explanations, such as scientific, medical, and policy documents, its prevalence and characteristics outside financial reporting remain to be established. Extending the evaluation to broader domains, languages, table structures, and document lengths is therefore an important direction for future work.
Finally, TGF relies on automated claim extraction and LLM-based judgments, which may introduce model-dependent errors and sensitivity to evidence presentation Fan et al. (2026). Our human study provides positive validation for Numeric Grounding and Relation Consistency, but is limited in scale, and the weaker association observed for Analysis Support suggests that analytical interpretations remain more difficult to evaluate consistently. In addition, TGF is precision-oriented and evaluates the faithfulness of expressed claims rather than the omission of salient content. Future work could combine relation-aware faithfulness with source-level coverage evaluation and validate the metric using larger and more diverse human annotations.
References
- Characterizing multimodal long-form summarization: a case study on financial reports. arXiv preprint arXiv:2404.06162. External Links: Link Cited by: §1.
- Logical natural language generation from open-domain tables. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 7929–7942. External Links: Link Cited by: §1.
- Fast abstractive summarization with reinforce-selected sentence rewriting. In Proceedings of the 56th annual meeting of the association for computational linguistics (volume 1: Long papers), pp. 675–686. Cited by: §1.
- RTSUM: relation triple-based interpretable summarization with multi-level salience visualization. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 3: System Demonstrations), pp. 53–60. External Links: Link Cited by: §1.
- TRACSUM: a new benchmark for aspect-based summarization with sentence-level traceability in medical domain. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pp. 844–864. External Links: Link Cited by: §1.
- Handling divergent reference texts when evaluating table-to-text generation. In Proceedings of the 57th annual meeting of the association for computational linguistics, pp. 4884–4895. External Links: Link Cited by: §1.
- Arc: argument representation and coverage analysis for zero-shot long document summarization with instruction following llms. In Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 3626–3643. External Links: Link Cited by: §1.
- QAFactEval: improved qa-based factual consistency evaluation for summarization. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 2587–2601. External Links: Link Cited by: §1.
- The moving target: a longitudinal audit of trust-benchmark score drift across open-source chat llm release lines. External Links: 2607.02587, Link Cited by: Limitations.
- Bottom-up abstractive summarization. In Proceedings of the 2018 conference on empirical methods in natural language processing, pp. 4098–4109. External Links: Link Cited by: §1.
- Agent-as-judge for factual summarization of long narratives. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pp. 23602–23619. External Links: Link Cited by: §1.
- Re-frame the meeting summarization scope: fact-based summarization and personalization via questions. In Findings of the Association for Computational Linguistics: EMNLP 2025, pp. 20087–20137. External Links: Link Cited by: §1.
- SummaC: re-visiting nli-based models for inconsistency detection in summarization. Transactions of the Association for Computational Linguistics 10, pp. 163–177. External Links: Link Cited by: §1, §4.1.
- BART: denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th annual meeting of the association for computational linguistics, pp. 7871–7880. External Links: Link Cited by: §4.1.
- Rouge: a package for automatic evaluation of summaries. In Text summarization branches out, pp. 74–81. External Links: Link Cited by: §4.1.
- Lost in the middle: how language models use long contexts. Transactions of the association for computational linguistics 12, pp. 157–173. External Links: Link Cited by: §1.
- Long text and multi-table summarization: dataset and method. In Findings of the Association for Computational Linguistics: EMNLP 2022, pp. 1995–2010. External Links: Link Cited by: Appendix C, §1, §1, §4.1, §4.1, §4.1.
- Textrank: bringing order into text. In Proceedings of the 2004 conference on empirical methods in natural language processing, pp. 404–411. External Links: Link Cited by: §4.1.
- Factscore: fine-grained atomic evaluation of factual precision in long form text generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 12076–12100. External Links: Link Cited by: §3.
- Data-to-text generation with macro planning. Transactions of the Association for Computational Linguistics 9, pp. 510–527. External Links: Link Cited by: §1, §2.3.
- On context utilization in summarization with large language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 2764–2781. External Links: Link Cited by: §1.
- Beyond factual accuracy: evaluating coverage of diverse factual information in long-form text generation. In Findings of the Association for Computational Linguistics: ACL 2025, pp. 13468–13482. External Links: Link Cited by: §1.
- Openai gpt-5 system card. arXiv preprint arXiv:2601.03267. External Links: Link Cited by: §4.1.
- FineSurE: fine-grained summarization evaluation using llms. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 906–922. External Links: Link Cited by: §1, §2.2.
- Beyond accuracy: measuring bias acknowledgment in chain-of-thought reasoning for responsible ai evaluation. arXiv preprint arXiv:2606.15127. External Links: Link Cited by: §1.
- On positional bias of faithfulness for long-form summarization. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pp. 8791–8810. External Links: Link Cited by: §1.
- Towards faithful neural table-to-text generation with content-matching constraints. In Proceedings of the 58th annual meeting of the association for computational linguistics, pp. 1072–1086. External Links: Link Cited by: §1.
- Beyond pure text: summarizing financial reports based on both textual and tabular data.. In IJCAI, pp. 5233–5241. External Links: Link Cited by: §1, §1, §4.1, §4.1.
- Less is more for long document summary evaluation by llms. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 2: Short Papers), pp. 330–343. External Links: Link Cited by: §1.
- Big bird: transformers for longer sequences. Advances in neural information processing systems 33, pp. 17283–17297. External Links: Link Cited by: §4.1.
- AlignScore: evaluating factual consistency with a unified alignment function. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 11328–11348. External Links: Link Cited by: §1, §4.1.
- Bertscore: evaluating text generation with bert. arXiv preprint arXiv:1904.09675. External Links: Link Cited by: §4.1.
- Bankruptcy prediction from 10-k narratives: evidence from interpretable text scores and accounting baselines. Risks 14 (8), pp. 179. External Links: Link Cited by: §1.
- Enhancing factual consistency of abstractive summarization. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 718–733. External Links: Link Cited by: §1.
Contents of Appendix
Appendix A Datasets Statistics
Table 5 summarizes the statistics of the test sets used in our experiments. We report the number of instances, the average source length, and the average reference summary length for each dataset. All lengths are measured in tokens using the GPT o200_base tokenizer. For FINDSum, the source length includes both the narrative content and the associated tabular information.
The two FINDSum subsets represent substantially longer inputs than USTT. FINDSum-Liquidity contains 4,204 test instances, with an average source length of 34,630 tokens and an average reference summary length of 538 tokens. FINDSum-ROO contains a larger test set of 6,201 instances and slightly longer source documents, averaging 36,447 tokens, while its reference summaries are somewhat shorter, with an average length of 444 tokens. These statistics highlight the long-context nature of FINDSum, where a summarization method must identify and integrate salient evidence distributed across extensive narrative passages and multiple tables.
In contrast, USTT contains 1,390 test instances, with an average source length of 1,680 tokens and an average summary length of 85 tokens. Although USTT is considerably shorter than FINDSum, it provides a complementary table–text summarization setting with a different language and input scale. Evaluating on both datasets therefore allows us to examine whether the proposed method is effective not only for long financial reports, but also for shorter table–text inputs.
| Dataset | # Inst. | Avg. Doc. | Avg. Sum. |
| FindSum-Liquidity | 4204 | 34630 | 538 |
| FindSum-ROO | 6201 | 36447 | 444 |
| USTT | 1390 | 1680 | 85 |
Appendix B Sensitive Analysis of Base Model
Due to resource limit, we conduct the sensitive analysis on 500 samples each, from both datasets.
Appendix C Additional Quantitative Results
Tables 6 and 7 report the complete quantitative results on the FINDSum-ROO and FINDSum-Liquidity test sets, respectively. In addition to the metrics reported in the main text, we include Numerical Precision (NP), Numerical Coverage (NC), AlignScore, SummaC, and Grounded Claim Coverage (G-Cov).
Grounded Claim Coverage (G-Cov). G-Cov is a reference-based grounded coverage metric that complements the precision-oriented TGF. TGF evaluates the faithfulness of the claims expressed by a method and averages over the generated claim set. Consequently, a short or conservative summary may obtain a high TGF by expressing only a small number of well-supported claims. G-Cov instead uses claims from the reference summary and measures how much salient reference content is both recovered by the generated summary and grounded in the source.
For each document , we decompose the reference and generated summaries into atomic claim sets and as we done in TGF. For each reference claim , we select its best-aligned generated claim:
| (7) |
where is an LLM-as-Judge-based semantic alignment rating. Alignment evaluates whether the generated claim preserves the content of the reference claim, including its entities, metrics, values, periods, directions, and asserted relations. The normalized alignment credit is
| (8) |
The matched generated claim is additionally evaluated for Numeric Grounding (NG), Analysis Support (AS), and Relation Consistency (RC) against the source evidence. Each applicable rating is normalized to :
| (9) |
According to the type of the reference claim, its faithfulness credit is defined as
| (10) |
Then, grounded coverage credit assigned to is
| (11) |
and the document-level G-Cov score is
| (12) |
The final dataset score is obtained by averaging over documents.
Alignment and faithfulness serve distinct roles in this definition. A generated claim may resemble a reference claim but remain unsupported by the source; conversely, it may be source-grounded but express content unrelated to the reference claim. A reference claim receives high G-Cov credit only when both semantic coverage and source grounding are strong. G-Cov should therefore be interpreted as grounded recall of the salient content represented by the reference summary, rather than exhaustive coverage of all information in the source document.
| Input | Method | R-1 | R-2 | R-L | BS | NP | NC | NS | AlignScore | SummaC | G-Cov | TGF |
| Text- only | TextRank | 35.04 | 7.34 | 14.94 | 76.81 | 24.65 | 40.90 | 24.80 | 90.68 | 54.96 | 12.26 | 63.19 |
| BART | 23.19 | 5.83 | 10.39 | 67.81 | 22.41 | 21.56 | 19.80 | 94.63 | 81.82 | 10.73 | 54.33 | |
| BigBird-Pegasus | 24.10 | 7.01 | 13.49 | 68.24 | 13.17 | 31.29 | 22.80 | 86.78 | 64.09 | 10.54 | 64.37 | |
| Table- only | BART | 18.37 | 4.62 | 8.96 | 67.21 | 13.24 | 6.70 | 6.50 | 56.96 | 34.38 | 10.51 | 46.69 |
| BigBird-Pegasus | 19.25 | 6.90 | 12.91 | 69.70 | 8.17 | 20.43 | 12.47 | 50.45 | 52.28 | 9.02 | 52.22 | |
| Table-RAG† | 31.80 | 7.32 | 14.91 | 69.72 | 10.46 | 18.86 | 12.27 | 22.67 | 44.84 | 11.71 | 51.93 | |
| Text + Table | GCG† | 29.25 | 8.89 | 16.58 | 80.53 | 22.64 | 30.49 | 23.75 | 95.31 | 62.07 | 13.57 | 64.47 |
| Direct-LC† | 30.19 | 8.83 | 16.52 | 80.55 | 21.27 | 29.02 | 24.10 | 92.66 | 85.92 | 14.72 | 68.21 | |
| RAG-Sum† | 27.82 | 8.04 | 15.80 | 80.70 | 25.57 | 42.54 | 27.08 | 88.21 | 85.93 | 13.10 | 69.47 | |
| LoomSum (Ours)† | 35.45 | 10.77 | 18.55 | 81.34 | 22.53 | 54.17 | 27.03 | 90.28 | 86.46 | 17.75 | 71.38 |
| Input | Method | R-1 | R-2 | R-L | BS | NP | NC | NS | AlignScore | SummaC | Cov. | TGF |
| Text- only | TextRank | 33.06 | 6.98 | 13.54 | 74.33 | 17.18 | 42.37 | 20.05 | 80.53 | 62.07 | 12.06 | 65.86 |
| BART | 21.93 | 6.22 | 10.78 | 72.92 | 21.89 | 24.20 | 20.44 | 85.74 | 72.49 | 10.56 | 63.70 | |
| BigBird-Pegasus | 25.84 | 7.86 | 12.59 | 73.80 | 22.33 | 23.97 | 21.40 | 83.17 | 60.20 | 10.24 | 64.28 | |
| Table- only | BART | 17.71 | 5.98 | 8.47 | 66.53 | 11.68 | 8.53 | 7.50 | 56.25 | 33.24 | 9.20 | 48.15 |
| BigBird-Pegasus | 23.16 | 6.11 | 11.50 | 68.50 | 9.43 | 31.25 | 13.34 | 59.53 | 54.68 | 10.05 | 50.92 | |
| Table-RAG† | 30.05 | 5.88 | 12.70 | 65.59 | 10.70 | 33.95 | 14.68 | 28.69 | 44.56 | 11.20 | 48.92 | |
| Text + Table | GCG† | 33.96 | 7.79 | 14.14 | 76.54 | 14.29 | 51.45 | 24.59 | 78.19 | 63.05 | 16.25 | 66.71 |
| Direct-LC† | 32.01 | 7.49 | 13.74 | 75.38 | 15.34 | 54.78 | 25.11 | 79.22 | 64.29 | 17.75 | 69.29 | |
| RAG-Sum† | 33.51 | 7.72 | 14.30 | 76.31 | 16.57 | 50.77 | 23.48 | 74.38 | 63.83 | 14.84 | 76.34 | |
| LoomSum (Ours)† | 32.74 | 9.51 | 16.35 | 76.78 | 18.17 | 56.76 | 25.98 | 78.73 | 70.53 | 17.96 | 79.43 |
We begin with the analysis of overall performance. LoomSum achieves the strongest overall performance across the two FINDSum subsets, although it does not outperform every individual metric. On FINDSum-ROO, it obtains the best result on 8/11 reported metrics, including all three ROUGE scores, BERTScore, NC, SummaC, G-Cov, and TGF. On FINDSum-Liquidity, it achieves the best R-2, R-L, BERTScore, NC, NS, G-Cov, and TGF. These results indicate that organizing source-grounded evidence and preserving cross-modal relations improve analytical faithfulness without requiring a corresponding loss in reference-based summary quality.
Numerical content selection. Next, we analyze an important set of metrics (NP, NC, NS) as introduced in Liu et al. (2022). LoomSum achieves the highest NC on both subsets, reaching 54.17 on ROO and 56.76 on Liquidity. It recovers a larger proportion of the numerical content selected by the reference summaries. Its NP is not always the highest, suggesting that broader numerical coverage may include additional source-supported quantities that do not appear in the single reference summary. On ROO, this precision–coverage balance yields an NS of 27.03, which is nearly identical to the best score of 27.08 obtained by RAG-Sum. On Liquidity, LoomSum achieves the highest NS of 25.98.
Effect of input modality. We then shift our focus to analyze the impact of input modalities. The table-only methods generally obtain the lowest TGF and G-Cov scores. Although tables provide precise quantitative values, they contain less of the explanatory and contextual information needed to reconstruct the reference summaries. Text-only methods perform more strongly, reflecting the importance of narrative evidence. TextRank, in particular, remains competitive on several metrics, showing that extractive selection can preserve locally supported and reference-relevant information. Methods using both modalities generally achieve the strongest joint performance. This supports their complementary roles: tables provide exact quantities, whereas narrative passages describe their drivers, qualifications, and implications. The advantage of LoomSum over other joint-input methods further indicates that access to both modalities is not sufficient by itself; related evidence must also be correctly associated and organized before generation.
Faithfulness–coverage trade-off. We argue that TGF and G-Cov capture complementary properties. TGF averages faithfulness over the generated claims and therefore does not penalize information that a system never attempts to express. G-Cov instead averages grounded coverage over the reference claims, so omitted reference content receives little or no credit. A method with high TGF but low G-Cov is therefore generating a relatively small or conservative set of reliable claims rather than providing comprehensive grounded coverage. RAG-Sum illustrates this distinction. On ROO, it obtains a relatively high TGF of 69.47 but a G-Cov of only 13.10. The pattern is more pronounced on Liquidity, where it reaches a TGF of 76.34 but a G-Cov of 14.84. Its retrieval stage may expose the generator to a compact set of highly relevant evidence, improving the faithfulness of the claims that are produced. At the same time, restricting generation to the retrieved evidence can leave other reference claims uncovered. Direct-LC exhibits a different trade-off on Liquidity. Its G-Cov of 17.75 is close to the best score, indicating that it expresses a relatively broad range of reference content. However, its TGF of 69.29 is substantially below that of RAG-Sum and LoomSum. Providing the full source makes more information available, but it also increases the difficulty of correctly grounding quantities and associating them with the appropriate narrative interpretations. These comparisons demonstrate why TGF and G-Cov should be interpreted jointly. Moreover, LoomSum achieves the highest TGF and G-Cov on both FINDSum subsets, indicating that its faithfulness gains do not arise merely from generating fewer or more conservative claims.
Comparison with generic factuality metrics. Lastly, we provide analysis around generic factuality metrics we utilize in this work. AlignScore and SummaC produce rankings that do not always agree with TGF or G-Cov. On ROO, GCG obtains the highest AlignScore of 95.31, but its TGF and G-Cov are only 64.47 and 13.57, respectively. Text-only BART similarly obtains high AlignScore and SummaC scores of 94.63 and 81.82, while achieving substantially lower TGF and G-Cov scores of 54.33 and 10.73. The discrepancy is also visible on Liquidity, where text-only BART obtains the highest AlignScore and SummaC scores but reaches only 63.70 TGF and 10.56 G-Cov. Generic factuality metrics can therefore assign favorable scores to locally consistent summaries without reflecting whether a broad set of salient reference claims is covered or whether quantitative facts are associated with their correct narrative interpretations. TGF and G-Cov therefore provide complementary relation-sensitive faithfulness and grounded-coverage signals that are not fully captured by these generic scores.
| Input | Method | R-1 | R-2 | R-L | BS | NP | NC | NS | AlignScore | SummaC | G-Cov | TGF |
| Text- only | TextRank | 24.77 | 10.20 | 13.83 | 60.24 | 12.02 | 30.36 | 21.44 | 77.35 | 52.20 | 19.12 | 66.41 |
| BART | 22.84 | 8.90 | 14.25 | 59.80 | 14.81 | 26.99 | 19.12 | 67.30 | 42.03 | 17.18 | 54.26 | |
| BigBird-Pegasus | 24.18 | 7.69 | 14.17 | 60.12 | 13.22 | 27.51 | 20.35 | 64.42 | 47.83 | 16.74 | 56.91 | |
| Table- only | BART | 20.46 | 5.21 | 13.72 | 60.01 | 10.45 | 19.52 | 13.32 | 36.76 | 47.91 | 11.24 | 52.53 |
| BigBird-Pegasus | 20.71 | 5.54 | 12.01 | 59.25 | 12.33 | 18.89 | 13.98 | 46.16 | 49.80 | 11.45 | 55.79 | |
| Table-RAG† | 19.26 | 5.66 | 10.77 | 61.50 | 13.31 | 12.74 | 12.18 | 73.28 | 42.44 | 10.16 | 56.13 | |
| Text + Table | GCG† | 24.83 | 8.93 | 12.92 | 63.84 | 19.50 | 24.21 | 18.76 | 64.77 | 42.31 | 16.19 | 59.71 |
| Direct-LC† | 25.14 | 9.37 | 13.26 | 63.71 | 17.76 | 25.78 | 18.63 | 74.55 | 54.70 | 18.43 | 58.51 | |
| RAG-Sum† | 24.56 | 9.91 | 12.93 | 64.47 | 16.04 | 23.89 | 15.89 | 69.47 | 42.93 | 15.28 | 60.64 | |
| LoomSum (Ours)† | 26.10 | 9.18 | 15.16 | 64.16 | 19.81 | 49.79 | 27.36 | 77.99 | 54.02 | 19.78 | 63.71 |
Results on USTT. On USTT, we observe very similar patterns as in FindSum. In Table 8, LoomSum achieves the best R-1, R-L, NP, NC, NS, AlignScore, and G-Cov results. Its advantage is particularly clear for numerical content selection, improving NC from 30.36 for the next-best method to 49.79 and NS from 21.44 to 27.36. TextRank obtains the highest TGF, likely benefiting from its conservative extractive behavior, but achieves a lower G-Cov. This contrast further demonstrates the complementarity of the two metrics: higher faithfulness among the claims expressed does not necessarily imply broader grounded coverage of the reference content. Among joint-input methods, LoomSum achieves the strongest results on both TGF and G-Cov.
Appendix D Additional Case Studies
We further examine two complementary case studies from FINDSum and USTT. The FINDSum example focuses on relation preservation when quantitative facts and their explanations are distributed across different parts of a report. The USTT example instead illustrates grounded omission, where methods generate source-supported content but fail to cover the claims emphasized by the reference summary.
Table 9 presents an example from Commerce Bancshares’ 2015 annual report. We show the source-supported relations, the reference summary, and verbatim excerpts from each method’s output. Ellipses indicate omitted sentences. Red highlights only the phrase responsible for an unsupported relation, whereas green highlights correctly preserved relations in the output.
| Source / Method | Evidence or Output Excerpt | Analysis |
| Source Document |
R1: Total revenue increased by $25.7 million, comprising
$11.6 million growth in non-interest income and $14.1 million growth
in net interest income.
R2: Loan interest income increased by $10.2 million because average loans grew by $609.0 million, partly offset by a 12-basis-point decline in rates. R3: Net yield declined from 3.00% to 2.94% because of continuing low interest rates. |
The source establishes the revenue decomposition, the complete loan-interest driver chain, and the cause of margin compression. |
| Reference Summary | Net income increased to $263.7 million from $261.8 million. The increase resulted from $14.1 million growth in net interest income and $11.6 million growth in non-interest income. … These gains were partly offset by higher non-interest expense and lower investment-security gains. | The reference preserves the main income decomposition, but omits the loan-interest driver chain and the explanation for the yield decline. |
| Direct-LC | Total revenue increased $25.7 million because continuing low interest rates depressed net yield to 2.94%. … Net interest income was $634.3 million, while non-interest income was $447.6 million, an increase of $11.6 million. … The output also reports provision, asset-quality, capital, dividend, and shareholder-return statistics. | The revenue relation is explicitly stated but incorrectly linked to the low-rate environment. Broad coverage introduces many secondary indicators without preserving a focused driver-based narrative. |
| GCG | Total revenue increased $25.7 million versus 2014, driven by $11.6 million growth in non-interest income and $14.1 million growth in net interest income. … Net yield declined from 3.00% to 2.94% because average loans grew by $609.0 million, or 5.4%. … Asset quality improved as net loan charge-offs and non-performing assets declined. | R1 is explicitly preserved, but loan growth is incorrectly attached to the yield decline. The positive and negative components of the loan-interest driver chain are not coherently distinguished. |
| RAG-Sum | Total revenue increased $25.7 million, reflecting $11.6 million growth in non-interest income and $14.1 million growth in net interest income. … Net interest income increased $14.1 million due to higher trust fees, bank-card transaction fees, and mortgage-banking revenue. … The output additionally discusses impaired securities, private-equity investments, dividends, and shareholder return. | The revenue decomposition is preserved, but valid non-interest-income drivers are assigned to net interest income. Retrieval broadens coverage, yet also introduces secondary details and an incorrect cross-category relation. |
| TextRank | Total revenue increased $25.7 million due to growth in non-interest income of $11.6 million and growth in net interest income of $14.1 million. … Loan interest income increased $10.2 million because rates declined by 12 basis points, while higher loan balances partly offset the increase. … Net interest income totaled $664.0 million on a tax-equivalent basis, an increase of $15.4 million. | R1 is retained through sentence extraction, but the polarity of R2 is reversed. The output preserves source wording and values, but remains extractive and provides limited cross-sentence synthesis. |
| LoomSum | Total revenue increased $25.7 million due to $11.6 million growth in non-interest income and $14.1 million growth in net interest income. … Loan interest income increased $10.2 million as average loans grew by $609.0 million; these improvements are partly offset by a 12-basis-point decline in rates and higher non-interest expense. … Net yield declined from 3.00% to 2.94% because of continuing low interest rates. | LoomSum explicitly preserves R1–R3, including the opposing effects of loan growth and lower rates. It synthesizes the evidence into a coherent narrative without introducing excessive secondary detail. |
This example shows that covering the relevant quantities is not sufficient for analytical faithfulness. Direct-LC recovers several major indicators but incorrectly connects the revenue increase to the low-rate environment and introduces many secondary statistics. GCG and RAG-Sum preserve the revenue decomposition, yet attach valid drivers to the wrong outcomes: loan growth is associated with the yield decline, while non-interest-income drivers are assigned to net interest income. TextRank preserves substantial source wording but reverses the opposing effects of higher loan balances and lower rates. In contrast, LoomSum preserves the revenue decomposition, the complete loan-interest driver chain, and the explanation for the yield decline. In particular, it distinguishes the positive effect of loan growth from the offsetting effect of lower rates. The case therefore illustrates that the principal baseline failure is often not an unsupported individual fact, but an unsupported composition of otherwise valid facts and explanations.
| Source / Method | Evidence or Output Excerpt | Analysis |
| Source |
R1: 美元兑人民币中间价上行405BP至6.7503。
R2: 在岸即期汇率上行619BP至6.7670。 R3: 在岸即期汇率较中间价高167BP。 R4: 离岸即期汇率上行750.5BP至6.75875,较在岸低82.5BP。 配套表格则记录M2、M1和M0货币供应量。 |
The reference-relevant FX evidence occurs in one source section; the paired table is a cross-modal distractor. |
| Reference | 截至7月15日,美元兑人民币中间价较7月8日上行405BP至6.7503; 美元兑人民币即期汇率上行619BP至6.7670,较中间价高167BP; 美元兑离岸人民币即期汇率上行750.5BP至6.75875,较在岸即期汇率低82.5BP。 | Preserves all four target claims and both cross-rate comparisons. |
| Direct-LC | 央行在7月11日-15日实施逆回购150亿元到期150亿元,同时以1000亿元MLF 投放并有同额MLF到期,净投放为零;…货币市场方面,7月15日银行间 同业拆借加权利率较7月8日下行1.28基点至1.33%,质押式回购下行0.58基点 至1.31%;… | The facts are source-supported but come from non-target sections; all four reference-relevant FX claims are omitted. |
| GCG | 央行在7月11日-15日实现净投放为零,逆回购投放150亿元、到期150亿元, MLF投放1000亿元、到期1000亿元;…货币市场方面,7月15日银行间 同业拆借加权利率较7月8日下降1.28基点至1.33%,质押式回购利率下降 0.58基点至1.31%;… | The output is grounded in the source, but verbalization does not repair content selection; all target FX relations are omitted. |
| RAG-Sum | 外汇方面,美元对人民币中间价和即期汇率均走高,提示短期汇率承压。 | Retrieves the FX topic but omits every target value and both pairwise relations. |
| TextRank | 货币市场:截至7月15日,银行间同业拆借加权平均利率较7月8日变动 -1.28BP至1.33%,质押式回购加权平均利率较7月8日变动-0.58BP至 1.31%;… | Extracts a source-faithful sentence from the wrong section; all four reference-relevant FX claims are omitted. |
| LoomSum | 截止至7月15日,央行调整美元兑人民币中间价上行405BP至6.7503;在岸即期汇率上行619BP至6.7670, 较中间汇率价高167BP;离岸即期汇率上行750.5BP至6.75875,较在岸低82.5BP,呈跨市场小幅价差。 | Preserves every value, direction, and relation. |
USTT case. The USTT example in Table 10 illustrates a different failure mode: an output may remain source-grounded while providing poor coverage of the reference-relevant content. Direct-LC, GCG, and TextRank primarily select monetary-policy or money-market information from non-target sections. These statements are supported by the source, but they omit the four FX claims emphasized by the reference. RAG-Sum retrieves the correct topic and captures the coarse upward trend, yet removes the exact values and both cross-rate comparisons. LoomSum, by contrast, preserves the directions, values, and pairwise relations of all four target claims despite the presence of source-supported distractors. This case clarifies the role of G-Cov: source faithfulness alone does not guarantee that a summary covers the salient claims represented in the reference. Taken together, the FINDSum case demonstrates incorrect relation composition after partial content recovery, whereas the USTT case demonstrates grounded but incomplete content selection. LoomSum addresses both errors by explicitly aligning related evidence and organizing salient claims before generation.
Appendix E Evaluation Prompts
In this section, we provide the complete prompt used to decompose generated summaries into atomic claims for TGF evaluation below.