arXiv is now an independent nonprofit! Learn more
License: CC BY 4.0
arXiv:2609.00858v2 [cs.AI] 13 Sep 2026

Verifiable Disaster Storylines and Causal Knowledge Graphs: A Citation-Grounded Pipeline from Heterogeneous Humanitarian Sources

Conference: International Conference on Information Technology for Social Good; September 02–04, 2026; Pisa, ItalyInternational Conference on Information Technology for Social Good (GoodIT ’26), September 02–04, 2026, Pisa, ItalyDOI: 10.1145/3794786.3830759ISBN: 979-8-4007-2483-1/2026/09CCS: Computing methodologies Artificial intelligenceCCS: Computing methodologies Natural language processingCCS: Information systems Information extraction
Ivan Decostanzi email: ivan.decostanzi@isi.it Affiliation: ISI Foundation, Turin, Italy , Michele Ronco email: michele.ronco@ec.europa.eu Affiliation: European Commission, Joint Research Centre (JRC), Ispra, Italy , Sergio Consoli email: sergio.consoli@ec.europa.eu Affiliation: European Commission, Joint Research Centre (JRC), Ispra, Italy , Christina Corbane email: christina.corbane@ec.europa.eu Affiliation: European Commission, Joint Research Centre (JRC), Ispra, Italy , Lorenzo Bertolini email: lorenzo.bertolini@ec.europa.eu Affiliation: European Commission, Joint Research Centre (JRC), Ispra, Italy , Indaco Biazzo email: indaco.biazzo@ec.europa.eu Affiliation: European Commission, Joint Research Centre (JRC), Ispra, Italy , Daria Mihaila email: daria.mihaila@ec.europa.eu Affiliation: European Commission, Joint Research Centre (JRC), Ispra, Italy , Manuel Garcia-Herranz email: mherranz@unicef.org Affiliation: UNICEF, New York, USA , Felix Schwebel email: fschwebel@unicef.org Affiliation: UNICEF, New York, USA , Yelena Mejova email: yelena.mejova@isi.it Affiliation: ISI Foundation, Turin, Italy and Kyriaki Kalimeri email: kyriaki.kalimeri@isi.it Affiliation: ISI Foundation & UNICEF, Turin, Italy
© othergov
Abstract.

Effective humanitarian response depends on the rapid synthesis of heterogeneous, high-volume information sources — a task that routinely exceeds human analytical capacity in the critical early hours of a crisis. We present a pipeline that combines structured disaster records from EM-DAT with unstructured documents from ReliefWeb and the European Media Monitor (EMM) to produce source-grounded disaster storylines and causal knowledge graphs supporting situational awareness for responders and analysts. Using Retrieval-Augmented Generation, the pipeline extracts structured storylines — tabular event profiles covering 17 fields, from severity and key drivers to child-sensitive impact indicators — and constructs causal knowledge graphs where each node and edge is enriched with citation-grounded explanatory narratives, enabling full traceability back to primary sources. We evaluate the system on three diverse crisis use cases through a human evaluation involving 9 domain expert and 9 non-expert evaluators. Results confirm high retrieval precision, strong faithfulness of extracted causal relations, and a clear expert preference for citation-grounded components over ungrounded alternatives. The pipeline is designed to scale to the full EM-DAT catalogue, with the goal of publicly releasing a new collection of disaster stories complementing the EM-DAT database.

Keywords: 
Disaster Risk Management, Retrieval-Augmented Generation, Knowledge Graphs, Situational Awareness, Humanitarian Response

1. Introduction

In the immediate aftermath of a disaster, the ability to gather, synthesise, and act upon factual information is the difference between a coordinated response and a chaotic one. Disaster risk management (DRM) frameworks, utilised by entities such as the European Civil Protection and Humanitarian Aid Operations (ECHO), the United Nations Office for the Coordination of Humanitarian Affairs (UNOCHA), and the International Federation of Red Cross and Red Crescent Societies (IFRC), rely heavily on secondary data sources — news reports, situational updates, and field observations — to identify relief priorities and allocate resources effectively. However, modern disaster response faces an information paradox: while data availability has increased exponentially, the high volume and velocity of unstructured textual information routinely exceed human analytical capacity (Imran and others, 2020). Manually reviewing thousands of articles and reports to identify specific causal links — such as the relationship between a flood event and subsequent displacement or disruption of health services — is frequently intractable within the necessary operational timeframes. Standard disaster databases such as EM-DAT11 1 https://www.emdat.be/, while globally recognised, record aggregate impact statistics and systematically exclude smaller-scale events that do not meet reporting thresholds (Puime Pedra et al., 2026), leaving significant gaps in situational awareness.

Complementary textual sources — news articles, field assessments, and humanitarian coordination documents — can fill these gaps by capturing qualitative and quantitative details on impacts, exposed populations, local vulnerabilities, and cascading effects that are often absent from conventional event catalogues (Ronco et al., 2026). When paired with Retrieval-Augmented Generation (RAG) (Lewis et al., 2021), these unstructured sources can be transformed into coherent, factually anchored narratives. Knowledge graphs (KGs), which represent entities and their relationships as machine-readable triples, offer a particularly effective formalism for organising the resulting information and enabling structured reasoning over it (Chen et al., 2024; Yao et al., 2025; Tárraga et al., 2024). Recent work has demonstrated the viability of this approach across diverse hazard types: LLM-driven KG pipelines have been applied to earthquake emergency management (Yao et al., 2025), typhoon tracking (Huang et al., 2026), compound urban crises (Hao et al., 2025), and flood impact reporting (Colverd et al., 2023; Puime Pedra et al., 2026). A convergent finding across these systems is that grounding LLM outputs in structured or retrieved knowledge — rather than relying on parametric memory alone — is essential for factual accuracy in high-stakes operational contexts (Chen et al., 2024; Yao et al., 2025; Hao et al., 2025). Broader surveys confirm both the momentum and the remaining challenges of deploying LLMs responsibly in humanitarian settings, stressing the need for human oversight, transparency, and explainability (Xu et al., 2025; Lei et al., 2025; Rafiezadeh Shahi et al., 2026).

However, a persistent limitation of existing LLM-driven KG systems is that both generated narratives and graph elements are typically produced without explicit links to the source evidence from which they are derived. In high-stakes operational contexts — where every claim must be verifiable — this lack of traceability undermines user trust and limits practitioner adoption. Moreover, standard databases under-represent dimensions of particular concern to humanitarian actors, including child-specific vulnerabilities such as displacement, casualties, and loss of access to education and health services (Kadir et al., 2025; Hasbini et al., 2026).

Our prior work  (Ronco et al., 2026) introduced a pipeline that constructs causal KGs from over 3,000 global disaster events by combining EM-DAT records with news articles from the European Media Monitor (EMM) 22 2 https://knowledge4policy.ec.europa.eu/europe-media-monitor-emm_en through RAG-based extraction. Concretely, given an EM-DAT event record — e.g., Flood, Pakistan, June 2023 — the system retrieves relevant news articles and synthesises them into a structured event profile, which we term a storyline: a fixed-schema tabular summary capturing dimensions such as severity, key drivers, and impacts on critical services. From this storyline, a causal knowledge graph is extracted, encoding the event’s dynamics as subject–predicate–object triples constrained to causes and prevents relations. The present paper extends this framework in several directions:

  • •

    Humanitarian reports integration and enriched evidence base: We incorporate ReliefWeb33 3 https://reliefweb.int/ as a complementary source, linking EM-DAT events to ReliefWeb disaster records via GLIDE identifiers (Asian Disaster Reduction Center, 2024) and merging humanitarian field reports with news-derived documents. To our knowledge, this is the first pipeline to combine these two sources within an LLM-driven causal KG workflow.

  • •

    Full source traceability We introduce a Multi-Shot RAG strategy in which each storyline field is extracted through an independent retrieval-generation cycle, and a secondary validation step that generates explanatory text for every KG node and edge. Both mechanisms ground every output element in explicit citations to the underlying source documents, providing a verifiable audit trail that directly addresses concerns about transparency and explainability in automated DRM pipelines (Rafiezadeh Shahi et al., 2026).

  • •

    Child-sensitive impact dimensions: The storyline extraction schema is expanded to cover child-specific indicators, including displacement, casualties, and loss of access to education and health services.

  • •

    Exploration platform: We provide an interactive dashboard enabling exploration and analysis of the enriched causal knowledge graphs, enhanced source-grounded storylines, and child-specific risk indicators. Additionally, users can interact with the underlying database through natural language queries.

  • •

    Rigorous multi-level evaluation: We conduct rigorous human evaluation — involving 18 independent annotators — across three diverse crisis use cases spanning public health, natural disaster, and armed conflict, including a citation-level assessment of precision and recall of the generated attributions.

All source code is publicly available,44 4 https://github.com/idecost/StoryLine_KG and an interactive dashboard for exploring storylines, knowledge graphs, and child-specific risk indicators is accessible at https://idecost.github.io/StoryLine_KG/Viewer. The remainder of this paper is organised as follows: Section 2 details the pipeline, Section 3 the evaluation protocol, Section 4 the results, and Section 5 discusses limitations and future directions.

2. Methods

Flowchart of the pipeline showing data flow from EM-DAT, EMM, and ReliefWeb through RAG-based retrieval, storyline generation, and knowledge graph construction.

Figure 1. End-to-end pipeline for source-grounded disaster storyline and KG generation. (a) Main pipeline integrating EM-DAT, EMM, and ReliefWeb through RAG-based retrieval. (b) One-Shot vs. Multi-Shot storyline generation strategies. (c) Per-element RAG enhancement of the causal knowledge graph with citation-grounded narratives.Flowchart of the pipeline showing data flow from EM-DAT, EMM, and ReliefWeb through RAG-based retrieval, storyline generation, and knowledge graph construction.

This section details each component of the pipeline illustrated in Figure 1a, highlighting the extensions introduced relative to (Ronco et al., 2026). We begin with the integration of ReliefWeb as a complementary humanitarian data source (Section 2.1), then describe the expanded storyline extraction schema and the two generation strategies (Section 2.2), the causal KG construction (Section 2.3), the citation-grounded validation layer (Section 2.4), and the natural language query interface (Section 2.5).

2.1. ReliefWeb Integration

To enrich the information available for each disaster event, we extend the document retrieval pipeline by incorporating humanitarian reports from ReliefWeb. This integration is motivated by the complementary nature of ReliefWeb’s content, which aggregates situation reports, assessments, and humanitarian coordination documents. These sources often capture operational details and granular field data absent from news-based channels, thereby providing a more comprehensive operational picture of each disaster.

Consistent with the RAG paradigm, where model performance is enhanced by surfacing relevant external context, we align ReliefWeb data with our existing event-based structure. To link events across the two databases, we rely on GLIDE numbers (Asian Disaster Reduction Center, 2024), a standardized disaster identification system jointly developed by ADRC, CRED, OCHA/ReliefWeb, and UNDRR. This system assigns a unique structured code to each disaster, enabling unambiguous cross-referencing across humanitarian data systems. Each EM-DAT event is matched to its corresponding ReliefWeb disaster record using this identifier.

The full set of extracted text is then embedded using the BAAI/bge-m3 model (Multi-Granularity, 2024), with a chunking strategy of four sentences per chunk and a one-sentence overlap. Candidate chunks are ranked by relevance using BAAI/bge-reranker-v2-m3 (Multi-Granularity, 2024), and the top 15 chunks are retained. These ReliefWeb-derived documents are merged with the document set retrieved from EMM, collectively forming the updated evidence base for all downstream tasks.

2.2. Storyline Generation

As illustrated in Figure 1a, the first analytical stage of the pipeline transforms the document corpus into a structured storyline: a set of 17 fields that capture the key dimensions of a disaster event. These fields span standard impact indicators — damage analytics, event mapping, vital service continuity, and contextual threat assessments — as defined by established humanitarian frameworks (Aitsi-Selmi et al., 2015; Hasbini et al., 2026). Compared to (Ronco et al., 2026), we expand the extraction schema to include child-specific indicators covering casualties, displacement, and loss of access to education and health services, a dimension that remains systematically underreported in standard disaster databases (Kadir et al., 2025). The complete list of fields is provided in Table 1. Not all fields are necessarily populated for every event, as their availability depends on the completeness and detail of the underlying source documents. If no information is available for a given field, it is left as ’Unknown’.

Table 1. Storyline extraction elements grouped by category.
Category Elements
Hazard Profile & Risk Assessment Key information; Severity; Key drivers; Main impacts, exposure, and vulnerability; Likelihood of multi-hazard risks
Temporal & Situational Context Temporal details; Phase classification; Non-events
Children & Education Impact Impact on children; Impact on schools
Critical Services Disruption Health facilities disrupted; Water and sanitation access disrupted
Risk Governance & Best Practices Best practices for managing this risk
Recovery & Response Recommendations and supportive measures for recovery
Source Assessment Source type; Confidence level of information; Potential reporting bias

We implement two alternative strategies for generating storylines, detailed in Figure 1b. They share the same extraction schema and the same underlying model (Meta-Llama-3-70B-Instruct (Grattafiori et al., 2024)), but differ in how the document corpus is presented to the model and, crucially, in whether the outputs are traceable to their source evidence.

2.2.1. One-Shot Approach

The One-shot approach serves as the baseline in our evaluation. All documents retrieved for a given event are provided to the model as a single context block, and the 17 storyline fields are extracted jointly in one generation step (Figure 1b, Mode A). This strategy is straightforward and efficient, as it requires a single model invocation per event. However, because all fields are produced in a single pass over the full corpus, there is no mechanism to trace individual outputs back to their supporting passages.

2.2.2. Multi-Shot Approach (RAG)

To address the traceability limitation of the One-shot strategy, we introduce the Multi-shot approach (Figure 1b, Mode B). Rather than extracting all fields at once, each of the 17 storyline fields is treated as an independent information need. For each field, a fixed natural language query is formulated. This query is used to retrieve the most relevant passages from the unified document corpus via the RAG infrastructure (BGE-M3 embeddings and BGE-Reranker (Multi-Granularity, 2024)). The model then generates the field value based on the retrieved passages, and the supporting source documents are recorded alongside the output.

This decomposition yields two advantages. First, by narrowing the retrieval scope to a single information need at a time, the model receives more focused context for each field, reducing the risk that relevant details are overlooked or conflated within long input sequences. Second, each output element is naturally accompanied by explicit references to the source passages that support it, enabling full traceability from the structured storyline back to the original EMM and ReliefWeb documents. The 17 individually grounded fields are then assembled into a complete structured storyline with source attribution.

A comparative evaluation of the One-shot and Multi-shot approaches is presented in Section 3, where human annotators assess both factual accuracy and the perceived value of source attribution.

2.3. Causal Knowledge Graph Construction

The generated storyline, produced by either the One-shot or the Multi-shot approach, serves as input for the construction of a causal knowledge graph, following the text-to-graph methodology in (Ronco et al., 2026) (Figure 1a, Causal Knowledge Graph Extraction). By construction, graph elements — nodes and triplets — are compact abstractions detached from the textual evidence from which they originate. While their accuracy can be assessed through manual inspection against the source documents, as demonstrated in the evaluation of (Ronco et al., 2026) and in Section 4, this process is labour-intensive and does not scale to the thousands of elements produced across a large event catalogue. To address this, we introduce an automated citation-grounded validation layer (Section 2.4) that enriches every KG element with explanatory text and explicit source references, enabling practitioners to verify the factual basis of each node and edge without revisiting the full document corpus.

2.4. Citation-Grounded KG Validation

Our validation approach follows the “attribute-then-generate” paradigm (Slobodkin et al., 2024; Qian et al., 2024), adapted to the KG setting drawing on LLM-based fact verification methods (Xue and Zou, 2022; Shami et al., 2025; Huaman et al., 2020). As illustrated in Figure 1c, the framework operates as a secondary RAG-based pipeline applied independently to every node and triplet in the KG. For each element, the process proceeds in three steps:

(1) Query formulation. The LLM converts the KG element — a node label or a subject–predicate–object triplet — into a natural language question designed to retrieve supporting evidence. Unlike the Multi-shot storyline generation, where queries are fixed and predefined, here they are generated dynamically, since KG elements emerge from the extraction process and cannot be anticipated in advance (e.g., a node labelled “infrastructure damage” might yield the query “What infrastructure was damaged during the event?”).

(2) Grounded retrieval. The generated query is executed against the unified embedding space containing both ReliefWeb and EMM documents, using the same BGE-M3 and BGE-Reranker models employed throughout the pipeline.

(3) Narrative synthesis. The retrieved passages are used to generate a concise explanatory narrative that contextualises the graph element within the broader disaster account, rather than leaving it as an isolated structural statement.

Each generated narrative is accompanied by explicit citations to the source documents and their metadata, producing the Enhanced Causal Knowledge Graph shown in Figure 1a. This audit trail allows practitioners to verify whether each node and relationship is factually supported by official sources, directly mitigating the propagation of hallucinated content through the knowledge graph. An example of the output is presented in Figure 2(b,c,d).

2.5. Natural Language Database Interaction

The structured outputs described above — storylines and knowledge graphs — capture the information the pipeline is designed to extract. However, users may have questions that fall outside the scope of the predefined 17 fields or the graph structure, such as cross-event comparisons or queries about contextual factors not covered by the extraction schema. To support such exploratory analysis, we provide a natural language query interface to the underlying database. Users can formulate free-text questions about disaster events, impacts, and contextual factors. Each question is translated into a structured database query by the LLM, executed against the stored records, and the resulting answer is supported by retrieved textual evidence when available, ensuring that responses remain grounded in the source material. This functionality enables non-technical stakeholders — including field coordinators and policy analysts — to discover event-level details and cross-event patterns that may not surface through the structured pipeline outputs alone.

Screenshot of the interactive dashboard showing a
storyline excerpt, causal knowledge graph, source-grounded narrative,
citation popup, and natural language query interface.
Figure 2. Example output for the Haiti cholera case study. (a) Excerpt of the generated storyline with inline citations. (b) Causal knowledge graph derived from the storyline. (c) Source-grounded narrative associated with a selected edge. (d) Citation detail popup showing the retrieved source passage. (e) RAG-based Q&A answering a user query beyond the scope of the storyline and knowledge graph.Screenshot of the interactive dashboard showing a storyline excerpt, causal knowledge graph, source-grounded narrative, citation popup, and natural language query interface.

3. Evaluation

This section presents the evaluation of the pipeline across multiple dimensions. We describe the three crisis use cases selected for assessment (Section 3.1), as well as the employed human evaluation protocol (Section 3.2). The pipeline outputs — including the causal knowledge graphs, source-grounded storylines, and risk indicators — are publicly accessible through an interactive exploration dashboard at https://idecost.github.io/StoryLine_KG/Viewer/, which showcases a set of precomputed humanitarian events alongside the three use cases employed in this evaluation.

3.1. Use Cases

We evaluate the pipeline on three events spanning distinct humanitarian crisis typologies: a public health emergency, a sudden-onset natural disaster, and a protracted conflict. The cases were selected for their humanitarian significance, the richness of their documentation across EM-DAT and ReliefWeb, and their diversity along dimensions relevant to pipeline performance — geographic context, temporal dynamics, and reporting density.

Event 1 — Haiti Cholera Outbreak (2022). The resurgence of cholera in Haiti in October 2022 occurred against a backdrop of political instability, gang violence disrupting aid delivery, and collapsed water and sanitation infrastructure55 5 https://www.cdc.gov/mmwr/volumes/72/wr/mm7202a1.htm. The outbreak spread to all ten departments, resulting in tens of thousands of suspected cases, with children under five disproportionately affected. Its multi-causal nature — intersecting public health, governance failure, and conflict — makes it a demanding test of the pipeline’s ability to extract complex causal chains.

Event 2 —Hurricane Melissa, Dominican Republic (2025). Hurricane Melissa66 6 https://en.wikipedia.org/wiki/Hurricane_Melissa caused widespread displacement and infrastructure damage across the Caribbean in late 2025. We focus on the Dominican Republic, where the storm disrupted roads, schools, and health facilities across southern and eastern regions. The event is well-documented through OCHA and Dominican Civil Defence situation reports on ReliefWeb, providing a strong reference for evaluating KG faithfulness. Its rapid-onset, geographically bounded character offers a useful contrast to the other two cases.

Event 3 — Syria Conflict Escalation (Late 2024). The armed conflict escalation beginning in November 2024 --- culminating in the fall of Aleppo and the collapse of the Assad government --- triggered the displacement of over one million people within weeks, alongside widespread destruction of civilian infrastructure and acute food insecurity77 7 https://www.hrw.org/world-report/2024/country-chapters/syria. As the most information-dense and structurally complex of the three cases, it tests the pipeline’s capacity to handle rapidly evolving conflict settings with fragmented and heterogeneous source material.

3.2. Human Evaluation

We conduct an extensive human evaluation of the full pipeline across the three use cases. The evaluation is carried out by 18 independent annotators — 9 domain experts and 9 non-experts — none of whom are affiliated with the project or have conflicts of interest. Each evaluated item has been judged by nine different annotators. To assess the reliability of the collected annotations, we compute multiple complementary agreement measures. As the primary measure, we report Krippendorff’s α\alpha (Krippendorff, 2011), a standard metric for multi-annotator settings that accommodates ordinal scales and missing data. We additionally compute Fleiss’ κ\kappa as a cross-check; consistent with the literature (Zapf et al., 2016), we find that Fleiss’ κ\kappa yields values very close to Krippendorff’s α\alpha across all tasks, and therefore do not report it separately. We also compute, for each pair of annotators, the raw percentage agreement and the Prevalence-Adjusted Bias-Adjusted Kappa (PABAK) (Byrt et al., 1993), a variant of Cohen’s κ\kappa that corrects for prevalence and response-bias artefacts, and report the mean values across all annotator pairs. In summary, we report Krippendorff’s α\alpha, mean pairwise percentage agreement, and mean pairwise PABAK as our agreement measures.

Different evaluation stages, detailed in the following, are assigned to the most appropriate annotator profile: non-experts assess retrieval quality, KG text quality, and citation quality as these tasks require no domain-specific knowledge; experts evaluate storyline quality, KG faithfulness, and the overall system through the interactive dashboard.

Retrieval Quality.

While automatic RAG evaluation frameworks exist (Es et al., 2024; Salemi and Zamani, 2024), their application to the humanitarian domain remains largely unexplored. We therefore perform a precision-oriented human assessment: for each retrieved chunk (from ReliefWeb or EMM), non-expert annotators judge (a) relevance — whether the paragraph refers to the specific target disaster, mentioning the correct location, timeframe, or hazard type — and (b) informativeness — whether it contains concrete facts or analysis, as opposed to boilerplate, fragments, or garbled text.

Storyline Quality.

Domain experts compare the two storyline generation approaches — One-shot and Multi-shot — on a field-by-field basis across the 17 categories in Table 1. For each field, evaluators select one of the options: One-shot is better, Multi-shot is better, or Cannot tell. Additionally, evaluators provide an overall quality rating for each storyline on a 1–5 Likert scale. This design enables fine-grained analysis of which information categories benefit from each approach, as well as a holistic assessment of the two generation strategies.

Knowledge Graph Text Quality.

For the explanatory texts generated for each KG node and link (Section 2.4), non-expert annotators evaluate two independent dimensions. Relevance: whether the generated text discusses the correct concept (for nodes) or the correct relationship between source and target (for links), as opposed to being off-topic or addressing unrelated concepts. Informativeness: whether the text provides concrete, useful information about the disaster event, rated on a three-point scale — very informative (rich in facts, figures, or details), quite informative (some useful content but limited in depth), or not informative (empty, boilerplate, or tautological). Nodes and links are assessed in separate sub-tasks.

Knowledge Graph Faithfulness.

Following (Ronco et al., 2026), domain experts evaluate whether the extracted triples are supported by the storyline from which they are derived. Each triple (Source →\rightarrow Relation →\rightarrow Target) is classified as: Fully supported (explicitly stated), Partially supported (one entity or an implicit relation is present), Not present, or Cannot determine.

Citation Quality.

We additionally assess the citations attached to storyline fields (Section 2.2.2) and to KG node and link narratives (Section 2.4), adopting the citation recall and citation precision definitions established in prior work (Gao et al., 2023; Decostanzi et al., 2025). Given a generated text composed of statements s1,…,sns_{1},\ldots,s_{n}, where each sis_{i} cites a set of passages Ci={ci,1,ci,2,…}C_{i}=\{c_{i,1},c_{i,2},\ldots\}, citation recall evaluates whether sis_{i} is supported by CiC_{i} as a whole, while citation precision evaluates whether each individual citation ci,jc_{i,j} is relevant to sis_{i}; the two coincide when a statement carries a single citation. Non-expert annotators judge both dimensions on a three-point scale — fully supports, partially supports, does not support — applied to the individual citation (precision) or to the citation set (recall).

Expert System Assessment.

Finally, domain experts interact with the full exploration dashboard — including the knowledge graph visualisation, source-grounded storylines, and the natural language database interface (Section 2.5) — and provide a holistic evaluation of the system. Following established usability assessment practices (Brooke, ; Ronco et al., 2026), experts respond to structured questions covering the perceived usefulness of individual components, the utility of the system for humanitarian workflows, and overall trust in the generated outputs. Responses are collected on a five-point Likert scale, except for overall trust which is rated on a 0–10 scale.

4. Evaluation Results

All agreement metrics are computed pooling annotations across the three events; per-event breakdowns are available upon request. Each task is assessed by all 9 annotators of the relevant profile.

A recurring pattern across tasks is that Krippendorff’s α\alpha remains modest, while mean pairwise percentage agreement and PABAK are considerably higher. This is expected given the strong label imbalance present in most tasks—for instance, the large majority of retrieved paragraphs are judged relevant, and most KG texts are rated as relevant to their concept. As discussed in Section 3.2, both κ\kappa- and α\alpha-family statistics are known to be deflated under such prevalence conditions (Marzi et al., 2024), making PABAK and raw percentage agreement more informative indicators of true annotator consistency in this setting.

Retrieval Quality

A total of 110 paragraphs are evaluated across the three events. For relevance, annotators reach 83.1% mean pairwise agreement (PABAK =0.662=0.662, Krippendorff’s α=0.305\alpha=0.305, fair agreement (Krippendorff, 2019)). On average, 85.8±4.785.8\pm 4.7% of retrieved paragraphs are judged relevant to the target event, confirming the high precision of the retrieval stage. For informativeness, agreement is somewhat lower (67.2% mean pairwise, PABAK =0.344=0.344, α=0.187\alpha=0.187, fair), reflecting the greater subjectivity of distinguishing substantive content from boilerplate; 72.0±15.172.0\pm 15.1% of paragraphs are nonetheless deemed meaningful.

Storyline Quality

A total of 51 field comparisons are evaluated across the three events (17 fields per event). For each field, evaluators choose Multi-shot is better, One-shot is better, or Cannot tell; they additionally rate each storyline as a whole on a 1–5 scale. Agreement on this task is substantially lower than on retrieval and KG quality: pooled across all three events, annotators reach only 51.6% mean pairwise agreement (PABAK =0.033=0.033, Krippendorff’s α=0.097\alpha=0.097, slight agreement), reflecting the inherent subjectivity of comparing narrative outputs.

Results vary considerably across events. For Event 1, the Multi-shot storyline is preferred in 62.7±9.062.7\pm 9.0% of fields, One-shot in 19.6±6.819.6\pm 6.8%, and 17.6±5.917.6\pm 5.9% are judged Cannot tell. Event 3 shows a similar pattern with a stronger preference for Multi-shot (68.6±9.068.6\pm 9.0% vs. 23.5±17.623.5\pm 17.6%). Event 2 diverges markedly, driven by sharp annotator disagreement—one expert preferred Multi-shot in 94.1% of fields while another preferred One-shot in 58.8%. Overall, the Multi-shot storyline is preferred in 62.1±20.062.1\pm 20.0% of fields, One-shot in 26.1±19.526.1\pm 19.5%, and 11.8±9.311.8\pm 9.3% are judged Cannot tell.

Overall quality ratings confirm a mild preference for the Multi-shot approach. Multi-shot receives a mean rating of 3.67±0.713.67\pm 0.71 out of 5 across all nine annotators, compared to 2.78±0.832.78\pm 0.83 for One-shot, and seven out of nine annotators express an overall preference for Multi-shot. Annotators consistently value the source citations present in Multi-shot storylines, deemed essential for verifiability, though some fields are better served by the more concise One-shot output.

Knowledge Graph Text Quality

A total of 35 node texts and 28 link texts are evaluated. For KG node texts, 94.0±2.794.0\pm 2.7% of node texts are judged relevant to their concept (mean pairwise agreement 90.8%, PABAK =0.816=0.816, α=0.190\alpha=0.190). The low α\alpha relative to the high percentage agreement is consistent with the near-ceiling label distribution. Informativeness agreement is more moderate (60.0% mean pairwise, PABAK =0.200=0.200, α=0.302\alpha=0.302, fair); 89.6±6.389.6\pm 6.3% of node texts are rated at least quite informative—35.6±13.235.6\pm 13.2% very informative and 54.0±9.054.0\pm 9.0% quite informative—while only 10.5±8.110.5\pm 8.1% are rated not informative. The high standard deviations on the two positive categories suggest that disagreement concentrates on the boundary between very and quite informative rather than on whether a text is informative at all.

For KG link texts, relevance follows a similar pattern (92.3±4.992.3\pm 4.9% judged relevant; mean pairwise agreement 88.6%, PABAK =0.771=0.771, α=0.201\alpha=0.201, fair). Informativeness agreement is lower (55.5% mean pairwise, PABAK =0.109=0.109, α=0.174\alpha=0.174, slight), reflecting the inherent difficulty of assessing the informational depth of short relational descriptions; 94.4±9.594.4\pm 9.5% of link texts are rated at least quite informative—54.4±14.154.4\pm 14.1% very informative and 40.1±14.640.1\pm 14.6% quite informative—while only 5.6±5.15.6\pm 5.1% are rated not informative. As with node texts, the high variability across annotators on the two positive categories confirms that the distinction between very and quite informative is inherently subjective.

Knowledge Graph Faithfulness

A total of 28 triples are assessed across the three events, with no substantial variation observed across individual disasters. Overall, 86.7%86.7\% of triples are judged as supported by the storyline—56.1±26.556.1\pm 26.5% fully supported and 30.6±20.730.6\pm 20.7% partially supported—while only 9.2±9.49.2\pm 9.4% are rated as not present and 4.2±12.54.2\pm 12.5% as cannot determine. These results confirm that the large majority of automatically extracted causal relations are grounded in the generated narrative. Inter-annotator agreement is modest (Krippendorff’s α=0.222\alpha=0.222, mean pairwise agreement =54.2%=54.2\%, PABAK =0.083=0.083). As in previous tasks, disagreement concentrates on the boundary between the two positive categories rather than on whether a triple is supported at all.

Citation Quality

A total of 391 unique citations and 261 claims are assessed across 70 items and the three events. Pooled over all components, 91.791.7% of citations are at least partially relevant to the statement they support (79.379.3% relevant, 12.412.4% partially relevant; Krippendorff’s α=0.515\alpha=0.515, mean pairwise agreement =77.4=77.4%, PABAK =0.662=0.662), and 93.593.5% of claims are at least partially supported by their citation set (87.187.1% supported, 6.46.4% partially supported; α=0.465\alpha=0.465, mean pairwise agreement =86.4=86.4%, PABAK =0.796=0.796). Performance is uneven across components: KG node and link citations are consistently reliable (95.795.7% and 92.692.6% at least partially relevant; 99.199.1% and 98.498.4% of claims at least partially supported), whereas storyline citations show substantially more variability (71.771.7% and 65.465.4%, respectively). This appears to stem from over-citation: in the storyline the model attaches citations even when it lacks the evidence to answer, so short fields often carry several irrelevant sources. This reluctance to abstain deserves dedicated investigation, as citation errors disproportionately affect trust in this domain.

Expert System Assessment

Nine domain experts interact with the exploration dashboard and, as described in the following, they provide holistic evaluations across three dimensions: feature utility, operational impact, and overall trust.

Bar charts showing expert evaluation scores for feature
utility, operational impact, and overall trust.
Figure 3. Expert evaluation results from 9 domain experts. (a) Feature utility scores (1–5 scale) for the four main components of the dashboard: causal graphs, storyline fields, knowledge graph citations, and interactive KG queries. (b) Perceived operational impact (1–5 scale) across 3 dimensions: efficiency in crisis analysis workflows, enhancement of situational awareness, and usefulness for future disaster preparation. (c) Overall trust in the system, rated on a 0–10 scale.Bar charts showing expert evaluation scores for feature utility, operational impact, and overall trust.

For assessing the feature utility dimension, experts rated the perceived usefulness of each individual component of the pipeline on a 1–5 Likert scale (Figure 3a). The highest-rated components are KG Citations (M=4.33M=4.33, S​D=1.41SD=1.41) and Storyline & Fields (M=4.00M=4.00, S​D=1.32SD=1.32), indicating that experts find the most value in the structured narrative summaries and in the source-grounded textual descriptions attached to knowledge graph elements. Interactive KG Queries receive a moderately positive rating (M=3.75M=3.75, S​D=1.75SD=1.75, N=8N=8), though the high standard deviation signals polarised opinions—some experts appreciate the exploratory capability while others find it less intuitive. Notably, Causal Graphs receive the lowest rating (M=2.78M=2.78, S​D=1.09SD=1.09), suggesting that the automatically extracted causal structures are not yet perceived as sufficiently reliable or actionable by domain practitioners. This is consistent with Ronco et al. (Ronco et al., 2026), where the usefulness of knowledge graphs was similarly rated as only “somewhat useful” by most evaluators. It is precisely this limitation that motivated the design choice, introduced here, of enriching each KG node and link with automatically generated textual descriptions grounded in source citations (Section 2.4). The effectiveness of this strategy is reflected in the markedly higher ratings received by KG Citations, the highest-scored component in the evaluation, suggesting that anchoring graph elements to retrievable, source-backed explanations substantially increases their perceived value and compensates for the interpretability limitations of the bare graph structure.

For the operational impact dimension, experts rated the system’s potential contribution to humanitarian workflows, providing moderately positive scores across all three features. Efficiency and Situational Awareness both receive a mean of 3.673.67 (S​D=1.22SD=1.22 and 1.001.00 respectively), indicating that experts see tangible potential for the system to accelerate information synthesis and support a more comprehensive understanding of evolving crises. Future Preparation is rated slightly lower (M=3.33M=3.33, S​D=1.50SD=1.50), with the higher variance reflecting uncertainty about whether the system’s outputs—derived primarily from past and ongoing events—can effectively inform preparedness for future disasters. The consistent placement of all three features above the scale midpoint is encouraging, though the moderate absolute values suggest that the system is viewed as a promising complement to, rather than a replacement for, existing analytical workflows.

Finally, experts assessed overall trust in the system, yielding a mean of 6.56 (SD=2.06) on a 0–10 scale (Figure 3c), with all individual ratings falling within the 4–9 range. While these scores indicate adequate rather than high trust, they are encouraging for a fully automated pipeline requiring no human intervention and it is coherent with the positive trends reported across preceding evaluation dimensions. The spread of ratings reflects diverse expert expectations: residual reservations concentrate on causal graph quality and factual inconsistencies in storyline generation. The lowest score (4/10) is particularly instructive — the evaluator independently traced citations back to their original sources and found discrepancies between the reported information and the cited documents. Although such errors were infrequent, this case illustrates how even isolated citation inaccuracies can disproportionately erode trust in high-stakes humanitarian contexts, where every claim is expected to be verifiable.

5. Discussion and Conclusion

We presented an end-to-end pipeline for generating source-grounded disaster storylines and causal knowledge graphs from heterogeneous humanitarian sources, integrating ReliefWeb as a complementary evidence base, a Multi-Shot RAG strategy with full source traceability, a citation-grounded validation layer for every KG element, and child-sensitive impact dimensions. Evaluation across three diverse crisis use cases involving 18 independent annotators confirms the pipeline’s effectiveness: 85.8% retrieval precision, 86.7% of causal triples grounded in source material, and an overall expert trust of 6.56 out of 10. Citation-grounded components received the highest utility ratings (M=4.33M{=}4.33), validating the design choice of anchoring graph elements to retrievable source evidence. Experts highlighted the tool’s potential for rapid situational overview, cross-agency comparison of reported figures, and preliminary impact assessment. The evaluation also surfaces clear limitations, which map onto our future work. Causal graphs received the lowest utility rating (M=2.78M{=}2.78), perceived as oversimplified and potentially misleading without human validation, calling for refined graph complexity and confidence ranges flagging inter-source disagreement. The absence of temporal provenance was identified as a critical gap — storylines lack timestamps showing when figures were reported and how they evolved — motivating timestamped provenance with cross-source reconciliation. Experts further asked for sector-aligned structuring consistent with the humanitarian cluster system and for ingesting user-supplied or restricted-access documents.

Finally, storyline citations proved markedly less reliable than KG ones (71.771.7% vs. above 9292%), as the model keeps citing even when it lacks the evidence to answer — a behaviour that explicit abstention could mitigate.

On the operational side, we plan to run the pipeline over the full EM-DAT catalogue and publicly release the resulting dataset: a source-grounded (EMM and ReliefWeb), narrative-enriched version of EM-DAT offering contextual detail beyond the aggregate statistics currently available per event.

Acknowledgements.
The European Union owns the copyright of this work. © European Union, 2026. The authors acknowledge support from the Lagrange Project of the ISI Foundation, funded by Fondazione CRT

References

  • Aitsi-Selmi et al. (2015) A. Aitsi-Selmi, S. Egawa, H. Sasaki, C. Wannous, and V. Murray The sendai framework for disaster risk reduction: renewing the global commitment to people’s resilience, health, and well-being. International journal of disaster risk science 6 (2), pp. 164–176. Cited by: §2.2.
  • Asian Disaster Reduction Center (2024) Asian Disaster Reduction Center GLobal IDEntifier Number (GLIDE). Note: https://glidenumber.net/glide/public/search/search.jspAccessed: 2024-05-22 Cited by: 1st item, §2.1.
  • [3] J. Brooke SUS - A quick and dirty usability scale. Cited by: §3.2.
  • Byrt et al. (1993) T. Byrt, J. Bishop, and J. B. Carlin Bias, prevalence and kappa. Journal of clinical epidemiology 46 (5), pp. 423–429. Cited by: §3.2.
  • Chen et al. (2024) M. Chen, Z. Tao, W. Tang, T. Qin, R. Yang, and C. Zhu Enhancing emergency decision-making with knowledge graphs and large language models. International Journal of Disaster Risk Reduction 113 (104804). Note: _eprint: https://www.sciencedirect.com/science/article/pii/S2212420924005661 External Links: ISSN 2212-4209, Link, Document Cited by: §1.
  • Colverd et al. (2023) G. Colverd, P. Darm, L. Silverberg, and N. Kasmanoff FloodBrain: flood disaster reporting by web-based retrieval augmented generation with an llm. In 6th Workshop on Artificial Intelligence for Humanitarian Assistance and Disaster Response (NeurIPS 2023), External Links: Link Cited by: §1.
  • Decostanzi et al. (2025) I. Decostanzi, Y. Mejova, and K. Kalimeri A large-language-model framework for automated humanitarian situation reporting. arXiv preprint arXiv:2512.19475. Cited by: §3.2.
  • Es et al. (2024) S. Es, J. James, L. Espinosa Anke, and S. Schockaert RAGAs: Automated Evaluation of Retrieval Augmented Generation. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations, N. Aletras and O. De Clercq (Eds.), St. Julians, Malta, pp. 150–158. External Links: Document Cited by: §3.2.
  • Gao et al. (2023) T. Gao, H. Yen, J. Yu, and D. Chen Enabling large language models to generate text with citations. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 6465–6488. Cited by: §3.2.
  • Grattafiori et al. (2024) A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783. Cited by: §2.2.
  • Hao et al. (2025) H. Hao, X. Chen, Y. Chen, and N. Li Uncovering compound urban crises with large language model-assisted knowledge graph construction. International Journal of Disaster Risk Reduction 127 (105669). Note: _eprint: https://www.sciencedirect.com/science/article/pii/S2212420925004935 External Links: ISSN 2212-4209, Link, Document Cited by: §1.
  • Hasbini et al. (2026) L. Hasbini, L. G. Severino, M. M. De Brito, G. C. Gesualdo, A. M. Rotaru, D. N. Bresch, E. Mühlhofer, J. Wang, and T. M. N. Carvalho A database of disaster impacts in the Global South using Red Cross reports and Large Language Models. In Review (en). External Links: Link, Document Cited by: §1, §2.2.
  • Huaman et al. (2020) E. Huaman, E. Kärle, and D. Fensel Knowledge Graph Validation. arXiv. Note: arXiv:2005.01389 [cs] External Links: Link, Document Cited by: §2.4.
  • Huang et al. (2026) Y. Huang, Y. Xia, R. Tao, D. Jiao, X. Min, J. Zheng, Y. Jiang, W. Wu, and P. Du A LLM-based agent for the construction of typhoon knowledge graphs. Environmental Modelling & Software 197 (106856). External Links: ISSN 1364-8152, Link, Document Cited by: §1.
  • Imran et al. (2020) M. Imran et al. Using AI and social media multimodal content for disaster response and management. Information Processing & Management. Cited by: §1.
  • Kadir et al. (2025) A. Kadir, A. J. Stevens, E. A. Takahashi, and S. Lal Child public health indicators for fragile, conflict-affected, and vulnerable settings: a scoping review. PLOS Global Public Health 5. External Links: Document Cited by: §1, §2.2.
  • Krippendorff (2011) K. Krippendorff Computing Krippendorff’s alpha-reliability. Cited by: §3.2.
  • Krippendorff (2019) K. Krippendorff Content Analysis: An Introduction to Its Methodology - Fourth Edition. SAGE Publications, Thousand Oaks, California. External Links: Document Cited by: §4.
  • Lei et al. (2025) Z. Lei, Y. Dong, W. Li, R. Ding, Q. R. Wang, and J. Li Harnessing large language models for disaster management: a survey. In Findings of the Association for Computational Linguistics: ACL 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, pp. 14528–14551. External Links: Link, Document, ISBN 979-8-89176-256-5 Cited by: §1.
  • Lewis et al. (2021) P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W. Yih, T. Rocktäschel, S. Riedel, and D. Kiela Retrieval-augmented generation for knowledge-intensive nlp tasks. External Links: 2005.11401, Link Cited by: §1.
  • Marzi et al. (2024) G. Marzi, M. Balzano, and D. Marchiori K-Alpha Calculator–Krippendorff’s Alpha Calculator: A user-friendly tool for computing Krippendorff’s Alpha inter-rater reliability coefficient. MethodsX 12. External Links: Document Cited by: §4.
  • Multi-Granularity (2024) M. M. Multi-Granularity M3-embedding: multi-linguality, multi-functionality, multi-granularity text embeddings through self-knowledge distillation. arXiv preprint arXiv:2402.03216. Cited by: §2.1, §2.2.2.
  • Puime Pedra et al. (2026) M. Puime Pedra, S. Elkady, F.M. Villar-Rosety, et al. From headlines to databases: leveraging llms for structured disaster event extraction. International Journal of Data Science and Analytics 22, pp. 65. External Links: Document, Link Cited by: §1, §1.
  • Qian et al. (2024) H. Qian, Y. Fan, R. Zhang, and J. Guo On the capacity of citation generation by large language models. In China Conference on Information Retrieval, pp. 109–123. Cited by: §2.4.
  • Rafiezadeh Shahi et al. (2026) K. Rafiezadeh Shahi, M. M. Kuglitsch, J. B. Bove, M. Ronco, P. Ghamisi, Y. Sun, G. Duca, M. V. Gargiulo, A. Berlin, J. Jäpölä, F. Pharand-Deschênes, B. D. Malamud, B. Sakschewski, J. Rockstrom, and H. Kreibich Governing generative ai in disaster risk management. Note: Preprint, version 2 External Links: Link Cited by: 2nd item, §1.
  • Ronco et al. (2026) M. Ronco, L. Bandelli, L. Bertolini, S. Consoli, D. Delforge, A. Spadaro, M. Verile, and C. Corbane Disaster Storylines and Knowledge Graphs from Global News with Large Language Models and Retrieval-Augmented Generation. Scientific Data (en). External Links: ISSN 2052-4463, Link, Document Cited by: §1, §1, §2.2, §2.3, §2, §3.2, §3.2, §4.
  • Salemi and Zamani (2024) A. Salemi and H. Zamani Evaluating Retrieval Quality in Retrieval-Augmented Generation. arXiv. External Links: 2404.13781, Document Cited by: §3.2.
  • Shami et al. (2025) F. Shami, S. Marchesin, and G. Silvello Fact Verification in Knowledge Graphs Using LLMs. In Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR ’25, New York, NY, USA, pp. 3985–3989. External Links: ISBN 979-8-4007-1592-1, Link, Document Cited by: §2.4.
  • Slobodkin et al. (2024) A. Slobodkin, E. Hirsch, A. Cattan, T. Schuster, and I. Dagan Attribute first, then generate: locally-attributable grounded text generation. arXiv preprint arXiv:2403.17104. Cited by: §2.4.
  • Tárraga et al. (2024) J. M. Tárraga, E. Sevillano-Marco, J. Muñoz-Marí, M. Piles, V. Sitokonstantinou, M. Ronco, M. T. Miranda, J. Cerdà, and G. Camps-Valls Causal discovery reveals complex patterns of drought-induced displacement. iScience 27 (9), pp. 110628 (en). External Links: ISSN 2589-0042, Document, Link Cited by: §1.
  • Xu et al. (2025) F. Xu, J. Ma, N. Li, and J. C.P. Cheng Large language model applications in disaster management: an interdisciplinary review. International Journal of Disaster Risk Reduction 127 (105642). Note: _eprint: https://www.sciencedirect.com/science/article/pii/S2212420925004662 External Links: ISSN 2212-4209, Link, Document Cited by: §1.
  • Xue and Zou (2022) B. Xue and L. Zou Knowledge Graph Quality Management: a Comprehensive Survey. IEEE Transactions on Knowledge and Data Engineering, pp. 1–1 (en). External Links: ISSN 1041-4347, 1558-2191, 2326-3865, Link, Document Cited by: §2.4.
  • Yao et al. (2025) L. Yao, F. Ren, K. Du, and Q. Du From knowledge graph construction to retrieval-augmented generation: a framework for comprehensive earthquake emergency support. Geo-spatial Information Science 0 (0), pp. 1–21. Note: _eprint: https://doi.org/10.1080/10095020.2025.2514813 External Links: ISSN 1009-5020, Link, Document Cited by: §1.
  • Zapf et al. (2016) A. Zapf, S. Castell, L. Morawietz, and A. Karch Measuring inter-rater reliability for nominal data – which coefficients and confidence intervals are appropriate?. BMC Medical Research Methodology 16. External Links: Document Cited by: §3.2.