arXiv is now an independent nonprofit! Learn more
License: CC BY 4.0
arXiv:2608.00012v1 [cs.CL] 24 Jun 2026

Obshazard-bench: Benchmarking Multimodal Foundation Models for Real-Time Disaster Intelligence from Raw Earth Observation Streams

Fengxiang Wang    Qiuyang Yu    Yueying Li    Mingshuo Chen    Chengchi Fei    Kaiyi Xu    Lixin Gu    Wangxu Wei    Junchao Gong    Lipeng Ma    Jiong Wang    Fenghua Ling    Wenlong Zhang    Xue Yang Affiliation:  National University of Defense Technology, China. Shanghai Artificial Intelligence Laboratory, China. Shanghai Jiao Tong University, China. * Corresponding authors: Ben Fei and Long Lan.    Wenjing Yang    Ben Fei    Long Lan
Abstract

Multimodal Large Language Models (MLLMs) are increasingly used to interpret Earth observation data, yet their capability to support real-world disaster emergency response remains insufficiently evaluated. Existing remote sensing benchmarks largely rely on static, post-hoc, and expert-processed products, such as gridded reanalysis data, which are difficult to align with operational disaster scenarios where hazards evolve rapidly and decisions must be made under strict time constraints. To bridge this gap, we introduce Obshazard-bench, a real-time, observation-driven benchmark for evaluating disaster intelligence in MLLMs. Unlike image-centric or post-event benchmarks, Obshazard-bench directly integrates raw, high-frequency satellite sounding streams from diverse satellite sensors with concurrent ground-station observations, historical disaster records, and socio-economic indicators, bypassing delayed expert-processing and physical-inversion pipelines. The benchmark covers 8 major disaster categories and 28 sub-categories across more than 60 countries, incorporating over 120 historically documented extreme-event cases and thousands of lifecycle-oriented VQA samples. Moreover, Obshazard-bench further defines a three-stage evaluation taxonomy aligned with the operational disaster workflow: Predictive Crisis Anticipation for pre-disaster risk detection and early forecasting, Active Evolution Reasoning for in-situ disaster tracking and termination prediction, and Multi-faceted Impact Quantification for post-disaster magnitude deduction, humanitarian burden estimation, and socio-economic impact assessment. Experiments on representative general-purpose and Earth-focused foundation models reveal substantial limitations in transforming raw multi-channel physical observations into temporally grounded and decision-relevant disaster reasoning. The dataset and evaluation code are available at: Obshazard-bench

Refer to caption
Fig. 1: System architecture of the proposed Disaster Intelligence Benchmark (conceptualized in Section I and structured in Section III). The framework integrates: (Top-Left) Multi-source Earth system data coupling raw atmospheric sounding streams with ground measurements; (Top-Right) Global scale distribution covering over 60 countries; (Bottom-Left) Two-level hierarchical hazard taxonomy; and (Bottom-Right) Three-stage disaster lifecycle-oriented evaluation supported by representative visual question-answering examples.

I Introduction

Recent breakthroughs in Multimodal Large Language Models (MLLMs) and Earth Foundation Models (EFMs) have advanced spatial-temporal reasoning for Earth observation (EO) data, spanning multi-modal geospatial understanding [1, 2, 3, 4, 5] and global atmospheric forecasting [6, 7, 8, 9, 10]. Evaluating these systems requires rigorous benchmarking. While existing benchmarks like XLRS-Bench [11], OmniEarth-Bench [12], VRSBench [13], and RSRSD-5M [14] assess various remote sensing (RS) vision tasks, evaluating MLLMs in real-world disaster contexts remains difficult due to rapid event evolution and complex physical variables. Specially, existing benchmarks are limited in three main areas:

Operational Processing Latency

Disaster emergency response requires tracking atmospheric and surface hazards that evolve rapidly in real time. For timely warnings, MLLMs must process continuous, live, and high-frequency observation streams [15]. However, current benchmarks mostly use static, expert-processed datasets like orthorectified imagery or gridded reanalysis products. These products require time-consuming steps like expert cleaning and physical inversion, which introduce delays of days or weeks (Fig. 2). This latency makes them impractical for active emergency response.

Geographical and Hazard Scope

Natural disasters vary widely across different climates and regions. To test model generalization, a benchmark needs global coverage with diverse hazard types. Many existing disaster datasets focus on specific areas or limited hazard types, often concentrating on post-event structural damage.

Disaster Lifecycle-Oriented Evaluation

Emergency decisions are made continuously as a disaster develops. However, evaluating models only on post-event imagery isolates the event from its temporal context. Practical disaster management requires MLLM that can perform reasoning throughout the event lifecycle. Existing datasets do not support this temporal dependency.

To address these challenges, we introduce a real-time, observation-driven disaster intelligence benchmark that integrates continuous satellite observations, ground station data, authentic extreme-event records, and socio-economic indicators. Instead of using pre-processed visual maps, our framework replicates the professional meteorologist’s workflow by directly providing raw, multi-channel physical sounding profiles from core weather satellite sensors: the Advanced Micowave Sounding Unit-A (AMSU-A, 15 channels), the High-Resolution Infrared Radiation Sounder (HIRS, 20 channels), and the Microwave Humidity Sounder (MHS, 5 channels) [16, 17]. In operational centers, meteorologists analyze these exact vertical sounding profiles to detect atmospheric instability and moisture transport. Spanning 8 major disaster categories and 28 sub-categories across more than 60 countries (encompassing over 120 historical extreme-event cases), our benchmark evaluates whether foundation models can directly ingest raw sensor streams and transform them into structured, decision-relevant reasoning across three progressive lifecycle phases: pre-event anomaly detection, during-event intensity tracking, and post-event consequence reasoning. Ultimately, this setup evaluates whether MLLMs can bypass delayed processing pipelines to support practical decision-making across the disaster lifecycle—from issuing early alerts to guiding real-time tracking and accelerating post-event recovery (Fig. 2).

TABLE I: Quantitative Comparison with Existing Disaster-Specific Benchmarks
Dataset Disaster Types Real Events VQA Samples Covered Countries Raw Stream
CrisisMMD [18] 7 7 5 ×\times
xBD [19] 6 19 6 ×\times
Sen12Floods [20] 1 11 11 ×\times
FloodNet [21] 1 1 4,500 1 ×\times
RescueNet [22] 1 1 1 ×\times
CRASAR-U-D. [23] 2 10 1 ×\times
DisasterM3 [24] 10 36 26,988 5 ×\times
ZeShot-VQA [25] 4 19 15,000 6 ×\times
DisasterVQA [26] 13 15 4,405 1 ×\times
DORA [27] 10 45 515 5 ×\times
BRIGHT [28] 7 14 14 ×\times
EBD [29] 12 12 8 ×\times
MONITRS [30] 8 9,996 1 ×\times
Obshazard-bench 28 127 4,202 62 \checkmark

The main contributions of this paper are summarized as follows:

  • Direct Sensor-Stream Evaluation Paradigm: We introduce a disaster benchmark built on raw, live-streaming satellite and ground data, bypassing the delayed processing pipelines of traditional reanalysis datasets.

  • Global-scale Multi-hazard Testbed: The benchmark provides global coverage with 8 major disaster categories and 28 sub-categories across more than 60 countries, incorporating over 120 historical extreme-event cases to test model generalization.

  • Lifecycle-aligned Evaluation Taxonomy: We organize tasks into a three-stage temporal framework spanning pre-event predictive anticipation, during-event active evolution tracking, and post-event impact quantification, enabling continuous evaluation.

  • Systematic Functional Gap Assessment: We evaluate several open- and closed-source foundation models on our benchmark tasks. Our analysis systematically exposes their functional deficiencies and performance bottlenecks in translating raw, multi-channel physical signals into temporal and logical reasoning, helping identify key limitations in driving practical disaster mitigation applications.

Refer to caption
Fig. 2: Conceptual comparison between traditional batch-style evaluation and our proposed real-time streaming paradigm. The comparison is organized across three critical levels: (Row 1) Data Handling, contrasting delayed offline archives with direct, raw satellite feeds (AMSU-A, HIRS, MHS) bypassing expert processing; (Row 2) Temporal Coverage, highlighting sparse retrospective analysis versus dense, continuous lifecycle monitoring; and (Row 3) Operational Decision Utility, showing how our paradigm enables proactive alerts and notifications pre-event, continuous real-time tracking and guidance in-situ, and faster response and recovery strategies post-event.

II Related Work

II-A Foundational Remote Sensing Datasets and Benchmarks

Remote sensing (RS) datasets have evolved from early archival time-series analysis [31, 32] to large-scale multimodal benchmarks. Recent works such as XLRS-Bench [11], VRSBench [13], and RSRSD-5M [14] have advanced ultra-high-resolution (UHR) remote sensing perception. However, most of these datasets are still built around static optical imagery. This limits their use in disaster scenarios, where cloud interference is common and where atmospheric structure cannot be captured from surface appearance alone [33, 34].

Several benchmarks have started to include temporal or multi-modal observations. DynamicEarthNet [35] provides daily multi-spectral sequences, and FoMo-Bench [36] introduces multi-modal forest monitoring. Other specialized datasets, such as TreeFinder [37] and SegMunich in SpectralGPT [4], further enrich remote sensing data coverage. Yet these datasets mainly describe surface land-cover or vegetation changes. They do not focus on fast-changing physical processes in the atmosphere, which are critical for many extreme events.

Recent Earth foundation models and benchmarks also point to a broader trend toward cross-sphere and heterogeneous Earth-system reasoning. Examples include AnySat [38], TerraMind [39], OmniEarth-Bench [12], OmniGAIA [40], Earth-Agent [41], and TerraBench [42]. These works expand Earth AI evaluation beyond isolated visual perception, covering broader data sources, cross-domain reasoning, and in some cases tool-augmented agent workflows. However, high-frequency physical observations, such as vertical temperature and humidity sounding signals, are still underrepresented in benchmark design. Existing benchmarks also rarely organize such observations around disaster response stages. Obshazard-bench complements these efforts by focusing on raw sounding streams and lifecycle-oriented disaster reasoning.

Obshazard-bench addresses this gap by using high-frequency sounding streams from AMSU-A, HIRS, and MHS. These instruments provide multi-channel physical observations of the atmosphere, including AMSU-A with 15 channels and HIRS with 20 channels. Compared with optical-centric benchmarks, Obshazard-bench offers a complementary physical view of disasters. It allows models to reason over atmospheric signals that are not visible in traditional surface imagery.

II-B Benchmarks for Extreme Disaster Response

Remote sensing benchmarks for disaster response have moved from basic perception toward more complex reasoning. Early datasets such as xBD [19], RescueNet [43], and CRASAR-U-DRoIDs [23] mainly support building damage assessment from bi-temporal satellite images or unmanned aerial vehicle (UAV) imagery. These datasets are important for post-disaster mapping, but they are less suited to early warning or event evolution analysis.

Later benchmarks extend the task scope and hazard coverage. DisasterM3 [24] and RSCC [44] introduce multi-modal reasoning and change captioning. ZeShot-VQA [25] and DisasterVQA [26] explore disaster-related visual question answering. Other works, including Anomaly-CD [45], Shield [46], BRIGHT [28], EBD [29], and MONITRS [30], further improve multi-hazard coverage and dataset scale.

Despite these advances, most disaster benchmarks still rely on static or bi-temporal imagery. Their tasks are often centered on damage classification, change detection, or basic VQA. They rarely use raw high-frequency physical streams, and they do not fully reflect the operational needs of disaster response. In practice, disaster management requires more than post-event perception. Models must support early warning, track event evolution, and estimate humanitarian or socio-economic impacts.

Meteorological foundation models provide useful evidence that high-frequency physical data can support such reasoning. Models such as Aurora [47], Aardvark Weather [48], and Prithvi WxC [49] show strong potential for weather forecasting and early warning. However, these works mainly focus on model development and prediction tasks. They are not designed as benchmarks for evaluating multimodal disaster intelligence across multiple hazards and lifecycle stages.

This leaves an operational gap between experimental Earth AI systems and professional disaster management [17, 5]. Recent reviews also note that current systems still struggle to assimilate high-frequency physical data for response-oriented reasoning [50, 51, 16]. Obshazard-bench is designed to address this gap. It covers 8 major disaster categories and 28 sub-categories across 62 countries. It aligns daily physical observation sequences with 4,202 lifecycle-oriented VQA samples, providing a realistic testbed for evaluating disaster-oriented MLLMs.

III Obshazard-bench

III-A Data Sources and Processing

Obshazard-bench is built upon a multi-source architecture that connects verified disaster records with high-frequency geophysical observations. We use the EM-DAT international disaster database 11 1 https://www.emdat.be/ as the event metadata source, from which historical disaster cases are selected and organized. To focus on rapidly evolving extreme events, we retain disasters with an active duration no longer than 9 days. The resulting benchmark contains 127 verified historical extreme-event cases across 62 countries, covering 8 major disaster categories and 28 disaster sub-categories.

Departing from conventional optical-centric or post-event disaster benchmarks, Obshazard-bench uses raw satellite sounding streams as the core observational input. Specifically, we incorporate a complementary tri-sensor constellation consisting of AMSU-A with 15 channels, HIRS with 20 channels, and MHS with 5 channels. These instruments provide multi-channel atmospheric sounding observations related to temperature, humidity, and other physical conditions, enabling models to reason from raw geophysical signals rather than relying only on surface-level visual appearance. In addition to satellite streams, the benchmark integrates disaster metadata, ground-station observations, historical disaster records, and socio-economic indicators when constructing lifecycle-oriented disaster reasoning tasks.

The processing procedure is designed to preserve the raw-stream nature of the observations while making heterogeneous sources usable for multimodal evaluation. For each selected disaster event, satellite observations are temporally aligned with the event window and spatially associated with the affected region. The benchmark does not replace raw observations with delayed reanalysis products, expert-interpreted maps, or physically inverted variables. Instead, it organizes raw multi-channel sensor observations into event-centric samples, allowing Obshazard-bench to evaluate whether MLLMs can transform raw physical observations into temporally grounded and decision-relevant disaster reasoning.

TABLE II: Disaster taxonomy and sample distribution (single column).
L1 Category L2 Sub-categories Samples Ratio Events
Flood Flash flood 95 2.3% 5
Coastal flood 68 1.6% 2
General flood 195 4.6% 5
Riverine flood 165 3.9% 5
Mass movement (wet) Landslide 66 1.6% 2
Mudslide 123 2.9% 4
Wildfire Land fire 170 4.0% 5
Forest fire 159 3.8% 5
General wildfire 137 3.3% 4
Earthquake Tsunami 200 4.8% 5
Ground movement 217 5.2% 5
Storm Hail 174 4.1% 5
Tornado 195 4.6% 5
Blizzard/Winter storm 145 3.5% 5
Derecho 128 3.0% 5
Extra-tropical storm 150 3.6% 5
General storm 184 4.4% 5
Lightning/Thunderstorm 160 3.8% 5
Sand/Dust storm 130 3.1% 5
Severe weather 190 4.5% 5
Storm surge 191 4.5% 5
Tropical cyclone 205 4.9% 5
Volcanic activity Ash fall 135 3.2% 5
Lava flow 123 2.9% 4
General activity 163 3.9% 5
Pyroclastic flow 29 0.7% 1
Heat-wave Heat-wave 165 3.9% 5
Cold-wave Cold-wave 140 3.3% 5
Total 28 sub-categories 4,202 100.0% 127

III-B Benchmark Construction Pipeline

Refer to caption
Fig. 3: Benchmark construction pipeline of Obshazard-bench. The benchmark is constructed through four stages: (1) disaster event screening from EM-DAT records with spatial and temporal metadata, (2) retrieval of multi-source observations including AMSU-A, HIRS, MHS, and auxiliary contextual records, (3) event-centered spatial-temporal alignment that associates observations with disaster locations and time windows, and (4) construction of standardized benchmark cases containing observation sequences, metadata, and quality-control information. This pipeline preserves the raw-stream characteristics of satellite sounding observations while organizing heterogeneous sources into event-centric benchmark instances.

Obshazard-bench is constructed through a multi-stage pipeline that transforms heterogeneous disaster records and raw satellite observations into reasoning-oriented benchmark samples. The pipeline begins with disaster event metadata. For each event, we extract the disaster identifier, start date, end date, longitude, latitude, and disaster type. Events with durations longer than 9 days are filtered out, ensuring that the benchmark focuses on rapidly evolving disaster processes. To support pre-disaster reasoning, the observation window is extended by shifting the start date 14 days earlier than the recorded event onset.

For each retained event, raw satellite observations are collected within the corresponding temporal window. The current benchmark uses AMSU-A, HIRS, and MHS observations. For each day in the event-centered time span, the pipeline reads observations within a 24-hour window, typically from 00:00 to 23:00 at hourly intervals. These observations are retrieved from the storage system according to the corresponding sensor type, time stamp, and event location.

After temporal retrieval, the observations are spatially cropped around the event center. Given the event longitude and latitude, a local region is extracted using a padding size of 72 pixels, resulting in a 144×144144\times 144 spatial crop for each channel and each time step. This produces a raw multi-channel observation tensor over the event region.

Because satellite observations within a 24-hour window may contain multiple valid measurements for the same spatial location, we apply a temporal compression step to organize the raw stream into a daily observation layer. For each channel and each pixel, the pipeline scans the hourly observations and retains the latest valid observation within the 24-hour window. This step removes the hourly dimension while preserving the raw sensor values, producing a tensor of shape (C,H,W)(C,H,W) for each sensor and each day. Importantly, this operation does not convert the observations into reanalysis variables or expert-derived physical products; it only organizes the raw sensor stream into a consistent event-level format.

Finally, each processed observation is saved together with its event metadata, including the disaster identifier, start date, end date, longitude, and latitude. The resulting event-centric observations are then used to construct lifecycle-oriented VQA samples covering pre-disaster anticipation, active event evolution, and post-disaster impact quantification. In total, Obshazard-bench contains 4,202 benchmark samples derived from 127 historical disaster events.

III-C VQA Task Generation

After building event-centric observation samples, we convert them into visual question answering (VQA) instances. Each instance is defined by three elements: a disaster event, an observation timestep, and a target variable. This makes each sample traceable to its source event, raw observations, and ground-truth record.

For each event, we first build a temporal sequence around the event date. We use t14t-14, t7t-7, and t3t-3 for pre-event observations. These correspond to 14, 7, and 3 days before the recorded event onset. We use t15t_{15} for event-stage observations. When available, we also use tmax1t_{\max-1} and tmaxt_{\max} for later event-related observations. Each timestep is paired with raw multi-channel observations from AMSU-A, HIRS, and MHS. Event metadata, such as location, event date, and disaster category, is also attached.

We then generate VQA samples by matching each timestep with eligible target variables. Pre-event timesteps are used for early-warning targets, including disaster occurrence, disaster type, arrival time, and initial duration. Event-stage and later timesteps are used for evolution and impact-related targets, including recovery time, magnitude, total deaths, total affected, number affected, number injured, number homeless, and Consumer Price Index (CPI)-related economic disruption. In this way, one disaster event can produce multiple VQA samples across different timesteps and target variables.

Each VQA instance contains four parts: visual input, question, answer space, and ground-truth answer. The visual input comes from the raw multi-channel observation tensor. The question is generated from task-specific templates and filled with event-related information. The answer space depends on the target type. Categorical targets use discrete labels. Temporal, numerical, and impact-related targets are converted into standardized answer intervals. The ground-truth answer is taken from verified disaster metadata and outcome records.

We also record metadata for each sample, including disaster category, disaster sub-category, lifecycle stage, subtask name, timestep, event identifier, and answer type. These fields are used for evaluation and analysis, not as shortcut inputs to the model. With this event–timestep–target construction, Obshazard-bench contains 4,202 VQA samples from 127 historical disaster events.

III-D Task Dimensions

TABLE III: Lifecycle reasoning dimensions and sample distribution in ObsHazard-Bench.
Predictive Crisis Anticipation Active Evolution Reasoning Multi-faceted Impact Quantification
Sub-task Samples Sub-task Samples Sub-task Samples
Risk Detection 381 Active Termination Prediction 213 Magnitude Deduction 185
Type Classification 381 Humanitarian Burden 1,670
Arrival Time Prediction 381 Socio-economic Intensity 610
Initial Duration Prediction 381
Total 1,524 Total 213 Total 2,465
Ratio 36.3% Ratio 5.1% Ratio 58.7%

Obshazard-bench defines its evaluation space through the Cartesian combination of two dimensions: the disaster taxonomy dimension and the lifecycle reasoning dimension. The disaster taxonomy dimension specifies the physical hazard type, while the lifecycle reasoning dimension specifies the operational stage and reasoning target. This design enables evaluation not only across different disaster categories, but also across different phases of disaster response.

Disaster Taxonomy Dimension. The disaster taxonomy dimension follows a two-level hazard hierarchy. At the first level, Obshazard-bench covers 8 major disaster categories: Flood, Mass movement (wet), Wildfire, Earthquake, Storm, Volcanic activity, Heat-wave, and Cold-wave. At the second level, these categories are divided into 28 disaster sub-categories. Flood includes flash flood, coastal flood, general flood, and riverine flood. Mass movement (wet) includes landslide and mudslide. Wildfire includes land fire, forest fire, and general wildfire. Earthquake includes tsunami and ground movement. Storm includes hail, tornado, blizzard or winter storm, derecho, extra-tropical storm, general storm, lightning or thunderstorm, sand or dust storm, severe weather, storm surge, and tropical cyclone. Volcanic activity includes ash fall, lava flow, general volcanic activity, and pyroclastic flow. Heat-wave and Cold-wave are retained as individual sub-categories.

Lifecycle Reasoning Dimension. The lifecycle reasoning dimension is designed according to the operational workflow of disaster management and is divided into three stages.

Predictive Crisis Anticipation

This stage evaluates whether a model can identify potential disaster risks before or near the onset of an event. Representative tasks include risk detection, disaster type classification, arrival time prediction, and initial duration prediction. These tasks require the model to interpret early physical signals from raw sounding streams and associated contextual information, rather than simply recognizing visible post-event damage.

Active Evolution Reasoning

This stage focuses on reasoning during the ongoing development of a disaster. The core task is Active Termination Prediction, which requires the model to utilize real-time sequential sounding data to forecast exactly when the ongoing event will conclude. This setting reflects the operational need for continuous monitoring and real-time guidance, where the key question is not only what disaster is happening, but also how the event is changing over time.

Multi-faceted Impact Quantification

This stage evaluates the model’s ability to connect physical disaster signals with downstream consequences. It contains three complementary task groups. (1) Magnitude Deduction: inferring the intrinsic severity or intensity level of a disaster event directly from raw, multi-source observations and event-related evidence. (2) Humanitarian Burden: quantifying social consequences, including Total Deaths, Total Affected, Number Injured, and Number Homeless. (3) Socio-economic Intensity: estimating economic disruption, including impacts reflected through disaster-induced changes in consumer and regional economic indicators. Together, these tasks test whether models can move beyond event recognition and transform raw observational evidence into consequence-oriented disaster reasoning.

By crossing the disaster taxonomy dimension with the lifecycle reasoning dimension, Obshazard-bench forms a structured evaluation matrix over hazard types and disaster-response stages. This matrix allows the benchmark to diagnose whether a model’s capability is specific to certain hazards, certain lifecycle stages, or their interaction, rather than relying only on a single overall performance score.

III-E Evaluation Metrics

Following the evaluation protocol of Earth AI [5], Obshazard-bench adopts a task-aware scoring scheme for heterogeneous disaster reasoning tasks. Each model response is normalized into a comparable answer format before scoring. For categorical tasks, such as risk detection and disaster type classification, predictions are evaluated according to their semantic or lexical consistency with the ground-truth answer. For numerical and interval-based tasks, including arrival time prediction, duration prediction, recovery time prediction, magnitude deduction, humanitarian burden estimation, and socio-economic intensity estimation, the score is computed according to the distance between the predicted value or interval and the ground truth. All scores are normalized to the range of [0,1][0,1].

To avoid bias toward sample-rich disaster categories or frequently occurring subtasks, we aggregate performance using an equal-weight protocol across timesteps, subtasks, lifecycle stages, and disaster categories. This metric design enables fair comparison across the three lifecycle stages of Obshazard-bench: Predictive Crisis Anticipation, Active Evolution Reasoning, and Multi-faceted Impact Quantification.

III-F Quality Control

To ensure the reliability of Obshazard-bench, we conduct quality control from both automated data construction and expert manual review.

Automated Data Quality Control: During data generation, we automatically check the consistency between disaster metadata, satellite observations, and constructed benchmark instances. Specifically, the pipeline verifies event time windows, location coordinates, satellite data availability, channel completeness, missing-value patterns, and the correspondence between generated samples and source records. Cases that fail these checks are removed or regenerated to ensure that each sample is grounded in valid event metadata and aligned multi-source observations.

Expert Manual Quality Control: In addition to automated checks, domain experts manually inspect the constructed cases and task annotations. They review whether the selected observations are relevant to the target disaster, whether the generated questions are scientifically meaningful, and whether the answers are consistent with disaster records and available evidence. Ambiguous, weakly supported, or scientifically invalid samples are revised or excluded from the final benchmark.

IV Experiments

TABLE IV: Overall performance across disaster categories. We report equal-weight scores for eight major disaster categories and the average score across categories. Mass Mov. denotes Mass Movement, and Avg. denotes average. The best result in each column is highlighted in bold.
Model Earthquake Flood Storm Wildfire Cold-wave Heat-wave Mass Mov. Volcanic Avg.
Claude Opus 4.8 [52] 0.2252 0.2389 0.2684 0.2377 0.1574 0.2839 0.2719 0.2557 0.2424
GPT-5.5 [53] 0.3501 0.2318 0.2626 0.3681 0.1376 0.2583 0.3496 0.3081 0.2833
Kimi-k2.6-1T [54] 0.2128 0.2344 0.2406 0.2194 0.2690 0.3022 0.2461 0.2805 0.2506
Qwen3.5-397B-A17B [55] 0.2091 0.1946 0.2154 0.1932 0.1422 0.2362 0.1281 0.2192 0.1922

IV-A Experimental Setup

We evaluate four representative foundation models on Obshazard-bench, including three closed-source frontier models and one open-source large language model. The closed-source models are Claude Opus 4.8 [52], GPT-5.5 [53], and Kimi-k2.6-1T [54]; the open-source model is Qwen3.5-397B-A17B [55]. All models are evaluated under a zero-shot setting with the same input representation and uniform prompts for fair comparison. As described in Section III-E, we report normalized task-aware scores following the Earth AI evaluation protocol [5]. Additional implementation details, prompt templates, and model configurations are provided in the Appendix.

IV-B Main Results

We evaluate representative general-purpose and domain-adapted models on Obshazard-bench, including frontier MLLMs, an open-source large language model, and a fine-tuned disaster-oriented model. All tasks are evaluated under the three lifecycle stages defined in Section III-D: Predictive Crisis Anticipation (PCA), Active Evolution Reasoning (AER), and Multi-faceted Impact Quantification (MIQ). Scores are normalized to [0,1][0,1] and aggregated with an equal-weight protocol across timesteps, subtasks, lifecycle stages, and disaster categories, preventing the final score from being dominated by sample-rich tasks or high-frequency disaster types. Following the evaluation practice in Earth AI [5], we use task-aware soft scoring rather than strict binary matching, where numeric answers are graded according to their distance from the ground truth and categorical or textual answers receive partial credit when semantically or lexically close.

Table IV reports the overall performance across the eight disaster categories. Obshazard-bench remains challenging for all evaluated models: even the best-performing model achieves an average score below 0.30. Among the evaluated general-purpose models, GPT-5.5 achieves the highest average score, followed by Kimi-k2.6-1T and Claude Opus 4.8. While GPT-5.5 provides the strongest overall performance, the remaining models exhibit complementary strengths across different disaster categories and lifecycle stages. These results show that current models still struggle to transform raw multi-channel Earth observations into lifecycle-aware disaster intelligence.

TABLE V: Performance across lifecycle stages. PCA denotes Predictive Crisis Anticipation, AER denotes Active Evolution Reasoning, and MIQ denotes Multi-faceted Impact Quantification. The best result in each column is highlighted in bold.
Model PCA AER MIQ Avg.
Claude Opus 4.8 [52] 0.4697 0.0278 0.2297 0.2424
GPT-5.5 [53] 0.3443 0.2078 0.2977 0.2833
Kimi-k2.6-1T [54] 0.3756 0.1488 0.2275 0.2506
Qwen3.5-397B-A17B [55] 0.3084 0.1092 0.1591 0.1922

IV-C Capability Differences across Lifecycle Stages and Disaster Types

To better understand model behavior, we analyze model performance along two benchmark dimensions: lifecycle stage and disaster type. Table V shows that different models exhibit distinct lifecycle-stage capability profiles. Claude Opus 4.8 achieves the highest score on Predictive Crisis Anticipation, indicating strong capability in pre-disaster risk detection and early forecasting. However, its performance drops sharply on Active Evolution Reasoning, suggesting that strong pre-event reasoning does not necessarily translate into reliable reasoning about event evolution or recovery time.

GPT-5.5 shows the strongest performance on Multi-faceted Impact Quantification and achieves the highest overall score among the evaluated models, indicating better capability in connecting event-related evidence with consequence-level estimates. Kimi-k2.6-1T exhibits a capability profile between GPT-5.5 and Claude Opus 4.8: it performs strongly on Predictive Crisis Anticipation and moderately on Multi-faceted Impact Quantification, but remains limited on Active Evolution Reasoning. Claude Opus 4.8 performs competitively on Predictive Crisis Anticipation, while Qwen3.5-397B-A17B remains consistently behind the stronger closed-source models. These results indicate that model capacity and general reasoning capability remain important for raw-observation-based disaster intelligence.

These lifecycle-stage results suggest that disaster intelligence is not a single homogeneous capability. PCA emphasizes early physical signal interpretation, AER requires temporal evolution modeling, and MIQ requires connecting physical observations with humanitarian and socio-economic consequences. A model that performs well in one stage may still fail in another, demonstrating the necessity of evaluating the full disaster lifecycle rather than reporting only a single aggregated score.

The disaster-level results in Table IV further reveal strong model–hazard interactions. GPT-5.5 performs best on earthquake, wildfire, and mass movement events, which may benefit from broader world knowledge and semantic reasoning about disaster mechanisms. Claude Opus 4.8 achieves the best score on storm, while Qwen3.5-397B-A17B shows weaker performance across most disaster categories. These results demonstrate that model capability varies substantially across hazard families, even under the same benchmark format.

These results demonstrate that model capability varies substantially across hazard families. A model that performs well on one disaster type may not generalize to another, even under the same benchmark format. This supports the need for a multi-hazard benchmark such as Obshazard-bench, where evaluation spans hydrological, meteorological, geophysical, and environmental disasters rather than focusing on a single hazard family. Overall, the lifecycle-stage and disaster-type analyses jointly show that current MLLMs exhibit structured capability gaps across both operational stages and physical hazard mechanisms.

Refer to caption
Fig. 4: Temporal-window ablation across lifecycle tasks. We compare model performance under different observation windows for three lifecycle tasks: Predictive Crisis Anticipation (PCA), Active Evolution Reasoning (AER), and Multi-faceted Impact Quantification (MIQ). PCA, AER, and MIQ denote Predictive Crisis Anticipation, Active Evolution Reasoning, and Multi-faceted Impact Quantification, respectively.

IV-D Temporal-window Ablation Study

TABLE VI: Temporal-window ablation for Predictive Crisis Anticipation. We report PCA scores using observations collected 14, 7, and 3 days before event onset.
Model Days before event onset
14 7 3
Claude Opus 4.8 [52] 0.3973 0.4385 0.5731
GPT-5.5 [53] 0.3893 0.3529 0.2908
Kimi-k2.6-1T [54] 0.3750 0.3453 0.4065
Qwen3.5-397B-A17B [55] 0.2717 0.3204 0.3331
TABLE VII: Temporal-window ablation for Active Evolution Reasoning. We report AER scores using two event-stage observation checkpoints: the event-day checkpoint and a late-stage checkpoint before the maximum timestep.
Model Event-stage observation checkpoint
Start day End-1 day
Claude Opus 4.8 [52] 0.0557 0.0000
GPT-5.5 [53] 0.2443 0.2318
Kimi-k2.6-1T [54] 0.1322 0.1842
Qwen3.5-397B-A17B [55] 0.1270 0.0925
TABLE VIII: Temporal-window ablation for Multi-faceted Impact Quantification. We report MIQ scores across five observation checkpoints, spanning pre-event lead times and event-stage observations.
Model Days before event onset Event-stage observation checkpoint
14 7 3 Start day End day
Claude Opus 4.8 [52] 0.2366 0.2280 0.2379 0.2329 0.2131
GPT-5.5 [53] 0.3138 0.3058 0.2916 0.2978 0.2794
Kimi-k2.6-1T [54] 0.2440 0.2127 0.2357 0.2324 0.2128
Qwen3.5-397B-A17B [55] 0.1475 0.1664 0.1592 0.1609 0.1614

We further conduct a temporal-window ablation study to examine how model performance changes when observations are taken at different timesteps. Unlike the previous analysis, which focuses on model capability differences across stages and hazard types, this study investigates whether different lifecycle tasks benefit from observations closer to the disaster event. This is important for raw-stream disaster intelligence, where practical systems must decide how much historical observation to retain, how frequently to run inference, and which time windows are most informative for different operational tasks.

Table VI shows the ablation results for Predictive Crisis Anticipation using observations collected 14, 7, and 3 days before event onset. Claude Opus 4.8 increases from 0.3973 with a 14-day lead time to 0.5731 with a 3-day lead time, suggesting that near-onset observations can contain more informative physical signals for early-warning tasks. However, GPT-5.5 shows a counterintuitive decline from 0.3893 to 0.2908 as the lead time shortens. Kimi-k2.6-1T follows a non-monotonic pattern, dropping at the 7-day lead time and recovering at the 3-day lead time, suggesting that its temporal sensitivity is weaker and less stable than Claude Opus 4.8. This indicates that temporal proximity alone does not guarantee improved reasoning if a model cannot effectively integrate sequentially updated observations. The result highlights the value of evaluating multiple pre-event lead times rather than relying on a single static pre-disaster snapshot.

Table VII further evaluates Active Evolution Reasoning using event-stage observations. GPT-5.5 performs best at both checkpoints and remains stable at the late-stage checkpoint, while Claude Opus 4.8 obtains near-zero performance at the late-stage checkpoint. Kimi-k2.6-1T improves from the event-day checkpoint to the late-stage checkpoint, indicating a mild benefit from additional event-tail observations. This highlights the importance of temporal grounding for reasoning about disaster evolution and recovery.

Table VIII shows a different pattern for Multi-faceted Impact Quantification. Across all models, MIQ performance remains relatively stable across timesteps. Kimi-k2.6-1T follows the same broadly time-invariant pattern, with only small fluctuations across the five checkpoints. This suggests that consequence-level estimation is less sensitive to temporal proximity than early-warning tasks. One possible explanation is that humanitarian burden and socio-economic intensity cannot be inferred from physical observations alone; they also require robust grounding in event metadata, exposure, vulnerability, and historical context. Therefore, additional observations closer to the event do not automatically improve consequence-level estimation.

Overall, the temporal-window ablation reveals that different lifecycle tasks respond differently to observation timing. PCA is generally time-sensitive and often benefits from near-onset signals, MIQ is relatively time-invariant and likely requires additional contextual grounding beyond raw observations, while AER can benefit from late-stage event-tail evidence. These findings demonstrate that Obshazard-bench can evaluate not only whether a model performs well, but also when observations become useful for different types of disaster reasoning.

IV-E Operational Implications

Beyond model ranking, the experimental results provide several implications for practical disaster-intelligence workflows. First, the temporal trend in PCA suggests the possibility of adaptive observation-window retention. If future models can reliably improve as event onset approaches, operational systems may preserve dense observations only within shorter high-risk rolling windows, reducing storage and inference costs while maintaining early-warning utility. However, this should be balanced against the value of longer lead time, since earlier but less accurate predictions may still be useful for evacuation planning and emergency resource pre-positioning.

Second, the weak temporal sensitivity of MIQ indicates that consequence-level estimation may not require continuous inference at every observation timestep. In practice, MIQ modules could be executed at selected checkpoints, while more system resources are allocated to collecting exposure, vulnerability, infrastructure, and socio-economic context. This suggests that improving impact assessment may depend less on simply increasing raw observation frequency and more on integrating decision-relevant contextual information.

Third, the results show that no single evaluated model performs consistently well across all lifecycle stages and disaster categories. General-purpose models, although strong in some settings, still exhibit clear stage-specific and hazard-specific capability gaps. This suggests that raw-observation-based disaster intelligence may involve multiple technical bottlenecks rather than a single unified reasoning ability. From an operational perspective, future disaster-response systems may benefit from combining complementary models or modules, where different models specialize in early warning, event evolution, or impact quantification.

Finally, given the low absolute scores, current MLLMs should be viewed as decision-support tools rather than autonomous disaster-response agents. For high-stakes tasks such as evacuation planning, humanitarian burden estimation, and recovery scheduling, model outputs should be coupled with uncertainty thresholds, expert review, and escalation mechanisms.

Overall, the experimental results show that current MLLMs still struggle to transform raw multi-channel Earth observations into lifecycle-aware disaster intelligence. The failures are not uniform: they depend on the lifecycle stage, temporal distance to the event, and disaster type. Obshazard-bench thus serves not only as a leaderboard, but also as a diagnostic benchmark for identifying stage-specific, time-dependent, and hazard-dependent limitations in current foundation models.

V Conclusion

We present Obshazard-bench, a real-time observation-driven benchmark for evaluating disaster intelligence from raw Earth observation streams. By integrating satellite sounding observations, disaster metadata, and auxiliary contextual information, Obshazard-bench covers 8 major disaster categories and 28 sub-categories across 127 historical disaster events. Furthermore, we introduce a lifecycle-oriented evaluation taxonomy consisting of Predictive Crisis Anticipation, Active Evolution Reasoning, and Multi-faceted Impact Quantification, enabling systematic assessment of disaster reasoning before, during, and after event occurrence.

Experimental results show that current MLLMs remain far from solving raw-observation-based disaster intelligence. Performance varies substantially across lifecycle stages, disaster categories, and observation windows, revealing distinct temporal, hazard-specific, and reasoning-related capability gaps. These findings suggest that disaster intelligence cannot be treated as a single homogeneous capability and requires stronger temporal grounding, event-evolution modeling, and consequence-level reasoning.

We hope Obshazard-bench can serve as a standardized testbed for future disaster foundation models and stimulate research on real-time Earth observation understanding, lifecycle-aware disaster reasoning, and decision-oriented geospatial intelligence.

References

  • [1] K. Kuckreja, M. S. Danish, M. Naseer et al., “Geochat: Grounded large vision-language model for remote sensing,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024.
  • [2] W. Zhang, M. Cai, T. Zhang et al., “Earthgpt: A universal multi-modal large language model for multi-sensor image comprehension in remote sensing domain,” 2024.
  • [3] X. Guo, J. Lao, B. Dang et al., “Skysense: A multi-modal remote sensing foundation model towards universal interpretation for earth observation imagery,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024, pp. 27 662–27 673.
  • [4] D. Hong, B. Zhang, X. Li et al., “Spectralgpt: Spectral remote sensing foundation model,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 8, pp. 5227–5244, 2024.
  • [5] A. Bell, A. Aides et al., “Earth ai: Unlocking geospatial insights with foundation models and cross-modal reasoning,” 2026.
  • [6] K. Bi et al., “Accurate medium-range global weather forecasting with pangu-weather,” Nature, vol. 619, no. 7970, pp. 533–538, 2023.
  • [7] R. Lam et al., “Learning skillful medium-range global weather forecasting,” Science, vol. 382, no. 6677, pp. 1416–1421, 2023.
  • [8] T. Nguyen et al., “Climax: A foundation model for weather and climate,” in International Conference on Machine Learning (ICML), 2023.
  • [9] K. Chen et al., “Fengwu: Pushing the frontiers of global medium-range weather forecasting,” 2023.
  • [10] J. Jakubik et al., “Foundation models for generalizable geospatial artificial intelligence,” 2023.
  • [11] F. Wang, H. Wang, Z. Guo, D. Wang, Y. Wang, M. Chen et al., “Xlrs-bench: Could your multimodal llms understand extremely large ultra-high-resolution remote sensing imagery?” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025, highlight, arXiv:2503.23771.
  • [12] F. Wang et al., “Omniearth-bench: Towards holistic evaluation of earth’s six spheres and cross-spheres interactions with multimodal observational earth data,” 2025, preprint, related to 2026 entry.
  • [13] X. Li, J. Ding, and M. Elhoseiny, “Vrsbench: A versatile vision-language benchmark dataset for remote sensing image understanding,” in Advances in Neural Information Processing Systems (NeurIPS), 2024.
  • [14] R. Shen et al., “Rsrsd-5m: A large-scale unlabeled dataset for pretraining earth foundation models,” 2026, preprint.
  • [15] M. Andrychowicz et al., “Deep learning for hourly geographical forecasting at kilometer scale with metnet-3,” 2023.
  • [16] G. Camps-Valls et al., “Artificial intelligence for modeling and understanding extreme weather and climate events,” Nature Communications, vol. 16, no. 1, p. 1919, 2025.
  • [17] M. Reichstein et al., “Early warning of complex climate risk with integrated artificial intelligence,” Nature Communications, 2025.
  • [18] F. Alam et al., “Crisismmd: Multimodal twitter datasets of natural disasters,” in Proceedings of the International AAAI Conference on Web and Social Media (ICWSM), 2018.
  • [19] R. Gupta et al., “Creating xbd: A dataset for assessing building damage from satellite imagery,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2019.
  • [20] D. Bonafilia et al., “Sen12-flood: A multi-spectral active-passive satellite dataset for flood detection,” in Proceedings of the IEEE International Geoscience and Remote Sensing Symposium (IGARSS), 2020.
  • [21] M. Rahnemoonfar et al., “Floodnet: A high-resolution aerial imagery dataset for post-flood scene understanding,” IEEE Access, vol. 9, pp. 89 644–89 659, 2021.
  • [22] R. Gupta and M. Shah, “Rescuenet: Joint building segmentation and damage assessment from satellite imagery,” 2020.
  • [23] D. Manzini et al., “Crasar-u-droids: A large-scale suas dataset for building damage assessment,” 2024, arXiv preprint.
  • [24] J. Wang, W. Xuan, H. Qi, Z. Liu et al., “Disasterm3: A remote sensing vision-language dataset for disaster damage assessment and response,” in Advances in Neural Information Processing Systems (NeurIPS), 2025, datasets & Benchmarks Track, arXiv:2505.21089.
  • [25] A. Karimi et al., “Zeshot-vqa: Zero-shot visual question answering for natural disaster damage assessment,” 2025.
  • [26] F. Al-Mohannadi et al., “Disastervqa: Social media-based visual question answering benchmark for crisis response,” 2026.
  • [27] A. Wang et al., “Dora: An end-to-end agentic benchmark for real-world disaster response,” 2026.
  • [28] H. Chen et al., “Bright: A globally distributed multimodal building damage assessment dataset with very-high-resolution for all-weather disaster response,” Earth System Science Data, 2025.
  • [29] Z. Wang et al., “Constructing an extensible building damage dataset via semi-supervised fine-tuning across 12 natural disasters,” Journal of Remote Sensing, 2025.
  • [30] A. of MONITRS et al., “Monitrs: Multimodal observations of natural incidents through remote sensing,” in Advances in Neural Information Processing Systems (NeurIPS), 2025, spotlight Poster.
  • [31] Z. Zhu et al., “Continuous monitoring of land cover changes using landsat time series,” Remote Sensing of Environment, vol. 185, pp. 1–3, 2017.
  • [32] S. Mayr and C. Kuenzer, “Validation of earth observation time-series: A review for large-area and temporally dense land surface products,” Remote Sensing, vol. 11, no. 22, p. 2616, 2019.
  • [33] A. Zhang, Z. Zhang, K. Shi, and P. Tang, “Benchmark datasets for satellite image time series classification: A review,” Remote Sensing, vol. 18, no. 10, p. 1581, 2026.
  • [34] Y. Fu, Z. Zhu, L. Liu et al., “Remote sensing time series analysis: A review of data and applications,” Journal of Remote Sensing, vol. 4, 2024.
  • [35] A. Toker et al., “Dynamicearthnet: Daily multi-spectral satellite dataset for semantic change segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  • [36] E. Rolf et al., “Fomo: Multi-modal, multi-scale and multi-task remote sensing foundation models for forest monitoring,” in Proceedings of the AAAI Conference on Artificial Intelligence, 2025.
  • [37] Y. Wang et al., “Treefinder: A us-scale benchmark dataset for individual tree mortality monitoring using high-resolution aerial imagery,” in Advances in Neural Information Processing Systems (NeurIPS), 2025.
  • [38] G. Astruc, N. Gonthier, C. Mallet, and L. Landrieu, “Anysat: One earth observation model for many resolutions, scales, and modalities,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025, pp. 19 530–19 540.
  • [39] J. Jakubik et al., “Terramind: Large-scale generative multimodality for earth observation,” 2025.
  • [40] K. Li et al., “Omnigaia: Towards native omni-modal ai agents,” 2026.
  • [41] P. Feng, Z. Lv, J. Ye, X. Wang, X. Huo, J. Yu, W. Xu, W. Zhang, L. Bai, C. He et al., “Earth-agent: Unlocking the full landscape of earth observation with agents,” arXiv preprint arXiv:2509.23141, 2025.
  • [42] D. T. Nguyen, T. Nguyen, F. A. Maani, H. M. Le, M. U. Sheikh, N. Saeed, M. H. Khan, and S. Khan, “Terrabench: Can agents reason over heterogeneous earth-system data?” 2026.
  • [43] M. Rahnemoonfar, T. Chowdhury, and R. Murphy, “Rescuenet: A high resolution uav semantic segmentation dataset for natural disaster damage assessment,” Scientific Data, vol. 10, no. 1, p. 913, 2023.
  • [44] C. Chen et al., “Rscc: A benchmark for remote sensing change captioning with rich human descriptions,” in Advances in Neural Information Processing Systems (NeurIPS), 2025.
  • [45] C. Li et al., “Anomaly-cd: Earth anomaly change detection with high-resolution time series,” 2024.
  • [46] B. N. U. Collaborators, “Shield: Unsupervised detection of disaster-affected areas,” Journal of Remote Sensing, 2026.
  • [47] C. Bodnar et al., “A foundation model for the earth system,” Nature, 2025.
  • [48] A. Allen et al., “End-to-end data-driven weather prediction,” Nature, 2025.
  • [49] S. K. Mukkavilli et al., “Ai foundation models for weather and climate,” 2023.
  • [50] B. Kim and T. Kim, “Ai in extreme weather events prediction and response: a systematic topic-model review (2015–2024),” Frontiers in Environmental Science, vol. 13, 2025.
  • [51] M. Preisser and P. Passalacqua, “Remote sensing improves multi-hazard flooding and extreme heat detection by fivefold over current estimates,” AGU Advances, vol. 6, no. 2, p. e2025AV001667, 2025.
  • [52] Anthropic, “Introducing claude opus 4.8,” https://www.anthropic.com/news/claude-opus-4-8, May 2026, accessed: 2026-06-23.
  • [53] OpenAI, “Introducing gpt-5.5,” https://openai.com/index/introducing-gpt-5-5/, Apr. 2026, accessed: 2026-06-23.
  • [54] Moonshot AI, “Kimi-k2.6,” 2026, model documentation / technical report.
  • [55] Qwen Team, “Qwen3.5: Towards native multimodal agents,” https://qwen.ai/blog?id=qwen3.5, Feb. 2026, accessed: 2026-06-23.