arXiv is now an independent nonprofit! Learn more
License: CC BY-NC-ND 4.0
arXiv:2609.00106v1 [cs.AI] 31 Aug 2026

Deploying and Evaluating a Smart-Agriculture Agentic Engine for Full-Season Soybean Farm Operations

DOI: XXXXXXX.XXXXXXXConference: Accepted to the 34th ACM International Conference on Advances in Geographic Information Systems; Nov; 2026ISBN: 978-1-4503-XXXX-X/2018/06CCS: Computing methodologies Intelligent agentsCCS: Applied computing Agriculture
Ao Qu Note: Equal contribution. email: aoqu@stu.hit.edu.cn Affiliation: Faculty of Computing, Harbin Institute of Technology, State Key Laboratory of Smart Farm Technologies and Systems, Harbin, China , Panagiotis Michelakis Note: Currently with new affiliation. email: panosg@synkrasis-labs.com Affiliation: School of Electrical and Computer Engineering, National Technical University of Athens, Athens, Greece , Linyuan Han email: linyuanhan26@stu.hit.edu.cn Affiliation: Faculty of Computing, Harbin Institute of Technology, State Key Laboratory of Smart Farm Technologies and Systems, Harbin, China , Yiannis Hadjiyianni email: yiannisha@synkrasis-labs.com Affiliation: School of Electrical and Computer Engineering, National Technical University of Athens, Athens, Greece , Kun Ouyang email: kunouyang@stu.hit.edu.cn Affiliation: Faculty of Computing, Harbin Institute of Technology, State Key Laboratory of Smart Farm Technologies and Systems, Harbin, China , Konstantinos Siskos email: siskos@synkrasis-labs.com Affiliation: School of Electrical and Computer Engineering, National Technical University of Athens, Athens, Greece , Feng Li Note: Corresponding authors. email: feng.li@hit.edu.cn Affiliation: Faculty of Computing, Harbin Institute of Technology, State Key Laboratory of Smart Farm Technologies and Systems, Harbin, China , Ran Meng email: mengran@hit.edu.cn Affiliation: Faculty of Computing, Harbin Institute of Technology, State Key Laboratory of Smart Farm Technologies and Systems, Harbin, China , Jingchi Jiang email: jiangjingchi@hit.edu.cn Affiliation: Faculty of Computing, Harbin Institute of Technology, State Key Laboratory of Smart Farm Technologies and Systems, Harbin, China , Dimitrios Stamoulis Note: Project Lead: FAIRY Platform, smart-farm multi-agent system and world models. email: dimi@hit.edu.cn Affiliation: Faculty of Computing, Harbin Institute of Technology, State Key Laboratory of Smart Farm Technologies and Systems, Harbin, China and Jie Liu email: jieliu@hit.edu.cn Affiliation: Faculty of Computing, Harbin Institute of Technology, State Key Laboratory of Smart Farm Technologies and Systems, Harbin, China
© none
Abstract.

This paper presents FAIRY, a full-stack smart-agriculture agent system developed for and deployed to an operating soybean research farm at Harbin Institute of Technology’s smart-agriculture site. We develop FAIRY to execute and evaluate agentic agronomic operations on full-season spatiotemporal workflows that span ridge preparation, planting, irrigation, fertilization, pest and disease treatment, harvest, grain handling, drying, and storage. FAIRY integrates APIs and infrastructure across production-grade machinery, fixed soil and canopy sensors, multispectral and thermal drones, satellite vegetation products, a weather station, calibrated crop-process models, agronomic records, and multi-season yield histories. The system is built around the novel “everything is an event” execution paradigm, which represents spatiotemporal world evolution, remote sensing and UAV observations, sensor readings, crop-growth transitions, machinery actions, and management interventions as state-changing events in a shared farm process engine. On top of this event-driven world model, FAIRY implements a complete agentic stack: a knowledge library of atomic agronomic skills; multi-agent controller and orchestration backends; frontier- and edge-model execution; full-path trace logging; and deployment profiling on local nodes. We use FAIRY to evaluate nine state-of-the-art agent controllers across one hundred full-season soybean scenarios that preserve the operational coupling between spatial observations in a 64-ridge field, temporal decision sequences, agronomic constraints, delayed effects, and final yield. We develop an evaluation suite that combines agentic success, full-path spatiotemporal correctness, token cost, and edge-device runtime. Our results show that our spatiotemporally grounded Kendall correctness (KTC) improves alignment with downstream yield outcomes compared to existing order-only and exact-match metrics. Our analyses further show that hierarchical agronomic skills and expert operational context substantially improve long-horizon behavior compared to geospatial in-context learning and LLM-as-an-Expert schemes.

Keywords: 
Spatiotemporal agents, Smart agriculture, Agent evaluation, Geospatial workflows, Full-season farm operations, World models

1. Introduction

Recent advances in agentic AI have made tool-augmented language models increasingly useful for geospatial analysis, where tasks often require data discovery, API selection, model invocation, map operations, and multi-step interpretation over spatial and temporal data. Prior geospatial agent systems have shown this potential across remote sensing, Earth observation, urban analysis, forestry, climate studies, agriculture monitoring, and satellite-vision workflows (Lee et al., 2025; Bhattaram et al., 2025). As reflected by recent releases in commercial geospatial platforms, such as Google Earth AI and Microsoft Planetary Computer offerings for agriculture and environmental monitoring, these settings are a natural fit for agents because geospatial workflows require selecting appropriate imagery or products, applying spatial filters, invoking detection or classification models, reasoning over intermediate outputs, and producing map-based results.

Full-process overview of the FAIRY agentic farm system.
Figure 1. FAIRY deploys and evaluates agentic controllers over a full soybean season (planting through grain storage) on a 64-ridge operating research farm, coupling spatial observations, temporal decisions, delayed agronomic effects, and final yield.Full-process overview of the FAIRY agentic farm system.

Despite these advances, deploying agentic systems for agricultural operations and complex agronomic spatiotemporal workflows remains difficult. Farm operations differ in three concrete ways (Yan et al., 2026; Zhang et al., 2026; Xu et al., 2025; Seo and Lee, 2026; Zuzuárregui et al., 2025; Qu et al., 2026): Feedback is delayed: a poor planting decision surfaces as weak emergence days later; a missed irrigation window appears as yield loss months later. Observations are partial and spatial: sensors are zone-level across ridges, drones are weather-gated, and ground inspection is sparse, so the agent reasons over a partially observed spatial field. Consequences propagate: operational errors compound across the season, so an action can be syntactically valid and operationally wrong, such as harvesting before grain moisture is suitable.

At the same time, practical deployments of geospatial agents often inherit orchestration and evaluation practices originally developed for domains such as web automation, coding, or OS control, where feedback is immediate and the action space is bounded by application APIs. However, unlike cloud-centric remote-sensing and geospatial big-data pipelines, agricultural applications introduce a broader evaluation surface: Crop state evolves through biological growth, weather, soil-water dynamics, sensing constraints, and management interventions, so agent decisions are coupled to physical processes that unfold over days, weeks, and seasons. Ultimately, agent performance depends not only on successful tool calls, but also on spatial coverage, temporal alignment, data-product choice, and the operational meaning of the produced analysis.

In this work, we present a full-stack smart-agriculture agent system developed for an operating soybean research farm at a university smart-agriculture research site (Figure 1) spanning a 64-ridge field. Real-life field operations span ridge preparation, planting, irrigation, fertilization, pest and disease treatment, harvest, grain handling, drying, and storage. The site has production-grade machinery; a digitized sensing layer covering fixed soil and canopy sensors, multispectral and thermal drones, satellite vegetation products, and a weather station; a calibrated, physics-grounded process model; and multi-season historical harvest and yield records. We are preparing a controlled portion of this field to be operated by agents alongside the existing human-operated workflow, with the goal of comparing agent-managed and human-managed operations in upcoming harvest seasons. This makes the evaluation question practical: we need to understand how agent systems behave on real farm operations before deciding what to deploy in the field.

To this end, we develop FAIRY as a full-stack agentic research engine for the operating farm. FAIRY first integrates the farm-facing infrastructure needed for deployment: production-grade machinery, fixed soil and canopy sensors, multispectral and thermal drones, satellite vegetation products, a weather station, calibrated crop-process models, agronomic records, and multi-season yield histories. The system then organizes these components through an “everything is an event” execution model, where weather updates, remote-sensing observations, UAV inspections, sensor readings, crop-growth transitions, machinery actions, and management interventions are represented as state-changing events in a shared farm process engine. On top of this event-driven world model, FAIRY implements the agentic stack used in our evaluation: a retrieval-based knowledge library of atomic agronomic skills, nine agent controller backends with optional agent-to-agent (A2A) orchestration, frontier- and edge-model execution, full-path trace logging, and edge-device deployment profiling.

Before deploying agents on the controlled farm portion, we need an evaluation setting that can exercise farm operations beyond the limited number of observed historical seasons. The calibrated farm world model allows us to construct realistic scenarios that remain tied to our field geometry, crop-process assumptions, sensing layer, and operational workflow, while also covering conditions that have not yet occurred in the deployment record. Specifically, we build one hundred scenarios across three levels of complexity: atomic tasks, episode chains, and full-season scenarios. We develop a comprehensive evaluation suite which covers task success, temporal correctness, coverage, full-path correctness, and yield preservation. We conduct extensive metric-calibration studies that assess which trace-level metric best tracks yield, identifying temporally grounded Kendall correctness (KTC) as the best-calibrated predictor and exposing the failure modes of order-only and exact-match alternatives. Finally, we include an operational demonstration that runs the same event-driven runtime end-to-end, from a user request through drone and satellite observation to a geospatial visualization in the FAIRY interface.

Our results lead to four practical lessons for deploying agents in farm operations. First, domain expertise is the dominant lever: expert operational context reduces the full-season yield shortfall from ∼\sim22% under zero context to ∼\sim5% for Qwen on held-out L3 scenarios. Second, agent and farm performance diverge under longer horizons: short tasks are nearly solved, with ≥\geq99% temporal correctness and near-zero yield loss, while full-season scenarios remain materially below the human oracle. Third, multi-agent A2A orchestration introduces coordination cost that degrades both correctness and yield in our setting. Fourth, agriculture-tuned LLM context is only a partial substitute for human-expert context. We report these as field observations from an operating farm; the engine, scenarios, and evaluation are documented for reproduction.

2. The FAIRY System: Architecture Overview

An overview of FAIRY is shown in Figure 1. The system couples an event-driven engine, a physics-grounded process stack, a tool/sensing layer, a knowledge library, and the agentic backend.

2.1. Farm site and APIs

Farm equipment and sensing infrastructure. The study site is an industry-grade university soybean research farm with a modeled 268​m×71​m268\,\mathrm{m}\times 71\,\mathrm{m} field organized into 64 ridges, which serve as the atomic spatial units for observation and intervention. We expose farm capabilities to the agent as function-calling tools derived from the installed APIs and operational interfaces. As summarized in Table 1, the site combines fixed soil, canopy, weather, light, radiation, and chlorophyll sensing; UAV and calibration assets for multispectral, thermal, and LiDAR observation; ground robotic platforms; production machinery; spraying equipment; and ridge-level irrigation and fertilization facilities.

Table 1. FAIRY agent APIs closely model and connect to the underlying farm equipment and on-field sensing assets.
Category Asset Qty.
Field Soybean ridges 64
Soil sensing DF-G3012 + DF-HRS 6
Canopy sensing Apogee canopy index sensor 6
Weather WX-CQ10 weather station 1
Light sensing DLS light sensor 1
Radiation Solar-radiation sensor 1
Crop sensing SPAD chlorophyll meter 1
UAV DJI Matrice 4T thermal UAV 1
UAV DJI Mavic 3 Multispectral 1
LiDAR DJI Zenmuse L2 1
UAV automation DJI drone dock station 2
Calibration Ground control markers 12
Calibration P4M reflectance panels 2
Calibration Blackbody radiation source 1
UAV sensing Drone-mounted radiation sensor 1
Machinery Tractor 1
Machinery Container trailer 1
Spraying Tractor-mounted spray boom 1
Spraying Backpack spray tank 1
Irrigation/fertilization Ridge-level facilities per ridge

Satellite imagery and map products. We use Google Earth Engine (GEE) to retrieve Sentinel-2 Surface Reflectance Harmonized imagery over the target area of interest (AOI), which we resolve either from a GEE asset or from a local AOI ZIP/shapefile. The retrieval pipeline filters scenes by AOI, date range, and cloudy-pixel percentage, sorts candidate scenes by cloud coverage, and exports the selected image as a 10 m GeoTIFF in EPSG:4326. We export bands B2, B3, B4, B5, B6, B7, B8, B8A, B11, and B12 through GEE export and Google Drive/API download logic. Downstream, we clip the raster to the AOI and georegister map placement from the GeoTIFF bounds in EPSG:4326. For visualization, we generate a satellite backdrop PNG from B4/B3/B2 RGB bands using percentile stretching and transparency masking, then overlay the result in the Leaflet-based map UI.

Satellite crop classification. We implement crop classification with an in-house XGBoost multiclass model. The classifier uses the gbtree booster with a multi:softmax objective and takes the ten Sentinel-2 bands together with vegetation indices including NDVI, kNDVI, GW1, GW2, LSWI, NDWI, EVI, EVI2, MSAVI, GNDVI, NDRE, GWCCI, and REP. We extract labeled pixels from crop polygons using geometry masks and use a stratified 70/30 train-test split. The model supports both training from scratch and loading a pretrained checkpoint with optional fine-tuning; in this study, we do not enable additional Heilongjiang regional fine-tuning. The operational output classes are rice, maize, soybean, other, and background.

Drone imagery and APIs. We collect field imagery at regular intervals with farm personnel and are also experimenting with direct UAV control through DJI Cloud API. The current deployment uses a DJI M300-RTK equipped with a DJI P1 full-frame RGB camera, producing 8192×\times5460 imagery for field inspection. We generate orthomosaic products through DJI Terra, with OpenDroneMap also supported as an agent-callable processing backend. For field identification, we extract field-edge information using a multi-self-adaptive-threshold Canny operator and use the excess green index (ExG) to filter plant-covered regions. For missing-seedling detection, we combine overall plant density with local plant density estimated through sliding-window counting to identify sparse or missing-seedling areas. The drone and examples of the imagery and collected results are illustrated in Figure 2.

Example visualizations of on-field drone inspections.
Figure 2. On-field drone inspections and UAV APIs.Example visualizations of on-field drone inspections.

2.2. Growth process dynamics

Weather model. The farm world model is driven by weather and remote-sensing observations that define the external conditions for crop growth and management. Weather is generated through a WGEN/Richardson-style daily weather model (Siler and Singh, 2022), providing precipitation, temperature, radiation, wind, and humidity at the resolution required by the farm event loop. The observation layer incorporates the remote-sensing and drone-based products available during operation. We treat these observations as state updates that provide spatial evidence about canopy vigor, water stress, and anomaly regions for follow-up inspection, irrigation, fertilization, or pest-management decisions.

Growth model. Plant growth dynamics are represented through a coupled soil, phenology, canopy, and biotic-pressure stack following (Wang et al., 2025). The soil component maintains water and temperature states using a bucket-style water balance over precipitation, irrigation, runoff, drainage, and evapotranspiration. Phenology follows a GDD-based soybean development model with seed-type-specific maturity targets (Akyuz et al., 2017), supporting cultivar-specific assumptions and stage-dependent management rules. The canopy and biomass model follows Monteith-style radiation-use and light-interception principles (Monteith, 1977; Monsi and Saeki, 2005). The biotic-pressure component tracks weed, insect, and disease pressure as weather- and stage-dependent processes with treatment effects (Steduto et al., 2009).

Action effects and yield. Together, these components define a process-level state space in which action effects depend on crop stage, soil condition, recent weather, and observed canopy state. Management actions modify the crop trajectory through delayed and stage-dependent effects: Planting establishes stand fraction and emergence timing; irrigation changes soil-water availability over subsequent days; fertilization affects nutrient stress and canopy development. At maturity, we employ a yield-recovery model to convert accumulated crop state into harvested grain (Humburg, 2019).

Physics-engine calibration. We validate the engine against historical data by reconstructing 18 plot-level soybean scenarios from the 2025 growing season (05-10 to 05-27) and harvest campaign (09-13 to 09-23). Different scenario IDs span different planting dates, density treatments, and cultivar type. Each FAIRY world model is initialized to closely match historical plot-specific management operations, cultivar type, planting density, fertilization schedule, and measured daily weather. We use observed phenology and yield records as evaluation targets, while the engine independently simulates emergence, crop growth, soil water and nutrient stress, biomass accumulation, harvest recovery, and final yield. As shown in Figure 3, our engine yields closely track the observed yields across cultivars and treatments, with nearly all scenarios within yield prediction accuracies of 2%2\% MAE. We leave integration with commercial tools, such as WOFOST, to future work.

FAIRY Physics Engine closely matches real-world farm yields from historical harvest data.
Figure 3. FAIRY Physics Engine yield vs. real-world farm yield across 18 Farm IDs.FAIRY Physics Engine closely matches real-world farm yields from historical harvest data.

2.3. Stateful event-driven engine

We build the farm execution engine on the ARE framework (Froger et al., 2025), which represents agent environments through stateful apps, an event queue, notifications, scenarios, and logged execution traces. Farm operation is naturally event-based: tools execute ridge preparation, planting, irrigation, fertigation, pesticide application, harvest, and storage as timestamped farm events, each with arguments (e.g., operation duration, target ridges) and preconditions (e.g., weather and equipment readiness). The environment maintains the farm state, event queue, notifications, and operation history, while the controller observes and acts through the tool interface. Tool invocation is emulated inside the digital twin, so operations retain their farm-state effects and timing dependencies without real-clock execution. This evaluation should therefore be interpreted as a deployment-readiness study rather than a completed agent-managed harvest trial. The physical infrastructure, sensing stack, machinery interfaces, and historical operation records are real, while the agent actions are replayed through the calibrated digital twin to evaluate safety and operational correctness before field deployment.

2.4. Agent families and LLMs

We integrate nine controller architectures and evaluate them over the scenario suite: ReAct (Yao et al., 2023), Plan-and-Act (Erdogan et al., 2025), Reflexion (Shinn et al., 2023), AutoGen (Wu et al., 2024), MMRL (Tan et al., 2026), ReWOO (Xu et al., 2023), LATS (Zhou et al., 2024), CRITIC (Gou et al., 2024), and GoT (Besta et al., 2024). Each controller supports two modes: direct tool access and agent-to-agent (A2A) (A2A Project, 2025), where weather, sensing, machinery, and operations queries are routed through specialist app agents. Underlying farm state and tool capabilities are kept consistent across controllers and modes. We run frontier API backends (Qwen3.6-35B-A3B (Yang et al., 2025), DeepSeek-V4-Flash (DeepSeek-AI, 2026)). To reflect realistic deployment constraints, we also support edge execution through a vLLM API running on the NVIDIA Thor SDK with recent Qwen-3 and Gemma-4 model variants.

2.5. Knowledge library

Farm management relies on tacit timing rules and operating discipline that are usually implicit in human expertise. FAIRY encodes this knowledge as a knowledge library of agronomic skills that controllers can retrieve at decision time, mirroring hierarchical skill/prompt libraries used for geospatial and remote-sensing agents (Bhattaram et al., 2025; Badmus et al., 2027; Singh et al., 2024). Each skill is a structured record with an identifier, a natural-language title and description, a set of keywords, and a reusable operation workflow template. At decision time, a retriever scores library entries against the current task and injects the top-kk (k=3k{=}3 by default) into the controller context. We expose two retrieval mechanisms: a lightweight lexical retriever scoring token overlap between the query and each entry’s searchable text, and a semantic retriever over sentence embeddings of the entire full-path workflow; both return ranked skills with scores.

This library underlies one of four interchangeable in-context regimes the controller can run under: Zero Context (operation primitives only), LLM-as-an-Expert (an agriculture-tuned LLM (Wang et al., 2025) generates scenario context from the task, tools, and world state), Skills Library (retrieved skills injected into context), and Expert Instruct (human-written scenario instructions). We vary how the library is organized for retrieval (i.e., a flat pool versus a tier-grouped organization that separates atomic from composite operations) and the ranking mechanism used to select entries (i.e., manual selection, lexical text similarity, and path similarity).

Table 2. Long-horizon analysis on atomic tasks (L1), episode chains (L2), and full-season scenarios (L3-test) for Qwen3.6-35B-A3B and DeepSeek-V4-Flash. Yield Loss is the % drop from the human-oracle biological yield; Succ. is BFCL tool-call success; KTC is temporally grounded correctness; Token cost is reported as the average tokens per task and per agent API call.
Task In-context Yield Succ. Path KTC Tkns/ Tkns/ Yield Succ. Path KTC Tkns/ Tkns/
Level Learning Loss Score Corr. Score Task Step Loss Score Corr. Score Task Step
Qwen3.6-35B-A3B DeepSeek-V4-Flash
L1 Zero Context 1.3% 42.0% 33.2% 75.1% 0.26M 13.4k 2.9% 46.1% 25.2% 78.1% 1.78M 48.8k
LLM-as-an-Expert 0.1% 86.6% 65.6% 98.7% 0.29M 13.5k 0.1% 89.0% 68.1% 98.8% 0.31M 13.8k
Expert Instruct 0.1% 85.6% 66.7% 99.2% 0.28M 13.6k 0.3% 88.2% 66.6% 99.2% 0.28M 13.2k
L2 Zero Context 4.3% 39.6% 22.3% 73.9% 0.66M 21.1k 6.5% 48.6% 17.1% 74.3% 2.03M 28.8k
LLM-as-an-Expert 0.8% 69.0% 49.4% 98.1% 0.92M 27.1k 0.9% 67.3% 45.3% 97.7% 2.28M 45.0k
Expert Instruct 0.8% 65.9% 51.2% 98.0% 1.20M 32.7k 0.8% 70.0% 52.4% 98.1% 1.41M 33.8k
L3 Zero Context 22.3% 40.2% 22.8% 87.9% 3.89M 33.0k 12.6% 45.1% 28.5% 88.0% 5.55M 40.5k
LLM-as-an-Expert 15.3% 46.3% 24.9% 89.8% 3.66M 33.0k 11.6% 52.9% 26.5% 89.0% 5.34M 39.6k
Skills Library 4.9% 59.7% 30.9% 93.7% 4.09M 37.4k 6.3% 65.7% 33.0% 93.6% 6.76M 50.3k
Expert Instruct 4.6% 62.0% 28.6% 92.1% 2.93M 29.0k 2.9% 67.6% 30.5% 92.7% 4.29M 35.5k

2.6. Realistic full-season scenarios

We curate scenarios from representative on-field procedures based on historical farm operations. We consider three levels of increasing complexity: L1 atomic tasks require one or two tool calls (e.g., checking weather before a drone flight); L2 episodes chain observation, diagnosis, and intervention (e.g., detecting an anomaly from a drone survey and applying targeted treatment); and L3 full-season scenarios span planting through post-harvest storage, with multiple interventions, delayed consequences, and accumulated effects on crop state and recovered yield. For each scenario we follow the ARE annotation protocol (Froger et al., 2025): three domain experts independently specify oracle solutions using the same operation primitives exposed to the agents; if the first two diverge we inspect the scenario to resolve ambiguity and revise, and the third expert confirms consistency.

Example scenarios by level: L1 (atomic). “The drone detected signs of aphids on ridges 15–25; verify the issue and apply an appropriate pesticide treatment.” L2 (episode). “Remote sensing shows a localized low-NDVI area during V4; diagnose drought vs. pest vs. nutrient deficiency and, if nutritional, treat via ridge-level fertigation.” L3 (full season). “Manage the full season from planting onward with weekly monitoring, addressing nutrient, water, pest, and disease issues from sensor/drone evidence, selecting a harvest window after maturity, and completing drying and storage.”

Replaying the oracle workflows in the engine produces human-oracle farm-state trajectories and target crop yields against which agent runs are compared. The evaluation reported here uses a full-season test set of 70 L3 scenarios, alongside the L1/L2 splits. We also construct a focused 20 L3-mini test set used for ablations, as well as a held-out validation set of 20 L3 scenarios for confirming the generalization of different library schemes. The oracle workflows are not assumed to be globally optimal; they represent expert-validated operational references for reproducible comparison. In future work, we will quantify inter-expert disagreement and test the sensitivity of KTC and Yield Loss rankings to alternative oracle choices.

2.7. Evaluation suite and metrics

We evaluate each run along two axes: how the agent acted (trace-level correctness) and what it achieved (yield outcome). Both compare an agent run against the human-oracle workflow for the same scenario, represented as an ordered set of timestamped farm events, each carrying a tool, its arguments, target ridges, and a season-day.

Trace-level correctness. As our primary correctness metric we use KTC (Kendall-style Temporal Correctness), a temporally grounded full-path score that extends the path-correctness view of (Michelakis et al., 2025) to farm-event order. After matching the agent’s executed operations to the oracle’s, KTC measures the order agreement of the matched operations as a normalized Kendall rank correlation, KTC=(τ+1)/2∈[0,1]\mathrm{KTC}=(\tau+1)/2\in[0,1], rewarding required operations executed in causally valid order and penalizing reordering. We complement KTC with two reference trace metrics computed from the same (agent, oracle) pair: BFCL, a tool-call success rate measured as set overlap of executed (tool,args)(\text{tool},\text{args}) signatures (order- and time-agnostic); Path Correctness (CORE), a normalized edit distance between the agent and oracle operation sequences (Michelakis et al., 2025).

Agronomic outcome. We report Yield Loss, the percentage shortfall of the agent’s biological yield relative to the human oracle (1−agentbio/oraclebio1-\text{agent}_{\text{bio}}/\text{oracle}_{\text{bio}}). We pair correctness with outcome because they catch different failure modes: trace metrics flag plausible-but-wrong sequences a final-state check would miss (Froger et al., 2025), while yield loss captures operational misjudgments (e.g., harvesting on day 87 instead of 89) that a trace-only metric may treat as minor deviations. We additionally report token cost and runtime per scenario.

Operational demonstration of the agentic farm.
Figure 4. FAIRY geospatial user interface. The web-based viewer displays the full-path agent workflow together with map-based visualizations.Operational demonstration of the agentic farm.

2.8. FAIRY system user interface

We adapt the ARE user interface (Froger et al., 2025) from a generic agent-workflow environment into a geospatial inspection interface for agricultural workflows, as shown in Figure 4. Notably, we extend the UI with domain-specific visualization support for satellite and UAV outputs, including zoomable OpenStreetMap-based map views, overlaid satellite imagery, crop-classification rasters with adjustable opacity, and drone-analysis tabs for UAV patrol paths and orthomosaic-derived products. We connect each application module to dedicated state and tool panels, so users can inspect the execution chain, intermediate outputs, and final geospatial products within the same workflow interface.

Table 3. Agent performance on L3-test-mini full-season scenarios with A2A enabled.
Agent In-context Yield Succ. Path KTC Tkns/ Tkns/ Yield Succ. Path KTC Tkns/ Tkns/
Routing Learning Loss Score Corr. Score L3 sc. Step Loss Score Corr. Score L3 sc. Step
Qwen3.6-35B-A3B DeepSeek-V4-Flash
Direct Zero Context 27.3% 38.7% 22.9% 88.9% 3.38M 31.9k 16.3% 43.9% 26.9% 87.4% 5.03M 39.4k
LLM-as-an-Expert 9.8% 47.8% 25.0% 89.8% 4.12M 35.2k 8.5% 56.1% 25.2% 90.5% 5.40M 39.8k
Skills Library 1.6% 61.4% 30.5% 93.8% 4.01M 38.1k 1.9% 67.2% 32.8% 93.5% 7.29M 50.9k
Expert Instruct 4.0% 63.1% 29.5% 92.8% 2.94M 29.0k 1.8% 67.2% 30.0% 92.7% 4.50M 36.5k
A2A Zero Context 43.3% 22.1% 15.3% 84.0% 1.26M 12.1k 54.4% 22.2% 14.6% 92.8% 4.97M 25.3k
LLM-as-an-Expert 24.6% 39.4% 21.2% 89.1% 1.61M 10.7k 33.6% 36.7% 16.7% 85.7% 2.16M 11.9k
Skills Library 18.0% 51.7% 22.5% 90.8% 2.22M 12.9k 27.5% 46.5% 27.6% 94.3% 3.81M 21.0k
Expert Instruct 31.8% 43.6% 21.4% 87.5% 1.10M 10.9k 9.5% 42.4% 22.0% 87.5% 1.69M 12.6k
Table 4. Agent performance on L3-validation full-season scenarios across library schemes. Flat structures correspond to directly retrieving against L3-level contexts, and hierarchical structures correspond to multi-level (L1, L2, L3) retrieval contexts.
Knowledge Library Similarity Yield Succ. Path KTC Tkns/ Tkns/ Yield Succ. Path KTC Tkns/ Tkns/
Library Structure Mechanism Loss Score Corr. Score L3 sc. Call Loss Score Corr. Score L3 sc. Call
Qwen3.6-35B-A3B DeepSeek-V4-Flash
Zero Context – – 21.6% 37.6% 22.2% 85.6% 3.54M 33.0k 22.3% 43.5% 27.8% 88.7% 5.12M 38.2k
Expert Instruct – – 3.1% 60.2% 26.7% 91.9% 3.55M 31.1k 1.3% 70.1% 28.6% 93.3% 5.74M 40.3k
GeoLLM-Engine (Singh et al., 2024) Flat Text (Manual) 3.0% 60.8% 34.3% 92.7% 4.74M 40.4k 3.7% 63.8% 34.4% 93.6% 6.19M 44.6k
GeoFlow (Bhattaram et al., 2025) Hierarchical Full-path (Manual) 2.9% 68.0% 34.8% 97.2% 4.57M 39.5k 2.8% 73.5% 41.1% 96.8% 7.60M 47.4k
PowerChain (Badmus et al., 2027) Flat Textual 7.9% 59.1% 35.9% 94.6% 4.41M 40.1k 2.4% 71.9% 38.1% 94.5% 7.84M 47.4k
PowerDAG (Badmus and Pandey, 2026) Flat Full-path 7.8% 60.6% 34.7% 95.3% 4.44M 40.2k 4.6% 67.6% 38.5% 95.8% 7.45M 44.2k
HTAM (Li et al., 2025) Hierarchical Full-path 3.5% 58.8% 31.0% 92.9% 4.73M 40.4k 8.1% 58.2% 32.1% 93.7% 5.65M 44.9k
FAIRY Skills Library Hierarchical Text + Grouped Paths 3.5% 61.1% 32.0% 93.8% 5.11M 41.6k 3.0% 63.1% 32.4% 93.9% 6.04M 46.1k
Figure summarizing agentic performance results across different controller and backbone implementations.
Figure 5. Agent-family and backbone comparison on L3-mini full-season scenarios. The left and right panels report Qwen3.6-35B-A3B-FP8 and DeepSeek-V4-Flash, respectively. For each agent family, colored bars denote different context settings, and rows show full-season yield, FAIRY KTC score, and token cost per step.Figure summarizing agentic performance results across different controller and backbone implementations.

3. Results

We organize results around the questions an operator would ask before deploying: what makes agents reliable (context and skills), how reliability scales with horizon, what multi-agent orchestration costs, which metric we should trust, and what deployment costs.

Expert context dominates the controller spread. Table 2 reports the held-out N=70N{=}70 L3 evaluation under four in-context regimes for both backbones. Context is the dominant lever. For Qwen, moving from Zero Context to Expert Instruct cuts Yield Loss from 22.3%22.3\% to 4.6%4.6\% and raises KTC from 87.9%87.9\% to 92.1%92.1\%; the Skills Library regime is close behind (4.9%4.9\% Yield Loss, 93.7%93.7\% KTC) without any hand-written per-scenario policy. DeepSeek shows the same ordering (Zero →\rightarrow Expert: 12.6%→2.9%12.6\%\rightarrow 2.9\% Yield Loss). Notably, LLM-as-an-Expert recovers only part of the gap (Qwen 15.3%15.3\% Yield Loss versus 4.6%4.6\% for human instructions), indicating that agriculture-tuned LLM context captures high-level decisions but not the full procedural discipline.

Table 5. Local vLLM performance on atomic tasks (L1), episode chains (L2) and full-season scenarios (L3-mini) under expert-instructed context. Runtime (seconds) is measured on NVIDIA Thor.
Tasks Model Yield Succ. Path KTC Tkns/ Tkns/ Runtime/ Runtime/
Level Type Loss Score Corr. Score Task Step Step Task
L1 Qwen3.6-35B-A3B-FP8 0.2% 85.5% 65.3% 98.9% 0.29M 13.8k 4.3s 101.1s
Qwen3.6-27B-FP8 0.2% 87.7% 68.3% 99.0% 0.29M 13.8k 20.2s 427.0s
Qwen3.5-9B 1.7% 87.5% 67.1% 99.3% 0.39M 16.6k 12.8s 312.6s
Qwen3.5-4B 0.6% 85.7% 59.0% 98.3% 0.82M 27.3k 11.0s 411.9s
Qwen3.5-2B 17.6% 46.3% 13.1% 97.8% 7.04M 56.6k 2.3s 735.7s
Gemma-4-31B-IT-NVFP4 0.1% 87.3% 69.1% 98.8% 0.31M 15.3k 16.7s 339.5s
L2 Qwen3.6-35B-A3B-FP8 1.1% 72.1% 49.2% 97.9% 0.87M 25.2k 4.4s 164.7s
Qwen3.6-27B-FP8 0.8% 70.1% 45.7% 97.6% 0.77M 22.2k 20.7s 720.4s
Qwen3.5-9B 1.2% 67.3% 47.2% 97.5% 0.85M 26.6k 15.7s 510.6s
Qwen3.5-4B 1.5% 67.5% 48.4% 96.4% 1.45M 30.3k 6.6s 931.4s
Qwen3.5-2B 11.7% 31.8% 11.2% 90.6% 5.53M 48.4k 2.5s 518.4s
Gemma-4-31B-IT-NVFP4 1.1% 69.6% 56.0% 98.6% 0.68M 23.3k 19.1s 567.5s
L3 Qwen3.6-35B-A3B-FP8 2.1% 66.6% 28.1% 94.2% 3.77M 31.9k 3.3s 407.1s
Qwen3.6-27B-FP8 2.3% 66.6% 35.1% 93.1% 2.56M 26.1k 14.7s 1451.8s
Qwen3.5-9B 3.0% 54.6% 26.2% 93.5% 2.60M 29.5k 11.1s 981.9s
Qwen3.5-4B 5.0% 50.7% 22.2% 88.1% 4.07M 36.5k 7.6s 861.2s
Gemma-4-31B-IT-NVFP4 2.1% 53.9% 26.6% 92.0% 1.93M 26.3k 17.1s 1267.2s
Gemma-4-26B-A4B-NVFP4 2.4% 45.6% 18.9% 88.2% 1.53M 25.1k 3.5s 222.5s
Gemma-4-26B-A4B-NVFP4-assistant 12.0% 42.6% 16.2% 86.1% 1.79M 31.2k 7.5s 444.1s
Gemma-4-12B-it 0.9% 59.5% 29.1% 92.4% 2.58M 30.6k 12.3s 1039.9s
Gemma-4-12B-it-assistant 7.1% 58.3% 27.8% 91.7% 3.20M 35.0k 9.0s 823.5s
Gemma-4-E4B-it 31.2% 41.7% 19.5% 80.8% 1.29M 23.6k 6.4s 382.8s
Gemma-4-E2B-it 40.0% 31.3% 9.8% 67.5% 1.76M 26.2k 3.4s 238.9s
Figure summarizing the alignment of agentic evaluation metrics with downstream agronomic performance.
Figure 6. Left: Assessing full-path agent metrics vs. biological yield preservation. Right: Evaluating the cost vs. downstream objective performance (yield) tradeoff across scenario runs.Figure summarizing the alignment of agentic evaluation metrics with downstream agronomic performance.

Knowledge-library structure and retrieval. Table 4 ablates how the knowledge library is organized (flat vs. tier-grouped) and how entries are ranked (manual, text similarity, path similarity) on the L3-mini set. Tier-grouped organization with similarity-based retrieval gives the strongest correctness (KTC up to 97.2%97.2\% for Qwen, 96.8%96.8\% for DeepSeek) and the best task success, while flat/manual retrieval is weaker and more variable. The effect on yield is smaller than the context effect, which is expected: retrieval quality refines how an already-grounded agent acts, whereas context grounding determines whether it acts correctly at all.

Horizon analysis: short tasks are nearly solved, full seasons are not. Table 2 reports atomic (L1) and episodic (L2) tasks. Under expert context, both horizons are essentially solved: L1 reaches ≥\geq99% KTC with near-zero yield loss, and L2 reaches ≥\geq98% KTC with <<1% yield loss. The contrast with the L3 results is the central horizon finding: short tasks mostly test whether the scaffold can select tools and satisfy local preconditions, whereas full-season scenarios test whether decisions remain coherent after their effects propagate through soil state, growth, stress accumulation, treatment residuals, harvest timing, and storage. Yield loss is where the horizon bites: it is small on L1/L2 but reaches double digits on L3 under zero context.

Operational demonstration of the proposed FAIRY framework on our research farm.
Figure 7. FAIRY Operational Demonstration: drone patrol and Sentinel-2 crop classification coordinated through the event-driven runtime, rendered as a geospatial inspection result in the system UI.Operational demonstration of the proposed FAIRY framework on our research farm.

Multi-agent orchestration introduces coordination cost. Table 3 compares direct tool access against A2A routing, where the controller coordinates with weather, sensing, machinery, and operations specialists. Averaged over contexts, A2A degrades both correctness and yield on both backbones: Yield Loss rises from 7.1%7.1\% to 31.2%31.2\% (DeepSeek) and from 10.7%10.7\% to 29.4%29.4\% (Qwen). Token cost per scenario drops because work is offloaded to specialists, but the main controller frequently loses track of what each specialist observed or executed. The effect is controller-dependent (tree search is comparatively robust), consistent with coordination-overhead observations in other expert-level multi-agent studies (Lee et al., 2025; Badmus et al., 2027). Early analysis suggests errors due to task decomposition issues, so we plan a detailed failure taxonomy as future work.

Which trace metric predicts yield? Figure 6 (left) and Table 6 regress each metric against biological yield preservation across the full sweep. Overall, we observe that KTC is the best-calibrated predictor as it lies almost on the unit line. Moreover, we note that BFCL’s high R2R^{2} performance comes from robustly identifying catastrophic “never harvested” runs (zero recovered yield), rather than grading quality among completing runs: it detects failure but may be poorly calibrated. Overall, the results show that existing practices of order-only and exact-match metrics might not fully capture downstream tasks. Overall, we consider the following for our practical deployment and future work: report KTC for trace-level correctness and Yield Loss for outcome, and treat exact-match success as a failure detector rather than a quality measure.

Table 6. Alignment between agent-evaluation metrics and the downstream agronomic objective, measured by yield. Lower RMSE and slope closer to 1 indicate better calibration, while higher R2R^{2} indicates stronger explanatory fit.
Full-path agent metric RMSE Slope aa R2R^{2}
Success rate BFCL (Patil et al., 2025) 37.1 1.63 0.81
Path correctness ARE (Froger et al., 2025) 62.5 3.04 0.41
FAIRY KTC score 4.2 1.01 0.70

Edge deployment profiling. Because field deployment cannot assume frontier-API availability, we assess performance and cost under edge deployment considerations. Table 5 shows that local vLLM execution can support L1 and L2 farm tasks with low yield loss across several mid-size models, but full-season L3 scenarios separate models more clearly: Qwen3.6-35B-A3B-FP8 and Qwen3.6-27B-FP8 keep yield loss near 2%, while smaller or assistant-tuned variants degrade sharply. In practice, this suggests that edge deployment is feasible for farm-agent evaluation, but full-season autonomy still requires models with enough planning capacity.

Per-controller robustness. Figure 5 summarizes the performance across the 9 different agentic back-ends on the L3-mini set across the four in-context regimes. Overall, we observe that in-context operational grounding is the dominant factor across backbones: moving from zero context to the skills library or expert instruction reduces average yield loss from 19.9% to 2.9%/2.1% for Qwen3.6-35B-A3B and from 12.8% to 2.2%/3.0% for DeepSeek-V4-Flash. In practice, this means that the farm deployment cannot rely on generic, state-of-the-art agent scaffolds alone; agents need explicit agronomic skills or expert operational context to preserve yield under full-season decision sequences.

4. FAIRY Operational Demonstration

We demonstrate FAIRY on a real-farm inspection episode at the HIT smart-agriculture site. As shown in Figure 7, the demonstration focuses on the following L2 episode: a user issues a field-monitoring request, and the agent must coordinate satellite and drone observations, crop-identification tools, post-processing, and the FAIRY UI to return an interpretable inspection result. The integrated episode exercises the full observation workflow: the agent first triggers the drone API, retrieves and classifies satellite imagery while the drone is in flight; it then retrieves the drone result, generates the orthomosaic, plots the patrol path, runs field/crop analysis, integrates satellite and drone findings, and updates the UI display tool to render base-map context, classification results, the patrol path, and image-analysis products. Overall, this demonstration provides us with a working deployment prototype and end-to-end path from user request, through event-driven agent execution, to farm-facing geospatial visualization.

5. Discussion

From expert instructions to reusable skills. Encoding tacit expert knowledge as scenario-specific instructions is effective for evaluation but does not scale: each new task or seasonal edge case would need another written policy. Organizing this knowledge into retrievable, composable skill libraries is the scalable alternative, consistent with the use of structured knowledge pools in power-grid (Badmus et al., 2027) and Earth-observation (Bhattaram et al., 2025; Shabbir et al., 2026; Chen et al., 2026; Feng et al., 2026) agents; our library ablation is a first step in this direction.

Full-path evaluation paired with farm objectives. A planting, irrigation, or spraying decision can look locally plausible while still causing downstream yield loss, so the downstream objective must be evaluated alongside the action trace. KTC and Yield Loss diagnose different failure modes: a high-KTC, high-yield-loss run shows that small operational differences can have physical consequences, while a lower-KTC, low-yield-loss run shows deviation from the reference trace that nonetheless preserved the outcome. Our metric-calibration study sharpens this: among trace metrics, temporally grounded correctness tracks yield, whereas flat order metrics and success measures do not. We note that the current analysis aggregates results at the scenario level. A natural next step is ridge-level spatial analysis, including whether errors cluster around sparsely instrumented ridges, low-observability zones, or operations that depend on UAV and satellite coverage.

Agent orchestration should preserve operational state. Planning, memory, retrieval, verification, and specialist decomposition can improve the execution trace, but they do not remove the need to maintain agronomic and operational assumptions across time. Splitting the farm interface into specialists matches the structure of real farm systems, yet introduces state-sharing requirements; if the controller loses track of what each specialist observed or executed, orchestration becomes error-prone. The design requirement is not only better decomposition but preserving shared farm state, timing constraints, and task context across agent boundaries.

LLM-as-an-Expert context. Even when a domain LLM (Wang et al., 2025) is given the scenario, tools, world state, and a matched response template, it does not fully match human experts on procedural ordering. This does not make the domain LLM unhelpful, as it recovers high-level crop-management choices, but it motivates structured knowledge libraries that complement both LLM-distilled and human-written guidance.

Quality of human-annotated solutions. Because every trace metric compares against a human-oracle workflow, the oracle must be a fixed, shared artifact; regenerating it across software versions or machines introduces drift that silently changes every metric. We recommend version-controlling and serializing the oracle workflows alongside each scenario. Moreover, a timing-aware metric can only be validated where mistiming actually costs yield, making such information available in the oracle traces particularly important. The next stage of deployment is to run an agent-managed plot alongside the human-operated workflow and compare both operational traces and harvested outcomes under real seasonal conditions.

6. Conclusion

We presented FAIRY, a deployable smart-agriculture agentic engine, and used it for a full-season spatiotemporal evaluation of contemporary agent practices on an operating soybean research farm. Expert context and retrievable agronomic skills are the dominant levers for long-horizon reliability; multi-agent orchestration and deployment cost are practical constraints; and among trace-level metrics, temporally grounded correctness is the one that tracks real yield. These results inform our next stage: operating a dedicated plot under agent management alongside the human-operated farm. We hope FAIRY provides the community a deployable, spatiotemporally grounded reference point for evaluating agents in other physical-process settings. Our entire working prototype can be found here: https://github.com/Fengrui-Lab/FAIRY

Acknowledgements.
This work was supported by the National Natural Science Foundation of China (grant No. 62350710797). We gratefully acknowledge the support of the National Key Research and Development Program [2025YFE0209200] and the Key Research and Development Program of Heilongjiang Province, China [2024ZX01A07, JD2023GJ01]. This work was also supported by the NSFC grant (No. 42471362). DS gratefully acknowledges the support of the NSFC Excellent Young Scientists Fund Program (Overseas).

References

  • A2A Project (2025) A2A Project Agent2Agent protocol (A2A). Note: https://github.com/a2aproject/A2AGitHub repository Cited by: §2.4.
  • Akyuz et al. (2017) F. A. Akyuz, H. Kandel, and D. Morlock Developing a growing degree day model for north dakota and northern minnesota soybean. Agricultural and Forest Meteorology 239, pp. 134–140. Cited by: §2.2.
  • Badmus and Pandey (2026) E. O. Badmus and A. Pandey PowerDAG: reliable agentic ai system for automating distribution grid analysis. External Links: 2603.17418, Link Cited by: Table 4.
  • Badmus et al. (2027) E. O. Badmus, P. Sang, D. Stamoulis, and A. Pandey PowerChain: a verifiable agentic ai system for automating distribution grid analyses. Electric Power Systems Research 262, pp. 113555. External Links: Document, Link Cited by: §2.5, Table 4, §3, §5.
  • Besta et al. (2024) M. Besta, N. Blach, A. Kubicek, R. Gerstenberger, M. Podstawski, L. Gianinazzi, J. Gajda, T. Lehmann, H. Niewiadomski, P. Nyczyk, and T. Hoefler Graph of thoughts: solving elaborate problems with large language models. In Proceedings of the Thirty-Eighth AAAI Conference on Artificial Intelligence and Thirty-Sixth Conference on Innovative Applications of Artificial Intelligence and Fourteenth Symposium on Educational Advances in Artificial Intelligence, AAAI’24/IAAI’24/EAAI’24. External Links: ISBN 978-1-57735-887-9, Link, Document Cited by: §2.4.
  • Bhattaram et al. (2025) A. Bhattaram, J. Chung, S. Chung, R. Gupta, J. Ramamoorthy, K. Gullapalli, D. Marculescu, and D. Stamoulis GeoFlow: agentic workflow automation for geospatial tasks. In Proceedings of the 33rd ACM International Conference on Advances in Geographic Information Systems, SIGSPATIAL ’25, New York, NY, USA, pp. 1150–1153. External Links: ISBN 9798400720864, Link, Document Cited by: §1, §2.5, Table 4, §5.
  • Chen et al. (2026) Z. Chen, H. Wang, J. Yao, J. Zhang, P. Ghamisi, J. Zhou, P. M. Atkinson, and B. Zhang CangLing-knowflow: a unified knowledge-and-flow-fused agent for comprehensive remote sensing applications. External Links: 2512.15231, Link Cited by: §5.
  • DeepSeek-AI (2026) DeepSeek-AI DeepSeek-v4: towards highly efficient million-token context intelligence. Cited by: §2.4.
  • Erdogan et al. (2025) L. E. Erdogan, N. Lee, S. Kim, S. Moon, H. Furuta, G. Anumanchipalli, K. Keutzer, and A. Gholami PLAN-and-act: improving planning of agents for long-horizon tasks. In Proceedings of the 42nd International Conference on Machine Learning, ICML’25. Cited by: §2.4.
  • Feng et al. (2026) P. Feng, Z. Lv, J. Ye, X. Wang, X. Huo, J. Yu, W. Xu, W. Zhang, L. Bai, C. He, and W. Li Earth-agent: unlocking the full landscape of earth observation with agents. In International Conference on Learning Representations, External Links: Link Cited by: §5.
  • Froger et al. (2025) R. Froger, P. Andrews, M. Bettini, A. Budhiraja, R. S. Cabral, V. Do, E. Garreau, J. Gaya, H. Laurençon, M. Lecanu, K. Malkan, D. Mekala, P. Ménard, G. M. Bertran, U. Piterbarg, M. Plekhanov, M. Rita, A. Rusakov, V. Vorotilov, M. Wang, I. Yu, A. Benhalloum, G. Mialon, and T. Scialom ARE: scaling up agent environments and evaluations. External Links: 2509.17158, Link Cited by: §2.3, §2.6, §2.7, §2.8, Table 6.
  • Gou et al. (2024) Z. Gou, Z. Shao, Y. Gong, yelong shen, Y. Yang, N. Duan, and W. Chen CRITIC: large language models can self-correct with tool-interactive critiquing. In The Twelfth International Conference on Learning Representations, External Links: Link Cited by: §2.4.
  • Humburg (2019) D. Humburg Chapter 38: determining harvest losses in soybeans. In South Dakota State University iGrow Soybean Best Management Practices, Cited by: §2.2.
  • Lee et al. (2025) C. Lee, V. Paramanayakam, A. Karatzas, Y. Jian, M. Fore, H. Liao, F. Yu, R. Li, I. Anagnostopoulos, and D. Stamoulis Multi-agent geospatial copilots for remote sensing workflows. In IGARSS 2025 - 2025 IEEE International Geoscience and Remote Sensing Symposium, Vol. , pp. 1084–1089. External Links: Document Cited by: §1, §3.
  • Li et al. (2025) K. Li, J. Wang, Z. Wang, H. Qiao, W. Zhang, D. Meng, and X. Cao Designing domain-specific agents via hierarchical task abstraction mechanism. External Links: 2511.17198, Link Cited by: Table 4.
  • Michelakis et al. (2025) P. Michelakis, Y. Hadjiyianni, and D. Stamoulis CORE: full-path evaluation of llm agents beyond final state. External Links: 2509.20998, Link Cited by: §2.7.
  • Monsi and Saeki (2005) M. Monsi and T. Saeki On the factor light in plant communities and its importance for matter production. Annals of botany 95 (3), pp. 549–567. Cited by: §2.2.
  • Monteith (1977) J. L. Monteith Climate and the efficiency of crop production in britain. Philosophical transactions of the royal society of London. B, Biological Sciences 281 (980), pp. 277–294. Cited by: §2.2.
  • Patil et al. (2025) S. G. Patil, H. Mao, F. Yan, C. C. Ji, V. Suresh, I. Stoica, and J. E. Gonzalez The berkeley function calling leaderboard (bfcl): from tool use to agentic evaluation of large language models. In Proceedings of the 42nd International Conference on Machine Learning, ICML’25. Cited by: Table 6.
  • Qu et al. (2026) A. Qu, P. Michelakis, Y. Hadjiyianni, F. Li, J. Jiang, D. Stamoulis, and J. Liu Full-season agent evaluation in soybean farm operations under real-world agricultural process dynamics. In Second Workshop on Agents in the Wild: Safety, Security, and Beyond, External Links: Link Cited by: §1.
  • Seo and Lee (2026) S. Seo and K. Lee Density-driven multidrone coordination for efficient farm coverage and management in smart agriculture. IEEE Transactions on Control Systems Technology 34 (2), pp. 711–724. External Links: ISSN 2374-0159, Link, Document Cited by: §1.
  • Shabbir et al. (2026) A. Shabbir, M. A. Munir, A. Dudhane, M. U. Sheikh, M. H. Khan, P. Fraccaro, J. B. Moreno, F. S. Khan, and S. Khan ThinkGeo: evaluating tool-augmented agents for remote sensing tasks. External Links: 2505.23752, Link Cited by: §5.
  • Shinn et al. (2023) N. Shinn, F. Cassano, A. Gopinath, K. Narasimhan, and S. Yao Reflexion: language agents with verbal reinforcement learning. In Proceedings of the 37th International Conference on Neural Information Processing Systems, NIPS ’23, Red Hook, NY, USA. Cited by: §2.4.
  • Siler and Singh (2022) T. B. Siler and M. P. Singh Optimal soybean maturity group selection is influenced by planting date in northern production systems. Crop Science 62 (6), pp. 2462–2475. Cited by: §2.2.
  • Singh et al. (2024) S. Singh, M. Fore, and D. Stamoulis GeoLLM-engine: a realistic environment for building geospatial copilots. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, pp. 585–594. Cited by: §2.5, Table 4.
  • Steduto et al. (2009) P. Steduto, T. C. Hsiao, D. Raes, and E. Fereres AquaCrop—the fao crop model to simulate yield response to water: i. concepts and underlying principles. Agronomy journal 101 (3), pp. 426–437. Cited by: §2.2.
  • Tan et al. (2026) R. Tan, B. Peng, Z. Yang, H. Cheng, O. Mees, T. Zhao, A. Tupini, I. Meijier, Q. Wu, Y. Yang, L. Liden, Y. Gu, S. Zhang, X. Liu, L. Wang, M. Pollefeys, Y. J. Lee, and J. Gao Multimodal reinforcement learning with adaptive verifier for ai agents. External Links: 2512.03438, Link Cited by: §2.4.
  • Wang et al. (2025) H. Wang, Y. Guan, F. Meng, C. Zhao, L. Yan, Y. Yang, and J. Jiang Agri-CM3{}^{3}: a Chinese massive multi-modal, multi-level benchmark for agricultural understanding and reasoning. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, pp. 11729–11754. External Links: Link, Document, ISBN 979-8-89176-251-0 Cited by: §2.2, §2.5, §5.
  • Wu et al. (2024) Q. Wu, G. Bansal, J. Zhang, Y. Wu, B. Li, E. Zhu, L. Jiang, X. Zhang, S. Zhang, J. Liu, A. H. Awadallah, R. W. White, D. Burger, and C. Wang AutoGen: enabling next-gen LLM applications via multi-agent conversations. In First Conference on Language Modeling, External Links: Link Cited by: §2.4.
  • Xu et al. (2023) B. Xu, Z. Peng, B. Lei, S. Mukherjee, Y. Liu, and D. Xu ReWOO: decoupling reasoning from observations for efficient augmented language models. External Links: 2305.18323, Link Cited by: §2.4.
  • Xu et al. (2025) Z. Xu, J. Xu, M. Zhang, P. Wang, C. Deng, and C. Liu Multimodal agricultural agent architecture (ma3): a new paradigm for intelligent agricultural decision-making. External Links: 2504.04789, Link Cited by: §1.
  • Yan et al. (2026) L. Yan, H. Wang, C. Tang, H. Liu, T. Sun, L. Liu, Y. Guan, and J. Jiang Agrieval: a comprehensive chinese agricultural benchmark for large language models. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 40, pp. 34205–34213. Cited by: §1.
  • Yang et al. (2025) A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv, C. Zheng, D. Liu, F. Zhou, F. Huang, F. Hu, H. Ge, H. Wei, H. Lin, J. Tang, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Zhou, J. Lin, K. Dang, K. Bao, K. Yang, L. Yu, L. Deng, M. Li, M. Xue, M. Li, P. Zhang, P. Wang, Q. Zhu, R. Men, R. Gao, S. Liu, S. Luo, T. Li, T. Tang, W. Yin, X. Ren, X. Wang, X. Zhang, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Zhang, Y. Wan, Y. Liu, Z. Wang, Z. Cui, Z. Zhang, Z. Zhou, and Z. Qiu Qwen3 technical report. External Links: 2505.09388, Link Cited by: §2.4.
  • Yao et al. (2023) S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao ReAct: synergizing reasoning and acting in language models. In International Conference on Learning Representations (ICLR), External Links: Link Cited by: §2.4.
  • Zhang et al. (2026) Z. Zhang, J. Zhang, H. Liu, Q. Lv, J. Yang, K. Cai, and K. Wang AgriWorld:a world tools protocol framework for verifiable agricultural reasoning with code-executing llm agents. External Links: 2602.15325, Link Cited by: §1.
  • Zhou et al. (2024) A. Zhou, K. Yan, M. Shlapentokh-Rothman, H. Wang, and Y. Wang Language agent tree search unifies reasoning, acting, and planning in language models. In Proceedings of the 41st International Conference on Machine Learning, ICML’24. Cited by: §2.4.
  • Zuzuárregui et al. (2025) M. A. Zuzuárregui, M. M. Toslak, and S. Carpin One for all: llm-based heterogeneous mission planning in precision agriculture. External Links: 2506.10106, Link Cited by: §1.