arXiv is now an independent nonprofit! Learn more
License: CC BY 4.0
arXiv:2609.01260v1 [cs.AI] 01 Sep 2026

Dual Process Motion Planning

Jiayi Yan    Francesco Fabiano    Alessandro Abate
Abstract

Robotic systems are deeply embedded in both industry and everyday life, where they are expected to act with speed, precision, and reliability. Classical control and planning methods have long delivered strong guarantees, but often at the cost of computational efficiency and adaptability. More recently, learning-based approaches have shown promise in overcoming these limitations, enabling agents to leverage experience to accelerate decision-making and address previously intractable problems. In this work, we bridge these two approaches through a neuro-symbolic perspective on nonlinear motion planning. Inspired by the Thinking Fast and Slow paradigm, we introduce a dual-process architecture that combines the strengths of robust reasoning and learning. Our framework integrates state-of-the-art symbolic solvers as a “System-2” component with experience-driven “System-1” modules. A metacognitive controller dynamically orchestrates their interaction, selecting when to rely on fast intuition versus slower, more precise reasoning. By evaluating the framework across diverse nonlinear benchmark environments, we demonstrate that this architecture yields consistent gains in planning efficiency, accuracy, and generalization, while promoting reuse across tasks. The results suggest that tightly coupling learning with structured reasoning offers a scalable path toward more capable and adaptive robotic systems.

1The Chinese University of Hong Kong, Shenzhen

jiayiyan@link.cuhk.edu.cn

2University of Oxford

francesco.fabiano@cs.ox.ac.uk, alessandro.abate@cs.ox.ac.uk

Introduction

Motion planning is a core component of robotic systems and underpins a wide range of applications, including autonomous driving (Paden et al. 2016) and multi-agent coordination (Hegde and Panagou 2016). Its primary objective is to compute a high-quality, collision-free trajectory that connects a start state to a goal state while respecting system dynamics and constraints. However, real-world environments are often high-dimensional, continuous, and dynamically constrained, which makes motion planning computationally challenging, especially under strict real-time requirements (LaValle 2006; Karaman and Frazzoli 2011).

Classical motion planning methods can be broadly categorized into experience-based planning using offline trajectory libraries and neural networks, and online optimization-based planning. Experience-based methods, such as those leveraging neural networks, amortize planning into a learned policy or trajectory generator that maps observations directly to controls or waypoints at runtime, enabling fast online inference and interpolation beyond a finite trajectory library, albeit typically with weaker guarantees (Ichter et al. 2017). Their effectiveness depends critically on selecting a trajectory that matches the current scenario, which is difficult in complex environments, hence often leading to suboptimal or unsafe performance.

In contrast, online optimization-based methods, including Model Predictive Control (MPC)  (Mayne et al. 2000) and Control Barrier Function (CBF)-based safety filters  (Ames et al. 2019), compute trajectories by solving constrained optimization problems that explicitly account for system dynamics and environmental constraints. While these methods reliably provide solutions, they incur in computational overhead, limiting their applicability in time-critical settings.

To bridge this gap, recent work has explored learning-augmented motion planning, where data-driven models are used to improve efficiency and performance (Mansard et al. 2018; Bjelonic et al. 2022; Carvalho et al. 2024). These methods improve convergence and solution quality, but they typically rely on a fixed pipeline that often hinders their efficiency. To address this limitation, we draw inspiration from a recent AI paradigm Fabiano et al. (2025) that is in turn informed by the dual-system theory Kahneman (2011). We propose Dual-MP, a dual-system motion-planning framework that exploits between fast experience-based and slower online solving.

We instantiate this idea within a SOFAI-style architecture (Pallagani et al. 2025). Dual-MP inherits the modular nature of SOFAI: it uses both fast, experience-based and slow, deliberate solvers. We refer to the former category as System-1 (S1) and to the latter as System-2 (S2). These are arbitrated through a metacognitive (MC) agent. Dual-MP is equipped with a Neural Network S1: a neural policy trained from successful trajectories, and two general S2 online solvers: one based on MPC and one on CBF. The code used in this work is available online11 1 https://github.com/verayannn/System-1-and-System-2-in-Motion-Planning.

We summarize our main contributions below:

  • •

    We propose Dual-MP, a modular S1/S2 architecture for nonlinear motion planning that arbitrates fast neural planning and deliberate symbolic solving.

  • •

    We instantiate the framework with a neural S1 policy and two nonlinear S2 solvers, based on MPC and CBFs, under a common MC planner interface.

  • •

    We add continual learning, where successful trajectories are reused to retrain S1 and improve future fast planning.

  • •

    We evaluate the framework across diverse nonlinear benchmark families, reporting success, runtime, S1/S2 usage, and trajectory quality.

Related work

Classical Motion Planning.

Current classical motion planning methods are primarily divided into two popular categories: experience-based planning using neural networks and online optimization-based planning.

Experience-based planning typically uses neural networks to amortize planning from previously solved problems, learning a direct mapping from observations, goals, and local state information to actions, waypoints, or full trajectories (Ichter et al. 2017; Fishman et al. 2023). Such methods can provide low-latency inference and strong empirical performance in complex environments by reusing structure learned from expert demonstrations, simulation data, or successful planning experience (Wang et al. 2021). While effective, such approaches typically do not preserve explicit models of dynamics, safety constraints, or solver confidence at inference time, which limits their interpretability and their ability to decide when the output should be trusted.

In contrast to experience-based planning methods, online optimization-based methods, including MPC and CBF-based safety filters, compute trajectories or controls by solving constrained optimization problems at runtime (Mayne et al. 2000; Ames et al. 2017). These methods directly incorporate system dynamics and environmental constraints, yielding high-quality and dynamically feasible solutions. However, they often incur significant computational overhead, which makes real-time deployment challenging in complex or high-dimensional scenarios.

Learning-Augmented Motion Planning.

To mitigate this, recent work seeks to combine the efficiency of offline planning with the accuracy of online optimization through learning-augmented hybrid approaches. A common strategy is to use learned models or offline datasets to warm-start online solvers. For example, Memory-of-Motion learns a mapping from task descriptors to state-control trajectories and uses the resulting memory to initialize nonlinear predictive control (Mansard et al. 2018). Bjelonic et al. (2022) use offline motion libraries as reference costs for online MPC, which allows long-horizon offline behaviors to be executed through short-horizon feedback optimization. Transformer-based motion planners learn to restrict or guide the search space from prior data (Johnson et al. 2022), while diffusion-based planners learn multimodal trajectory priors that can be sampled or adapted during planning (Carvalho et al. 2024). Closely related to our dual-system motivation, Fridovich-Keil et al. (2018) propose a “Planning, Fast and Slow” framework in which offline safety computation enables safe switching among online planners. Their method provides a strong safety-aware planning module; in our terminology, such a framework can be viewed as a possible instantiation of S2.

Despite these advances, a key limitation of existing hybrid approaches is that they usually invoke the online solver regardless of how well the offline or learned prior matches the current scenario. This leads to unnecessary computation in cases where a retrieved or neural trajectory is already sufficient. More fundamentally, these methods lack a principled mechanism for deciding when online optimization is required. Dual-MP is designed to address exactly this allocation problem.

Background

Nonlinear Model Predictive Control

Model Predictive Control (MPC) is a widely used optimal control framework for motion planning under dynamic and environmental constraints (Mayne et al. 2000). At each time step, MPC solves a finite-horizon optimization problem to compute a control sequence that minimizes a cost function while satisfying system dynamics and constraints.

Consider nonlinear dynamics x˙=f⁡(x,u),\dot{x}=f(x,u), or, after discretization, xk+1=Fd​(xk,uk),x_{k+1}=F_{d}(x_{k},u_{k}), where xk∈ℝnx_{k}\in\mathbb{R}^{n} is the state and uk∈ℝmu_{k}\in\mathbb{R}^{m} is the control. Given the current state x0x_{0} and goal xgx_{g}, nonlinear MPC solves minu0:Nh−1∑k=0Nh−1ℓ(xk,uk;xg)+Vf(xNh;xg)\min_{u_{0:N_{h}-1}}\sum_{k=0}^{N_{h}-1}\ell(x_{k},u_{k};x_{g})+V_{f}(x_{N_{h}};x_{g}) subject to xk+1=Fd​(xk,uk),xk∈𝒳free,uk∈𝒰.x_{k+1}=F_{d}(x_{k},u_{k}),\qquad x_{k}\in\mathcal{X}_{\mathrm{free}},\qquad u_{k}\in\mathcal{U}. The stage cost is typically chosen as ℓ⁡(xk,uk,xg)=‖xk−xg‖Q2+‖uk‖R2,\ell(x_{k},u_{k};x_{g})=\|x_{k}-x_{g}\|_{Q}^{2}+\|u_{k}\|_{R}^{2}, where Q⪰0Q\succeq 0 and R≻0R\succ 0. Obstacle avoidance is encoded through nonlinear state constraints defining the collision-free set 𝒳free\mathcal{X}_{\mathrm{free}} MPC is accurate because it optimizes over future trajectories, but repeatedly solving the resulting nonlinear program can be computationally expensive.

Control Barrier Functions

Control Barrier Functions (CBFs) provide a complementary approach for enforcing safety constraints in control systems (Ames et al. 2019). Unlike MPC, the CBF-QP does not optimize an entire future trajectory; instead, it acts as a safety filter that enforces local obstacle-avoidance constraints.

For a nonlinear control-affine system x˙=f⁡(x)+g⁡(x)​u,\dot{x}=f(x)+g(x)u, let the safe set for obstacle ii be 𝒞i={x:hi​(x)≥0},\mathcal{C}_{i}=\{x:h_{i}(x)\geq 0\}, where hi​(x)h_{i}(x) is positive outside the obstacle and negative inside it. Forward invariance of 𝒞i\mathcal{C}_{i} can be encouraged by enforcing Lf​hi​(x)+Lg​hi​(x)​u+γ​hi​(x)≥0,L_{f}h_{i}(x)+L_{g}h_{i}(x)u+\gamma h_{i}(x)\geq 0, where Lf​hiL_{f}h_{i} and Lg​hiL_{g}h_{i} are Lie derivatives and γ>0\gamma>0 controls how aggressively the controller moves away from the safety boundary.

At each timestep, the CBF controller solves the quadratic program minu⁡‖u−uref‖2\min_{u}\|u-u_{\mathrm{ref}}\|^{2} subject to Lf​hi​(x)+Lg​hi​(x)​u+γ​hi​(x)≥0,∀i,u∈𝒰.L_{f}h_{i}(x)+L_{g}h_{i}(x)u+\gamma h_{i}(x)\geq 0,\quad\forall i,\qquad u\in\mathcal{U}. Here, urefu_{\mathrm{ref}} is a nominal control input, such as a goal-directed command or a command proposed by the neural System-1 policy.

Thinking Fast and Slow in AI

Kahneman (2011) describe human decision making as the interaction between two complementary processes: a fast, intuitive, experience-driven System-1 and a slower, deliberative, reasoning-based System-2. Recent AI architectures, including SOFAI (Pallagani et al. 2025; Fabiano et al. 2025), adapt this dual-process principle to machine decision-making by combining fast and slow solvers with a metacognitive agent that arbitrates them (Ganapini et al. 2021). These architectures have proven successful in tackling settings closely related to motion planning, such as classical planning and constrained grid navigation, as shown in Fabiano et al. (2025).

Formally, SOFAI defines a decision architecture composed of three components: (i) a set of fast S1 solvers, typically data-driven and experience-based; (ii) a set of slow S2 solvers, based on explicit symbolic reasoning; and (iii) a centralized metacognitive (MC) controller. Incoming problem instances automatically trigger one or more S1 solvers, which produce candidate solutions together with confidence estimates. The MC controller then decides whether to accept the S1 proposal or invoke an S2 solver. This decision is performed in two stages: a lightweight assessment that evaluates whether the expected solution quality satisfies a task-dependent threshold under resource constraints, followed, when necessary, by a cost-benefit comparison between S1 and S2 execution. S2 reasoning is activated only if its expected improvement compensates for the additional computational cost. A key aspect of this framework is that S1 behavior is not static. Through metacognition, solutions produced or validated by S2 can be used to improve the fast solver over time, effectively distilling deliberative reasoning into reactive policies. In this sense, the architecture supports an iterative refinement process in which expensive symbolic reasoning is gradually amortized into efficient inference, enabling the system to adapt to recurring problem distributions while reducing reliance on S2 computation.

Dual-MP

As mentioned above, Dual-MP solves motion planning through metacognitive arbitration. A planning query is defined as q=(f,𝒪,x0,xg,𝒳,𝒰),q=(f,\mathcal{O},x_{0},x_{g},\mathcal{X},\mathcal{U}), where ff denotes the nonlinear system dynamics, 𝒪\mathcal{O} is the obstacle map, x0x_{0} and xgx_{g} are the start and goal states, 𝒳\mathcal{X} is the workspace, and 𝒰\mathcal{U} is the admissible control set. In this work, we consider discrete-time nonlinear reach–avoid problems of the form xt+1=Fd​(xt,ut),x_{t+1}=F_{d}(x_{t},u_{t}), where FdF_{d} is the discretized nonlinear dynamics used by the planner.

In the following, we define the main components of Dual-MP, namely the S1 and S2 solvers, as well as the functionalities required from the MC module. In particular, as detailed in Section Results, we consider two basic Dual-MP configurations obtained by combining the S1 solver with two S2 solvers.

Neural S1

Our System-1 is implemented as a neural reactive policy. This follows the common approach of imitation learning to predict low-level controls from local environment observations and goal information (Ichter et al. 2017).

Given the current rollout context, local obstacle information, nonlinear dynamics features, and the goal direction, the policy predicts a control input directly: utS1=πθ​(ct,s,dt,gt),u_{t}^{\mathrm{S1}}=\pi_{\theta}(c_{t},s,d_{t},g_{t}), where ctc_{t} is a fixed-length window of recent states in a local coordinate frame, ss encodes the local obstacle situation, dtd_{t} encodes the nonlinear dynamics at the current state, and gtg_{t} is the local goal vector.

The policy uses a lightweight convolutional architecture. A one-dimensional convolutional encoder processes the recent state context ctc_{t}, while separate multilayer encoders process the situation vector, dynamics features, and goal vector. The resulting features are concatenated and passed through a feedforward control head:

ht=[ϕc​(ct),ϕs​(s),ϕd​(dt),ϕg​(gt)],utS1=πθ​(ht).h_{t}=[\phi_{c}(c_{t}),\phi_{s}(s),\phi_{d}(d_{t}),\phi_{g}(g_{t})],\qquad u_{t}^{\mathrm{S1}}=\pi_{\theta}(h_{t}).

This design keeps S1 inexpensive at runtime while still conditioning the control prediction on both geometry and nonlinear dynamics.

As mentioned, S1 is trained by imitation from successful System-2 trajectories. In particular, the bootstrap demonstrations are solved offline by System-2, and the System-2 recoveries are collected online during continual learning. For each state-control pair, the supervised target is the expert control ut⋆u_{t}^{\star}, with auxiliary single-step losses that encourage accurate one-step motion, correct direction, and progress toward the goal, and a multi-step rollout term that penalizes the drift accumulated when the policy is run in closed loop:

ℒ⁡(θ)=\displaystyle\mathcal{L}(\theta)={} λu​ℓu+λx​ℓx+λdir​ℓdir+λspd​ℓspd\displaystyle\lambda_{u}\ell_{u}+\lambda_{x}\ell_{x}+\lambda_{\mathrm{dir}}\ell_{\mathrm{dir}}+\lambda_{\mathrm{spd}}\ell_{\mathrm{spd}}
+λprog​ℓprog+λroll​ℓroll+λsmo​ℓsmo+λobs​ℓobs.\displaystyle+\lambda_{\mathrm{prog}}\ell_{\mathrm{prog}}+\lambda_{\mathrm{roll}}\ell_{\mathrm{roll}}+\lambda_{\mathrm{smo}}\ell_{\mathrm{smo}}+\lambda_{\mathrm{obs}}\ell_{\mathrm{obs}}.

Here, ℓu\ell_{u} is a Huber control-imitation loss, ℓx\ell_{x} penalizes one-step prediction error under the nonlinear dynamics, ℓdir\ell_{\mathrm{dir}} encourages alignment with the expert displacement, ℓspd\ell_{\mathrm{spd}} matches step magnitude, and ℓprog\ell_{\mathrm{prog}} penalizes insufficient progress toward the goal. The remaining three terms are evaluated on a differentiable HH-step rollout of the policy under the nonlinear dynamics, run once every NN supervised batches (H=N=8H=N=8): ℓroll\ell_{\mathrm{roll}} penalizes deviation of the rolled-out positions from the expert positions, ℓsmo\ell_{\mathrm{smo}} penalizes large control changes, and ℓobs\ell_{\mathrm{obs}} penalizes proximity to obstacles along the rollout.

At execution time, S1 rolls out the learned policy under the nonlinear dynamics. A lightweight one-step safety filter selects among the predicted control, goal-directed controls, scaled controls, and zero control, keeping only candidates whose next segment is collision-free. The resulting trajectory is then verified for collision freedom and goal reach. In the SOFAI setting, S2 is invoked only when the S1 rollout fails this verification.

S2 Solvers: CBF and MPC

Dual-MP can be instantiated with two state-of-the-art S2 solvers:

Let us note that both solvers are implemented using their standard formulations from the literature.

The Metacognitive Module

Dual-MP follows the SOFAI principle of coordinating fast S1 solvers and slower deliberative S2 solvers through a metacognitive arbitration module (Pallagani et al. 2025). Depending on the Dual-MP configuration, the S1 and S2 components are instantiated differently; in our experiments, the two configurations are obtained by pairing the S1 solver with two S2 solvers introduced above.

For each new scenario, the metacognitive module first invokes the configured S1 solver. The proposed rollout is accepted only if it passes verification (collision freedom and goal reach - cf. above). If the S1 proposal is rejected, Dual-MP invokes the configured S2 fallback, provided that the available time budget is sufficient relative to the estimated problem difficulty. If S2 does not return a satisfactory solution within the budget, Dual-MP falls back to the best available S1 proposal. Successful S2 rollouts are stored in memory and periodically used to improve S1.

The metacognitive module relies on three quantities to properly arbitrate the two systems: problem difficulty, solution correctness, and solver confidence. Difficulty is a scenario-level quantity used to characterize the planning instance; correctness measures the quality of a proposed rollout; and confidence estimates whether an S1 proposal is worth verifying. Following the SOFAI principle, S2 solvers are treated as deliberative solvers and assigned confidence 11, while S1 confidence is computed as described above.

Correctness.

Given a rollout τ={xt}t=0T\tau=\{x_{t}\}_{t=0}^{T}, correctness is set to 11 if the trajectory reaches the goal and remains collision-free. Otherwise, we assign a soft correctness score based on path length, terminal goal error, and collision penalty: Correct⁡(τ,q)=11+L⁡(τ)+eg​(τ)+λcol​Ncol​(τ),\mathrm{Correct}(\tau;q)=\frac{1}{1+L(\tau)+e_{g}(\tau)+\lambda_{\mathrm{col}}N_{\mathrm{col}}(\tau)}, where L⁡(τ)=∑t=0T−1‖xt+1−xt‖1,eg​(τ)=‖xT−xg‖2.L(\tau)=\sum_{t=0}^{T-1}\|x_{t+1}-x_{t}\|_{1},\qquad e_{g}(\tau)=\|x_{T}-x_{g}\|_{2}. Here, Ncol​(τ)N_{\mathrm{col}}(\tau) is a bounded collision penalty and λcol\lambda_{\mathrm{col}} controls the relative cost of unsafe states. In the experiments reported in this paper, we use a strict correctness threshold of 1.01.0, requiring both goal reaching and collision avoidance. More permissive thresholds are supported by the SOFAI framework and may be useful under strict time constraints.

Difficulty.

We define the difficulty of a query scenario using simple features to maintain a lightweight process. Let 𝒪\mathcal{O} denote the set of rectangular obstacles and let 𝒲\mathcal{W} denote the workspace. We first compute the obstacle occupancy ηocc=∑Oj∈𝒪area⁡(Oj)area⁡(𝒲)+ϵ.\eta_{\mathrm{occ}}=\frac{\sum_{O_{j}\in\mathcal{O}}\mathrm{area}(O_{j})}{\mathrm{area}(\mathcal{W})+\epsilon}. Let xmid=(x0+xg)/2x_{\mathrm{mid}}=(x_{0}+x_{g})/2 be the midpoint between start and goal, and let dclear=minOj∈𝒪⁡dist⁡(xmid,Oj)d_{\mathrm{clear}}=\min_{O_{j}\in\mathcal{O}}\mathrm{dist}(x_{\mathrm{mid}},O_{j}) be a clearance estimate from this midpoint to the nearest obstacle. The difficulty is then

D⁡(q)\displaystyle D(q) =ηocc​(dclear+ϵ)−1​(1+log⁡(1+T))\displaystyle=\eta_{\mathrm{occ}}\left(d_{\mathrm{clear}}+\epsilon\right)^{-1}\left(1+\log(1+T)\right)
×(1+min⁡{10,(ϵg+ϵ)−1})\displaystyle\quad\times\left(1+\min\left\{10,\left(\epsilon_{g}+\epsilon\right)^{-1}\right\}\right)
×(1+min⁡{10,(ϵsafe+ϵ)−1}),\displaystyle\quad\times\left(1+\min\left\{10,\left(\epsilon_{\mathrm{safe}}+\epsilon\right)^{-1}\right\}\right),

where TT is the planning horizon, ϵg\epsilon_{g} is the goal tolerance, and ϵsafe\epsilon_{\mathrm{safe}} is the collision margin. The first factor captures geometric clutter and clearance, while the remaining factors account for horizon length, goal precision, and safety strictness.

SOFAI Continual Learning

A key advantage of SOFAI-inspired architectures is that S1 solvers can improve over time. Dual-MP supports this by storing successfully generated S2 trajectories and using them to enrich the corresponding S1 solver. Specifically, when S2 successfully solves a query, the resulting trajectory is added to the online dataset. For Neural S1, accumulated S2 trajectories are distilled periodically after a fixed number of episodes by fine-tuning the policy on both the original offline dataset and the newly collected S2 successes.

Experimental Evaluation

All experiments were run on an Intel Xeon Gold 6248 CPU server with 40 cores and 125 GiB of memory, running Ubuntu 22.04.5 LTS. The source code, benchmark data, and additional details, including offline training times, will be provided in the appendix for space reasons.

We evaluate Dual-MP on reach–avoid motion-planning tasks in two-dimensional continuous environments. Each instance is defined by nonlinear dynamics xt+1=Fd​(xt,ut),x_{t+1}=F_{d}(x_{t},u_{t}), a start state x0x_{0}, a goal state xgx_{g}, rectangular obstacles 𝒪\mathcal{O}, control bounds 𝒰\mathcal{U}, and workspace bounds 𝒳\mathcal{X}. A rollout counts as successful only if it is both collision-free and goal-reaching; no partial credit is given. Our experiments address four questions:

  • •

    Q1: Can Dual-MP improve state-of-the-art planning?

  • •

    Q2: Does Dual-MP preserve reliability?

  • •

    Q3: How does continual learning affect S1?

  • •

    Q4 Does warm-starting System-2 from the rejected S1 trajectory help System-2 solving?

We evaluate our approach on six benchmark families, each stressing a different aspect of reach–avoid planning:

  • •

    Large sparse (LS): a few large obstacles; tests long-horizon obstacle avoidance in an open workspace.

  • •

    Dense clutter (DC): many small rectangles; tests local collision avoidance under frequent re-planning pressure.

  • •

    Serial walls (SW): several wall-with-gap structures in sequence; tests repeated constrained passages, where a single missed gap is unrecoverable.

  • •

    Maze branching (MB): intersecting wall structures with multiple candidate routes; tests global route commitment.

  • •

    Long slalom (LSM): a long corridor with alternating offset barriers; tests sustained tracking over a horizon far longer than the MPC prediction window.

  • •

    Bugtrap (BT): concave trap layouts; tests robustness to locally attractive but globally poor routes.

Instances within a family share the workspace scale, the start/goal pair, and the reach–avoid objective, and vary in obstacle layout and in the parameters of the nonlinear drift field. For every family, we generate three disjoint instance sets from separate seeds: a training set of 100100 instances used only to bootstrap S1, an evaluation set of 500500 instances used for the results in Tables 2 and 3, and a held-out probe set of 500500 instances used exclusively to measure the effect of continual learning. The probe set is never trained on and is identical across all blocks and all methods, so probe curves are directly comparable over time.

All of the generated instances are run on six different solvers/Dual-MP configurations, listed in Table 1. The three Dual-MP variants differ only in the System-2 they escalate to and in whether that solver is warm-started from the rejected S1 trajectory.

Table 1: Planners configuration.
Name System-1 System-2
NN Neural policy –
CBF – CBF
MPC – MPC
DMP-CBF Neural policy CBF
DMP-MPC Neural policy MPC
DMP-Warm Neural policy MPC (S1 warm start)

The various approaches are then evaluated on three main metrics: (i) Success rate, which is the fraction of instances whose returned rollout is collision-free and reaches the goal. (ii) Runtimewhich represents the time to find a solution: for Dual-MP it is the sum of the S1 call and, when escalation occurs, the S2 call, so escalated instances are charged for both. We report the mean and the 9090th percentile. (iii) Trajectory Qualityevaluated only on successful rollouts, so that it never trades off against the success rate. We use a duration-invariant index that is the geometric mean of three sub-scores, each in (0,1](0,1] and each normalised against a scenario-defined rather than a solver-defined reference: Q=(ηpath⋅ηsmooth⋅ηclear)1/3.Q\;=\;\bigl(\eta_{\text{path}}\cdot\eta_{\text{smooth}}\cdot\eta_{\text{clear}}\bigr)^{1/3}. ηpath=Lref/L\eta_{\text{path}}=L_{\text{ref}}/L is the standard path-optimality ratio of the executed length LL against the shortest collision-free path LrefL_{\text{ref}}, obtained by grid search around the obstacles (Sucan et al. 2012). ηsmooth=SPARCmin-jerk/SPARC\eta_{\text{smooth}}=\mathrm{SPARC}_{\text{min-jerk}}/\mathrm{SPARC} is the spectral arc length of the speed profile expressed relative to an ideal minimum-jerk movement (Flash and Hogan 1985), so a minimum-jerk trajectory scores 1.01.0; spectral arc length is used because it is the only smoothness measure that is simultaneously valid and reliable (Balasubramanian et al. 2012). ηclear=min⁡(dmin/rbody,1)\eta_{\text{clear}}=\min(d_{\min}/r_{\text{body}},1) scores the worst obstacle clearance against one body radius, capped so that excess conservatism earns nothing. The geometric mean prevents a near-collision or a chattering control from being compensated elsewhere. Two properties matter for the comparisons below. First, QQ is invariant to how fast the trajectory is traversed, so an aggressive controller cannot buy quality by saturating its actuator. Second, because the reference is the scenario’s own shortest path, QQ is comparable across families of very different difficulty.

For Dual-MP we additionally report how many successes were produced by S1 alone versus by the S2 fallback.

Training Protocol

Bootstrap.

For each family and each System-2 solver, we solve the 100100 training instances with that solver and retain only the collision-free, goal-reaching trajectories. The base S1 policy is trained on those demonstrations by behaviour cloning augmented with a short differentiable rollout loss (horizon 88, applied every 88 supervised batches). Crucially, the bootstrap teacher is matched to the System-2 the arm will escalate to: DMP-CBF starts from a CBF-taught policy and DMP-MPC from an MPC-taught one.

Continual learning.

The 500500 evaluation instances are shuffled once (fixed seed) and presented in five sequential blocks of 100100. After each block, S1 is retrained, warm-started from the previous block’s weights, on the union of the bootstrap demonstrations and all System-2 recoveries collected so far. Recoveries are obtained in DAgger fashion (Ross et al. 2010): on instances where S1 was rejected, we relabel states visited by the S1 rollout with the System-2 solver, so the policy receives corrective actions on its own state distribution rather than only on the expert’s. A replay fraction of 0.60.6 keeps the fixed bootstrap demonstrations in the mixture and limits drift toward the increasingly hard, failure-biased recovery set. The retrained model is evaluated on the frozen probe set before the next block begins.

Results

We now present the experimental results of our paper and use these to answer the four research questions we introduced above. The main results are presented in Table 2, which reports results macro-averaged over the six families for the six configurations, and in Table 3 that breaks the same runs down per family.

Table 2: Aggregate results for all methods averaged over the six environment families (500500 instances per family).
Method Succ. (%) Mean RT (ms) P90 RT (ms) Mean QQ P90 QQ S1/S2 Solves
NN 36.4 230 305 0.601 0.817 182 / 0
CBF 82.0 657 1448 0.795 0.844 0 / 475
MPC 84.8 5708 7092 0.844 0.875 0 / 496
DMP-CBF 84.1 983 2279 0.734 0.850 204 / 217
DMP-MPC 85.1 4979 8086 0.780 0.877 141 / 285
DMP-Warm 83.6 5008 8161 0.782 0.877 140 / 278
Table 3: Per-family results on instance sets (n=500n=500). LS: large sparse, DC: dense clutter, SW: serial walls, MB: maze branching, LSM: long slalom, BT: bugtrap.
Method LS DC SW MB LSM BT
Success rate (%)
NN 82.8 74.4 12.0 19.2 13.0 17.0
CBF 96.0 94.0 59.7 65.0 98.2 78.8
MPC 98.2 94.0 77.5 69.5 76.2 93.2
DMP-CBF 97.0 94.8 65.6 70.0 98.2 79.2
DMP-MPC 98.2 95.2 77.8 69.6 76.6 93.4
DMP-Warm 98.2 95.2 77.2 67.8 74.6 88.8
Mean planning runtime (ms)
NN 112 135 230 255 460 185
CBF 121 234 934 961 1258 437
MPC 3439 3688 4886 6140 12316 3781
DMP-CBF 349 449 1131 1212 1974 783
DMP-MPC 1473 1670 4866 5988 12490 3385
DMP-Warm 1440 1715 4909 6048 12458 3476
Mean trajectory quality QQ
NN 0.781 0.750 0.580 0.484 0.425 0.586
CBF 0.817 0.817 0.778 0.772 0.801 0.787
MPC 0.860 0.861 0.836 0.826 0.827 0.853
DMP-CBF 0.794 0.763 0.696 0.679 0.741 0.729
DMP-MPC 0.761 0.716 0.793 0.802 0.817 0.792
DMP-Warm 0.762 0.709 0.814 0.790 0.817 0.796

(Q1) Can Dual-MP improve state-of-the-art planning?

Yes. Dual-MP improves the MPC baseline by filtering out easy instances with S1 and invoking MPC only when needed, which reduces the number of expensive optimisation calls while improving overall success. This yields the strongest success-runtime trade-off among the compared methods. The largest gains occur in large sparse and dense clutter environments, where S1 frequently solves instances without escalation. Gains are limited in serial walls, maze branching, and long slalom, where most instances still require MPC. By contrast, DMP-CBF is slower than CBF alone because the cost of an S1 rollout outweighs the savings from avoiding an already inexpensive fallback. Thus, the value of dual-process arbitration is determined primarily by the runtime gap between S1 and S2, rather than by S1 accuracy alone.

(Q2) Does Dual-MP preserve reliability?

Yes. Across the completed families, both Dual-MP variants match or improve the success rate of their respective S2 fallback (Tables 2 and 3). This follows from the verification gate: S1 is accepted only when its rollout is collision-free and reaches the goal, while rejected rollouts are delegated to S2. S1 can also solve some instances missed by the corresponding fallback solver, particularly in constrained environments.

The main trade-off is trajectory quality. MPC produces the highest-quality trajectories because it explicitly optimises path efficiency and obstacle clearance, whereas accepted S1 rollouts prioritise fast, feasible completion. Consequently, Dual-MP may return trajectories that are longer or pass closer to obstacles than solutions produced by MPC alone. Dual-MP is therefore most suitable when reducing online planning time is more important than retaining the trajectory quality of every MPC solution.

(Q3) How does continual learning affect S1?

Table 4 reports the cumulative changes in S1 success rate and trajectory quality, while Figures 1 and 2 show their evolution across training blocks.

Continual learning improves S1, but the type of improvement depends on the S2 teacher. The CBF-taught policy achieves substantial gains in success rate as fallback experience accumulates, increasing the fraction of instances solved directly by S1 and reducing reliance on CBF. The MPC-taught policy primarily improves trajectory quality and increases success in relatively open environments, such as large sparse and dense clutter, but provides limited or inconsistent success gains in more constrained families.

This difference is in the “teacher” capabilities: MPC demonstrations often contain long-horizon, globally route-dependent plans that a reactive policy with local observations cannot reliably reproduce after small closed-loop deviations. CBF, by contrast, provides short, local corrective actions that are easier for S1 to imitate.

Table 4: Continual learning on the held-out (500)-instance probe set. “Base” is the bootstrap policy before any online experience. B0–B4 show cumulative performance after each (100)-instance block. 𝚫\boldsymbol{\Delta} is the change from Base to B4, and 𝚫​𝑸\boldsymbol{\Delta Q} is the corresponding change in mean S1 trajectory quality.
Family Arm Base B0 B1 B2 B3 B4 𝚫\boldsymbol{\Delta} 𝚫​𝑸\boldsymbol{\Delta Q}
LS CBF 80.4 80.8 86.0 85.0 90.0 90.0 +9.6+9.6 +0.007+0.007
MPC 54.2 64.8 71.8 73.8 75.2 72.4 +18.2+18.2 +0.113+0.113
DC CBF 72.8 74.6 82.0 84.2 85.4 85.2 +12.4+12.4 −0.008-0.008
MPC 58.2 65.6 64.6 66.8 68.8 70.8 +12.6+12.6 +0.104+0.104
SW CBF 11.2 15.8 20.8 24.6 27.6 30.4 +19.2+19.2 −0.028-0.028
MPC 12.6 9.2 5.2 5.0 6.2 4.4 −8.2-8.2 +0.044+0.044
MB CBF 21.2 19.4 24.0 26.8 23.8 24.2 +3.0+3.0 +0.062+0.062
MPC 7.0 6.2 4.0 5.2 5.6 4.2 −2.8-2.8 +0.061+0.061
LSM CBF 13.0 16.8 16.6 18.8 19.0 24.8 +11.8+11.8 +0.006+0.006
MPC 0.6 0.8 0.6 2.0 2.0 2.6 +2.0+2.0 +0.041+0.041
BT CBF 19.2 23.0 28.8 30.6 36.4 42.4 +23.2+23.2 −0.020-0.020
MPC 22.2 19.4 23.6 25.0 24.6 27.2 +5.0+5.0 +0.085+0.085
Mean CBF 36.3 38.4 43.0 45.0 47.0 49.5 +13.2+13.2 +0.003+0.003
MPC 25.8 27.7 28.3 29.6 30.4 30.3 +4.5+4.5 +0.075+0.075
Refer to caption
Figure 1: Success rate of S1 on the 500-instance probe set as continual learning accumulates experience. Solid curves show the CBF- and MPC-taught policies.
Refer to caption
Figure 2: Mean quality QQ of successful S1 rollouts on the 500-instance probe set as continual learning accumulates experience. Quality is evaluated only on successful rollouts.

(Q4) Does warm-starting System-2 from the rejected S1 trajectory help System-2 solving?

Naive warm-starting from a rejected S1 rollout is not useful as shown in Table 5. Initialising MPC with the rejected trajectory does not give a consistent runtime benefit, and it can even steer the optimiser toward the same locally poor region that caused S1 to fail. Although warm-starting occasionally helps, the effect is unreliable and can reduce fallback success on the hardest instances. We therefore do not use rejected S1 rollouts as unconditional warm starts. While Dual-MP does compute an S1 confidence score, we did not tune it for warm-start selection because it was not a sufficiently reliable proxy for feasibility, so we omitted that mechanism. We leave this as future work.

Table 5: Warm-start ablation on instances escalated by both MPC arms. Times are mean System-2 solver time; “S2 solves” counts successful fallbacks.
Family nn Cold (s) Warm (s) 𝚫\boldsymbol{\Delta} S2 solves
LS 137 3.489 3.492 −0.003-0.003 128 / 128
DC 150 3.764 3.771 −0.007-0.007 126 / 126
SW 449 4.813 4.803 +0.010+0.010 339 / 337
MB 461 5.935 6.067 −0.133-0.133 310 / 300
LSM 477 11.096 11.057 +0.040+0.040 370 / 362
BT 364 3.870 4.001 −0.131-0.131 331 / 311

Discussion

Although Dual-MP is only an initial step toward dual-process motion planning, it delivers promising results; it outperforms the standalone baselines in the reported success and efficiency metrics. This suggests that verified arbitration is not just a conceptual framework, but a practical way to combine learned policies with optimization-based solvers.

The results show that Dual-MP is most effective when the two systems have distinct costs and complementary strengths. A fast S1 policy is most valuable when it avoids expensive MPC calls, but it offers little benefit when the fallback is already cheap, as with CBF. This means that success rate alone is not sufficient to evaluate a hybrid planner: the runtime balance between S1 and S2 largely determines whether arbitration is worthwhile. Exact rollout verification is central to this behavior. It allows S1 to produce low-latency feasible solutions without letting unsafe or incomplete predictions replace the fallback solver. When S1 fails, S2 takes over, and successful recoveries can then be reused for continual learning. This makes the architecture practical without requiring the neural policy to be perfectly safe in isolation, as long as it is paired with a correct fallback and a verification gate.

Similarly, we note that continual learning is most effective when the teacher’s behavior is compatible with the student’s policy class. CBF provides short, local corrections that improve S1 coverage, whereas MPC often produces longer, route-dependent behaviors that are harder for a reactive policy to reproduce reliably in closed loop. This suggests several promising directions for future work: (i) training S1 on MPC trajectories and pairing it with CBF, (ii) exploring other combinations of fast and slow components, and (iii) using a portfolio metacontroller to select among multiple S2 planners rather than relying on a single fallback. More generally, future work could optimize different objectives depending on the application, including planning time, trajectory quality, success rate, or even accepting slightly suboptimal solutions when that yields a better overall trade-off. In the same spirit, better knowledge transfer across environment families could reduce bootstrap cost, either by sharing an S1 policy across tasks or by making continual learning more selective when offline retraining budget is limited. Rejected S1 rollouts should likewise be used more selectively as MPC initializations; a stronger confidence or feasibility estimate could help determine when a proposal is close enough to serve as a useful warm start and when MPC should instead solve the problem from scratch.

Overall, this work provides an initial but promising step toward optimizing not only the individual tools, but also their composition. It shows that the interaction between learned policies, symbolic solvers, verification, and retraining can itself be designed and improved.

Conclusion

We presented Dual-MP, a dual-process motion-planning architecture that combines a fast neural System-1 policy with a symbolic System-2 fallback. Dual-MP accepts S1 rollouts only when they are collision-free and goal-reaching; otherwise, it delegates planning to either MPC or CBF. Successful S2 recoveries are then reused to improve S1 through continual learning.

Dual-MP with MPC strikes the best balance between success rate and runtime: it achieves the highest success rate while remaining substantially faster than standalone MPC. DMP-CBF is also improved in success rate, but it is slightly slower than CBF because the added S1 rollout overhead outweighs part of the savings. This supports the case for dual-process as it can deliver measurable gains in planning efficiency and performance. At the same time, the results show that continual learning can reduce dependence on System-2. Conversely, naive warm-starting of MPC from rejected S1 trajectories provides no consistent benefit. Overall, Dual-MP offers a promising demonstration that fast-and-slow architecture is a practical mechanism for combining the speed of learned policies with the robustness of optimization-based planning.

References

  • Ames et al. (2019) A. D. Ames, S. Coogan, M. Egerstedt, G. Notomista, K. Sreenath, and P. Tabuada Control barrier functions: theory and applications. In 2019 18th European Control Conference (ECC), Vol. , pp. 3420–3431. External Links: Document Cited by: Introduction, Control Barrier Functions.
  • Ames et al. (2017) A. D. Ames, X. Xu, J. W. Grizzle, and P. Tabuada Control barrier function based quadratic programs for safety critical systems. IEEE Transactions on Automatic Control 62 (8), pp. 3861–3876. External Links: Document Cited by: Classical Motion Planning..
  • Balasubramanian et al. (2012) S. Balasubramanian, A. Melendez-Calderon, and E. Burdet A robust and sensitive metric for quantifying movement smoothness. IEEE Transactions on Biomedical Engineering 59 (8), pp. 2126–2136. External Links: Document Cited by: item (iii).
  • Bjelonic et al. (2022) M. Bjelonic, R. Grandia, M. Geilinger, O. Harley, V. Medeiros, V. Pajovic, E. Jelavic, S. Coros, and M. Hutter Offline motion libraries and online mpc for advanced mobility skills. The International Journal of Robotics Research 41, pp. 903–924. External Links: Document Cited by: Introduction, Learning-Augmented Motion Planning..
  • Carvalho et al. (2024) J. Carvalho, A. T. Le, M. Baierl, D. Koert, and J. Peters Motion planning diffusion: learning and planning of robot motions with diffusion models. External Links: 2308.01557, Link Cited by: Introduction, Learning-Augmented Motion Planning..
  • Fabiano et al. (2025) F. Fabiano, M. B. Ganapini, A. Loreggia, N. Mattei, K. Murugesan, V. Pallagani, F. Rossi, B. Srivastava, and K. B. Venable Thinking fast and slow in human and machine intelligence. Commun. ACM 68 (8), pp. 72–79. External Links: Link, Document Cited by: Introduction, Thinking Fast and Slow in AI.
  • Fishman et al. (2023) A. Fishman, A. Murali, C. Eppner, B. Peele, B. Boots, and D. Fox Motion policy networks. In Proceedings of The 6th Conference on Robot Learning, K. Liu, D. Kulic, and J. Ichnowski (Eds.), Proceedings of Machine Learning Research, Vol. 205, pp. 967–977. External Links: Link Cited by: Classical Motion Planning..
  • Flash and Hogan (1985) T. Flash and N. Hogan The coordination of arm movements: an experimentally confirmed mathematical model. In Journal of Neuroscience, External Links: Link Cited by: item (iii).
  • Fridovich-Keil et al. (2018) D. Fridovich-Keil, S. L. Herbert, J. F. Fisac, S. Deglurkar, and C. J. Tomlin Planning, fast and slow: a framework for adaptive real-time safe trajectory planning. External Links: 1710.04731, Link Cited by: Learning-Augmented Motion Planning..
  • Ganapini et al. (2021) M. B. Ganapini, M. Campbell, F. Fabiano, L. Horesh, J. Lenchner, A. Loreggia, N. Mattei, F. Rossi, B. Srivastava, and K. B. Venable Thinking fast and slow in ai: the role of metacognition. External Links: 2110.01834, Link Cited by: Thinking Fast and Slow in AI.
  • Hegde and Panagou (2016) R. Hegde and D. Panagou Multi-agent motion planning and coordination in polygonal environments using vector fields and model predictive control. In 2016 European Control Conference (ECC), Vol. , pp. 1856–1861. External Links: Document Cited by: Introduction.
  • Ichter et al. (2017) B. Ichter, J. Harrison, and M. Pavone Learning sampling distributions for robot motion planning. CoRR abs/1709.05448. External Links: Link, 1709.05448 Cited by: Introduction, Classical Motion Planning., Neural S1.
  • Johnson et al. (2022) J. J. Johnson, U. S. Kalra, A. Bhatia, L. Li, A. H. Qureshi, and M. C. Yip Motion planning transformers: a motion planning framework for mobile robots. External Links: 2106.02791, Link Cited by: Learning-Augmented Motion Planning..
  • Kahneman (2011) D. Kahneman Thinking, fast and slow. Farrar, Straus and Giroux. Cited by: Introduction, Thinking Fast and Slow in AI.
  • Karaman and Frazzoli (2011) S. Karaman and E. Frazzoli Sampling-based algorithms for optimal motion planning. CoRR abs/1105.1186. External Links: Link, 1105.1186 Cited by: Introduction.
  • Kim et al. (2025) T. Kim, R. W. Beard, and D. Panagou How to adapt control barrier functions? a learning-based approach with applications to a vtol quadplane. In IEEE Conference on Decision and Control (CDC), Cited by: 1st item.
  • LaValle (2006) S. M. LaValle Planning algorithms. Cambridge University Press. External Links: Link, Document, ISBN 9780511546877 Cited by: Introduction.
  • Mansard et al. (2018) N. Mansard, A. DelPrete, M. Geisert, S. Tonneau, and O. Stasse Using a memory of motion to efficiently warm-start a nonlinear predictive controller. In 2018 IEEE International Conference on Robotics and Automation (ICRA), Vol. , pp. 2986–2993. External Links: Document Cited by: Introduction, Learning-Augmented Motion Planning..
  • Mayne et al. (2000) D.Q. Mayne, J.B. Rawlings, C.V. Rao, and P.O.M. Scokaert Constrained model predictive control: stability and optimality. Automatica 36 (6), pp. 789–814. External Links: ISSN 0005-1098, Document, Link Cited by: Introduction, Classical Motion Planning., Nonlinear Model Predictive Control.
  • Paden et al. (2016) B. Paden, M. Cáp, S. Z. Yong, D. S. Yershov, and E. Frazzoli A survey of motion planning and control techniques for self-driving urban vehicles. IEEE Trans. Intell. Veh. 1 (1), pp. 33–55. External Links: Link, Document Cited by: Introduction.
  • Pallagani et al. (2025) V. Pallagani, A. Loreggia, F. Fabiano, B. Srivastava, F. Rossi, and L. Horesh SOFAI lab: a hands-on guide to building neurosymbolic systems with metacognitive control. In AAAI Conference on Artificial Intelligence, Cited by: Introduction, Thinking Fast and Slow in AI, The Metacognitive Module.
  • Ross et al. (2010) S. Ross, G. J. Gordon, and J. A. Bagnell No-regret reductions for imitation learning and structured prediction. CoRR abs/1011.0686. External Links: Link, 1011.0686 Cited by: Continual learning..
  • Sucan et al. (2012) I. A. Sucan, M. Moll, and L. E. Kavraki The open motion planning library. IEEE Robotics and Automation Magazine 19 (4), pp. 72–82. External Links: Document Cited by: item (iii).
  • Verschueren et al. (2021) R. Verschueren, G. Frison, D. Kouzoupis, J. Frey, N. van Duijkeren, A. Zanelli, B. Novoselnik, T. Albin, R. Quirynen, and M. Diehl Acados – a modular open-source framework for fast embedded optimal control. Mathematical Programming Computation. Cited by: 2nd item.
  • Wang et al. (2021) J. Wang, T. Zhang, N. Ma, Z. Li, H. Ma, F. Meng, and M. Q.-H. Meng A survey of learning-based robot motion planning. IET Cyber-Systems and Robotics 3 (4), pp. 302–314. External Links: Document, Link, https://ietresearch.onlinelibrary.wiley.com/doi/pdf/10.1049/csy2.12020 Cited by: Classical Motion Planning..

Appendix A Computational Resources and Experimental Configuration

Computational resources

All experiments were run on an Intel Xeon Gold 6248 CPU server with 40 cores and 125 GiB of memory, running Ubuntu 22.04.5 LTS.

Overall, generating the six nonlinear benchmark suites, training the System 1 policies, collecting DAgger recovery demonstrations, and evaluating all System 1/System 2 configurations required several CPU-hours of computation across the experimental batches. A detailed training-time breakdown is provided in Table 6.

Environment Family Baseline S1 Training DAgger Collection CL S1 Training
Large sparse (LS) ∼\sim6 s 11 s (8–14) ∼\sim66 s
Dense clutter (DC) ∼\sim84 s 12 s (11–13) ∼\sim96 s
Serial walls (SW) ∼\sim18 s 32 s (28–35) ∼\sim24 s
Maze branching (MB) ∼\sim27 s 39 s (37–43) ∼\sim18 s
Long slalom (LSM) ∼\sim90 s 80 s (75–87) ∼\sim102 s
Bugtrap (BT) ∼\sim2 s 23 s (21–24) ∼\sim36 s
Table 6: Approximate training and dagger collecting time per environment family.

Parameters and Hyperparameters

Tables 7 and 8 summarize the model parameters and the training, benchmark, and reproducibility hyperparameters used in our experiments.

Parameter Description Value
Channel dimensions of the three-layer Conv1d context encoder. dctx→→→128d_{\mathrm{ctx}}\!\rightarrow\!64\!\rightarrow\!128\!\rightarrow\!128
Kernel size of the Conv1d context encoder. 3
Pooling operation applied to the encoded context. Adaptive average pooling
Widths of the situation and dynamics encoders, respectively. 128 / 128
Width of the goal encoder. 64
Dimensions of the multilayer perceptron prediction head. →→→2320\!\rightarrow\!256\!\rightarrow\!256\!\rightarrow\!2
Hidden-layer width of the neural System 1 policy. 256
Number of channels in the final convolutional layer. 128
Dropout probability used throughout the policy. 0.05
Table 7: Values and descriptions of the neural System 1 model parameters.
Hyperparameter Description Value
Optimizer used to train the neural System 1 policy. AdamW
Learning rate and weight-decay coefficient. 10−410^{-4} / 10−510^{-5}
Training batch size and validation-data fraction. 64 / 0.10
Maximum gradient norm used for gradient clipping. 5.0
Supervised training objective and its transition parameter. Huber loss, δ=1.0\delta=1.0
Number of continual-learning training epochs. 12
Number of baseline training epochs. 35
Random seeds used for training, evaluation, and probing. 7 / 8 / 700
Continual-learning block size and scenario-ordering seed. 100 / 42
Numbers of evaluation and DAgger workers. 16 / 16
System 1 integration step and nominal rollout horizon. 0.075 s / 900 steps
Situation-grid resolution and corridor buffer. 25×2525\times 25 / 2 cells
Differentiable rollout horizon and application frequency. 8 steps / every 8 batches
Safety-filter mode used during runtime and differentiable rollouts. Policy
Replay fraction allocated to bootstrap demonstrations. 0.60
Number of DAgger states sampled from each scenario. 4
Relative weights assigned to bootstrap and DAgger samples. 1.0 / 1.0
Planning timeout for each scenario. 60 s
Table 8: Key training, benchmark, and reproducibility hyperparameters.

Appendix B Visualising Trajectory Quality

The trajectory quality is described by a score QQ that takes the geometric mean of path efficiency, smoothness, and obstacle-clearance score. A higher score corresponds to a shorter, smoother, and/or better-cleared trajectory. Figure 3 shows the neural System 1 (NN), System 2 CBF, and System 2 MPC trajectories and their corresponding scores on the same scenario.

Refer to caption

(a) S1-NN. Q=0.592Q=0.592.

Refer to caption

(b) S2-CBF. Q=0.802Q=0.802.

Refer to caption

(c) S2-MPC. Q=0.870Q=0.870.

Figure 3: Trajectory-quality comparison on Bugtrap scenario 6. Higher QQ indicates a more efficient, smoother, and better-cleared trajectory.

Appendix C Visualising the Benchmark Families

Figure 4 provides a visual index of the six nonlinear two-dimensional obstacle-avoidance families and the trajectory computed by System 2 MPC. Large Sparse and Dense Clutter vary obstacle density; Serial Walls and Maze Branching introduce structured barriers and route choices; Long Slalom creates a long alternating corridor; and Bugtrap places the robot in a concave trapping geometry.

Refer to caption

(a) Large Sparse (LS).

Refer to caption

(b) Dense Clutter (DC).

Refer to caption

(c) Serial Walls (SW).

Refer to caption

(d) Maze Branching (MB).

Refer to caption

(e) Long Slalom (LSM).

Refer to caption

(f) Bugtrap (BT).

Figure 4: Representative nonlinear motion-planning environments used in the evaluation. Each panel shows an example map from one benchmark family.