Dual Process Motion Planning
Abstract
Robotic systems are deeply embedded in both industry and everyday life, where they are expected to act with speed, precision, and reliability. Classical control and planning methods have long delivered strong guarantees, but often at the cost of computational efficiency and adaptability. More recently, learning-based approaches have shown promise in overcoming these limitations, enabling agents to leverage experience to accelerate decision-making and address previously intractable problems. In this work, we bridge these two approaches through a neuro-symbolic perspective on nonlinear motion planning. Inspired by the Thinking Fast and Slow paradigm, we introduce a dual-process architecture that combines the strengths of robust reasoning and learning. Our framework integrates state-of-the-art symbolic solvers as a “System-2” component with experience-driven “System-1” modules. A metacognitive controller dynamically orchestrates their interaction, selecting when to rely on fast intuition versus slower, more precise reasoning. By evaluating the framework across diverse nonlinear benchmark environments, we demonstrate that this architecture yields consistent gains in planning efficiency, accuracy, and generalization, while promoting reuse across tasks. The results suggest that tightly coupling learning with structured reasoning offers a scalable path toward more capable and adaptive robotic systems.
1The Chinese University of Hong Kong, Shenzhen
jiayiyan@link.cuhk.edu.cn
2University of Oxford
francesco.fabiano@cs.ox.ac.uk, alessandro.abate@cs.ox.ac.uk
Introduction
Motion planning is a core component of robotic systems and underpins a wide range of applications, including autonomous driving (Paden et al. 2016) and multi-agent coordination (Hegde and Panagou 2016). Its primary objective is to compute a high-quality, collision-free trajectory that connects a start state to a goal state while respecting system dynamics and constraints. However, real-world environments are often high-dimensional, continuous, and dynamically constrained, which makes motion planning computationally challenging, especially under strict real-time requirements (LaValle 2006; Karaman and Frazzoli 2011).
Classical motion planning methods can be broadly categorized into experience-based planning using offline trajectory libraries and neural networks, and online optimization-based planning. Experience-based methods, such as those leveraging neural networks, amortize planning into a learned policy or trajectory generator that maps observations directly to controls or waypoints at runtime, enabling fast online inference and interpolation beyond a finite trajectory library, albeit typically with weaker guarantees (Ichter et al. 2017). Their effectiveness depends critically on selecting a trajectory that matches the current scenario, which is difficult in complex environments, hence often leading to suboptimal or unsafe performance.
In contrast, online optimization-based methods, including Model Predictive Control (MPC) (Mayne et al. 2000) and Control Barrier Function (CBF)-based safety filters (Ames et al. 2019), compute trajectories by solving constrained optimization problems that explicitly account for system dynamics and environmental constraints. While these methods reliably provide solutions, they incur in computational overhead, limiting their applicability in time-critical settings.
To bridge this gap, recent work has explored learning-augmented motion planning, where data-driven models are used to improve efficiency and performance (Mansard et al. 2018; Bjelonic et al. 2022; Carvalho et al. 2024). These methods improve convergence and solution quality, but they typically rely on a fixed pipeline that often hinders their efficiency. To address this limitation, we draw inspiration from a recent AI paradigm Fabiano et al. (2025) that is in turn informed by the dual-system theory Kahneman (2011). We propose Dual-MP, a dual-system motion-planning framework that exploits between fast experience-based and slower online solving.
We instantiate this idea within a SOFAI-style architecture (Pallagani et al. 2025). Dual-MP inherits the modular nature of SOFAI: it uses both fast, experience-based and slow, deliberate solvers. We refer to the former category as System-1 (S1) and to the latter as System-2 (S2). These are arbitrated through a metacognitive (MC) agent. Dual-MP is equipped with a Neural Network S1: a neural policy trained from successful trajectories, and two general S2 online solvers: one based on MPC and one on CBF. The code used in this work is available online11 1 https://github.com/verayannn/System-1-and-System-2-in-Motion-Planning.
We summarize our main contributions below:
- •
We propose Dual-MP, a modular S1/S2 architecture for nonlinear motion planning that arbitrates fast neural planning and deliberate symbolic solving.
- •
We instantiate the framework with a neural S1 policy and two nonlinear S2 solvers, based on MPC and CBFs, under a common MC planner interface.
- •
We add continual learning, where successful trajectories are reused to retrain S1 and improve future fast planning.
- •
We evaluate the framework across diverse nonlinear benchmark families, reporting success, runtime, S1/S2 usage, and trajectory quality.
Related work
Classical Motion Planning.
Current classical motion planning methods are primarily divided into two popular categories: experience-based planning using neural networks and online optimization-based planning.
Experience-based planning typically uses neural networks to amortize planning from previously solved problems, learning a direct mapping from observations, goals, and local state information to actions, waypoints, or full trajectories (Ichter et al. 2017; Fishman et al. 2023). Such methods can provide low-latency inference and strong empirical performance in complex environments by reusing structure learned from expert demonstrations, simulation data, or successful planning experience (Wang et al. 2021). While effective, such approaches typically do not preserve explicit models of dynamics, safety constraints, or solver confidence at inference time, which limits their interpretability and their ability to decide when the output should be trusted.
In contrast to experience-based planning methods, online optimization-based methods, including MPC and CBF-based safety filters, compute trajectories or controls by solving constrained optimization problems at runtime (Mayne et al. 2000; Ames et al. 2017). These methods directly incorporate system dynamics and environmental constraints, yielding high-quality and dynamically feasible solutions. However, they often incur significant computational overhead, which makes real-time deployment challenging in complex or high-dimensional scenarios.
Learning-Augmented Motion Planning.
To mitigate this, recent work seeks to combine the efficiency of offline planning with the accuracy of online optimization through learning-augmented hybrid approaches. A common strategy is to use learned models or offline datasets to warm-start online solvers. For example, Memory-of-Motion learns a mapping from task descriptors to state-control trajectories and uses the resulting memory to initialize nonlinear predictive control (Mansard et al. 2018). Bjelonic et al. (2022) use offline motion libraries as reference costs for online MPC, which allows long-horizon offline behaviors to be executed through short-horizon feedback optimization. Transformer-based motion planners learn to restrict or guide the search space from prior data (Johnson et al. 2022), while diffusion-based planners learn multimodal trajectory priors that can be sampled or adapted during planning (Carvalho et al. 2024). Closely related to our dual-system motivation, Fridovich-Keil et al. (2018) propose a “Planning, Fast and Slow” framework in which offline safety computation enables safe switching among online planners. Their method provides a strong safety-aware planning module; in our terminology, such a framework can be viewed as a possible instantiation of S2.
Despite these advances, a key limitation of existing hybrid approaches is that they usually invoke the online solver regardless of how well the offline or learned prior matches the current scenario. This leads to unnecessary computation in cases where a retrieved or neural trajectory is already sufficient. More fundamentally, these methods lack a principled mechanism for deciding when online optimization is required. Dual-MP is designed to address exactly this allocation problem.
Background
Nonlinear Model Predictive Control
Model Predictive Control (MPC) is a widely used optimal control framework for motion planning under dynamic and environmental constraints (Mayne et al. 2000). At each time step, MPC solves a finite-horizon optimization problem to compute a control sequence that minimizes a cost function while satisfying system dynamics and constraints.
Consider nonlinear dynamics or, after discretization, where is the state and is the control. Given the current state and goal , nonlinear MPC solves subject to The stage cost is typically chosen as where and . Obstacle avoidance is encoded through nonlinear state constraints defining the collision-free set MPC is accurate because it optimizes over future trajectories, but repeatedly solving the resulting nonlinear program can be computationally expensive.
Control Barrier Functions
Control Barrier Functions (CBFs) provide a complementary approach for enforcing safety constraints in control systems (Ames et al. 2019). Unlike MPC, the CBF-QP does not optimize an entire future trajectory; instead, it acts as a safety filter that enforces local obstacle-avoidance constraints.
For a nonlinear control-affine system let the safe set for obstacle be where is positive outside the obstacle and negative inside it. Forward invariance of can be encouraged by enforcing where and are Lie derivatives and controls how aggressively the controller moves away from the safety boundary.
At each timestep, the CBF controller solves the quadratic program subject to Here, is a nominal control input, such as a goal-directed command or a command proposed by the neural System-1 policy.
Thinking Fast and Slow in AI
Kahneman (2011) describe human decision making as the interaction between two complementary processes: a fast, intuitive, experience-driven System-1 and a slower, deliberative, reasoning-based System-2. Recent AI architectures, including SOFAI (Pallagani et al. 2025; Fabiano et al. 2025), adapt this dual-process principle to machine decision-making by combining fast and slow solvers with a metacognitive agent that arbitrates them (Ganapini et al. 2021). These architectures have proven successful in tackling settings closely related to motion planning, such as classical planning and constrained grid navigation, as shown in Fabiano et al. (2025).
Formally, SOFAI defines a decision architecture composed of three components: (i) a set of fast S1 solvers, typically data-driven and experience-based; (ii) a set of slow S2 solvers, based on explicit symbolic reasoning; and (iii) a centralized metacognitive (MC) controller. Incoming problem instances automatically trigger one or more S1 solvers, which produce candidate solutions together with confidence estimates. The MC controller then decides whether to accept the S1 proposal or invoke an S2 solver. This decision is performed in two stages: a lightweight assessment that evaluates whether the expected solution quality satisfies a task-dependent threshold under resource constraints, followed, when necessary, by a cost-benefit comparison between S1 and S2 execution. S2 reasoning is activated only if its expected improvement compensates for the additional computational cost. A key aspect of this framework is that S1 behavior is not static. Through metacognition, solutions produced or validated by S2 can be used to improve the fast solver over time, effectively distilling deliberative reasoning into reactive policies. In this sense, the architecture supports an iterative refinement process in which expensive symbolic reasoning is gradually amortized into efficient inference, enabling the system to adapt to recurring problem distributions while reducing reliance on S2 computation.
Dual-MP
As mentioned above, Dual-MP solves motion planning through metacognitive arbitration. A planning query is defined as where denotes the nonlinear system dynamics, is the obstacle map, and are the start and goal states, is the workspace, and is the admissible control set. In this work, we consider discrete-time nonlinear reach–avoid problems of the form where is the discretized nonlinear dynamics used by the planner.
In the following, we define the main components of Dual-MP, namely the S1 and S2 solvers, as well as the functionalities required from the MC module. In particular, as detailed in Section Results, we consider two basic Dual-MP configurations obtained by combining the S1 solver with two S2 solvers.
Neural S1
Our System-1 is implemented as a neural reactive policy. This follows the common approach of imitation learning to predict low-level controls from local environment observations and goal information (Ichter et al. 2017).
Given the current rollout context, local obstacle information, nonlinear dynamics features, and the goal direction, the policy predicts a control input directly: where is a fixed-length window of recent states in a local coordinate frame, encodes the local obstacle situation, encodes the nonlinear dynamics at the current state, and is the local goal vector.
The policy uses a lightweight convolutional architecture. A one-dimensional convolutional encoder processes the recent state context , while separate multilayer encoders process the situation vector, dynamics features, and goal vector. The resulting features are concatenated and passed through a feedforward control head:
This design keeps S1 inexpensive at runtime while still conditioning the control prediction on both geometry and nonlinear dynamics.
As mentioned, S1 is trained by imitation from successful System-2 trajectories. In particular, the bootstrap demonstrations are solved offline by System-2, and the System-2 recoveries are collected online during continual learning. For each state-control pair, the supervised target is the expert control , with auxiliary single-step losses that encourage accurate one-step motion, correct direction, and progress toward the goal, and a multi-step rollout term that penalizes the drift accumulated when the policy is run in closed loop:
Here, is a Huber control-imitation loss, penalizes one-step prediction error under the nonlinear dynamics, encourages alignment with the expert displacement, matches step magnitude, and penalizes insufficient progress toward the goal. The remaining three terms are evaluated on a differentiable -step rollout of the policy under the nonlinear dynamics, run once every supervised batches (): penalizes deviation of the rolled-out positions from the expert positions, penalizes large control changes, and penalizes proximity to obstacles along the rollout.
At execution time, S1 rolls out the learned policy under the nonlinear dynamics. A lightweight one-step safety filter selects among the predicted control, goal-directed controls, scaled controls, and zero control, keeping only candidates whose next segment is collision-free. The resulting trajectory is then verified for collision freedom and goal reach. In the SOFAI setting, S2 is invoked only when the S1 rollout fails this verification.
S2 Solvers: CBF and MPC
Dual-MP can be instantiated with two state-of-the-art S2 solvers:
- •
safe_control: a control barrier function (CBF)-based safety filter (Kim et al. 2025)
- •
acados: a nonlinear model predictive control (MPC) planner (Verschueren et al. 2021)
Let us note that both solvers are implemented using their standard formulations from the literature.
The Metacognitive Module
Dual-MP follows the SOFAI principle of coordinating fast S1 solvers and slower deliberative S2 solvers through a metacognitive arbitration module (Pallagani et al. 2025). Depending on the Dual-MP configuration, the S1 and S2 components are instantiated differently; in our experiments, the two configurations are obtained by pairing the S1 solver with two S2 solvers introduced above.
For each new scenario, the metacognitive module first invokes the configured S1 solver. The proposed rollout is accepted only if it passes verification (collision freedom and goal reach - cf. above). If the S1 proposal is rejected, Dual-MP invokes the configured S2 fallback, provided that the available time budget is sufficient relative to the estimated problem difficulty. If S2 does not return a satisfactory solution within the budget, Dual-MP falls back to the best available S1 proposal. Successful S2 rollouts are stored in memory and periodically used to improve S1.
The metacognitive module relies on three quantities to properly arbitrate the two systems: problem difficulty, solution correctness, and solver confidence. Difficulty is a scenario-level quantity used to characterize the planning instance; correctness measures the quality of a proposed rollout; and confidence estimates whether an S1 proposal is worth verifying. Following the SOFAI principle, S2 solvers are treated as deliberative solvers and assigned confidence , while S1 confidence is computed as described above.
Correctness.
Given a rollout , correctness is set to if the trajectory reaches the goal and remains collision-free. Otherwise, we assign a soft correctness score based on path length, terminal goal error, and collision penalty: where Here, is a bounded collision penalty and controls the relative cost of unsafe states. In the experiments reported in this paper, we use a strict correctness threshold of , requiring both goal reaching and collision avoidance. More permissive thresholds are supported by the SOFAI framework and may be useful under strict time constraints.
Difficulty.
We define the difficulty of a query scenario using simple features to maintain a lightweight process. Let denote the set of rectangular obstacles and let denote the workspace. We first compute the obstacle occupancy Let be the midpoint between start and goal, and let be a clearance estimate from this midpoint to the nearest obstacle. The difficulty is then
where is the planning horizon, is the goal tolerance, and is the collision margin. The first factor captures geometric clutter and clearance, while the remaining factors account for horizon length, goal precision, and safety strictness.
SOFAI Continual Learning
A key advantage of SOFAI-inspired architectures is that S1 solvers can improve over time. Dual-MP supports this by storing successfully generated S2 trajectories and using them to enrich the corresponding S1 solver. Specifically, when S2 successfully solves a query, the resulting trajectory is added to the online dataset. For Neural S1, accumulated S2 trajectories are distilled periodically after a fixed number of episodes by fine-tuning the policy on both the original offline dataset and the newly collected S2 successes.
Experimental Evaluation
All experiments were run on an Intel Xeon Gold 6248 CPU server with 40 cores and 125 GiB of memory, running Ubuntu 22.04.5 LTS. The source code, benchmark data, and additional details, including offline training times, will be provided in the appendix for space reasons.
We evaluate Dual-MP on reach–avoid motion-planning tasks in two-dimensional continuous environments. Each instance is defined by nonlinear dynamics a start state , a goal state , rectangular obstacles , control bounds , and workspace bounds . A rollout counts as successful only if it is both collision-free and goal-reaching; no partial credit is given. Our experiments address four questions:
- •
Q1: Can Dual-MP improve state-of-the-art planning?
- •
Q2: Does Dual-MP preserve reliability?
- •
Q3: How does continual learning affect S1?
- •
Q4 Does warm-starting System-2 from the rejected S1 trajectory help System-2 solving?
We evaluate our approach on six benchmark families, each stressing a different aspect of reach–avoid planning:
- •
Large sparse (LS): a few large obstacles; tests long-horizon obstacle avoidance in an open workspace.
- •
Dense clutter (DC): many small rectangles; tests local collision avoidance under frequent re-planning pressure.
- •
Serial walls (SW): several wall-with-gap structures in sequence; tests repeated constrained passages, where a single missed gap is unrecoverable.
- •
Maze branching (MB): intersecting wall structures with multiple candidate routes; tests global route commitment.
- •
Long slalom (LSM): a long corridor with alternating offset barriers; tests sustained tracking over a horizon far longer than the MPC prediction window.
- •
Bugtrap (BT): concave trap layouts; tests robustness to locally attractive but globally poor routes.
Instances within a family share the workspace scale, the start/goal pair, and the reach–avoid objective, and vary in obstacle layout and in the parameters of the nonlinear drift field. For every family, we generate three disjoint instance sets from separate seeds: a training set of instances used only to bootstrap S1, an evaluation set of instances used for the results in Tables 2 and 3, and a held-out probe set of instances used exclusively to measure the effect of continual learning. The probe set is never trained on and is identical across all blocks and all methods, so probe curves are directly comparable over time.
All of the generated instances are run on six different solvers/Dual-MP configurations, listed in Table 1. The three Dual-MP variants differ only in the System-2 they escalate to and in whether that solver is warm-started from the rejected S1 trajectory.
| Name | System-1 | System-2 |
|---|---|---|
| NN | Neural policy | – |
| CBF | – | CBF |
| MPC | – | MPC |
| DMP-CBF | Neural policy | CBF |
| DMP-MPC | Neural policy | MPC |
| DMP-Warm | Neural policy | MPC (S1 warm start) |
The various approaches are then evaluated on three main metrics: (i) Success rate, which is the fraction of instances whose returned rollout is collision-free and reaches the goal. (ii) Runtimewhich represents the time to find a solution: for Dual-MP it is the sum of the S1 call and, when escalation occurs, the S2 call, so escalated instances are charged for both. We report the mean and the th percentile. (iii) Trajectory Qualityevaluated only on successful rollouts, so that it never trades off against the success rate. We use a duration-invariant index that is the geometric mean of three sub-scores, each in and each normalised against a scenario-defined rather than a solver-defined reference: is the standard path-optimality ratio of the executed length against the shortest collision-free path , obtained by grid search around the obstacles (Sucan et al. 2012). is the spectral arc length of the speed profile expressed relative to an ideal minimum-jerk movement (Flash and Hogan 1985), so a minimum-jerk trajectory scores ; spectral arc length is used because it is the only smoothness measure that is simultaneously valid and reliable (Balasubramanian et al. 2012). scores the worst obstacle clearance against one body radius, capped so that excess conservatism earns nothing. The geometric mean prevents a near-collision or a chattering control from being compensated elsewhere. Two properties matter for the comparisons below. First, is invariant to how fast the trajectory is traversed, so an aggressive controller cannot buy quality by saturating its actuator. Second, because the reference is the scenario’s own shortest path, is comparable across families of very different difficulty.
For Dual-MP we additionally report how many successes were produced by S1 alone versus by the S2 fallback.
Training Protocol
Bootstrap.
For each family and each System-2 solver, we solve the training instances with that solver and retain only the collision-free, goal-reaching trajectories. The base S1 policy is trained on those demonstrations by behaviour cloning augmented with a short differentiable rollout loss (horizon , applied every supervised batches). Crucially, the bootstrap teacher is matched to the System-2 the arm will escalate to: DMP-CBF starts from a CBF-taught policy and DMP-MPC from an MPC-taught one.
Continual learning.
The evaluation instances are shuffled once (fixed seed) and presented in five sequential blocks of . After each block, S1 is retrained, warm-started from the previous block’s weights, on the union of the bootstrap demonstrations and all System-2 recoveries collected so far. Recoveries are obtained in DAgger fashion (Ross et al. 2010): on instances where S1 was rejected, we relabel states visited by the S1 rollout with the System-2 solver, so the policy receives corrective actions on its own state distribution rather than only on the expert’s. A replay fraction of keeps the fixed bootstrap demonstrations in the mixture and limits drift toward the increasingly hard, failure-biased recovery set. The retrained model is evaluated on the frozen probe set before the next block begins.
Results
We now present the experimental results of our paper and use these to answer the four research questions we introduced above. The main results are presented in Table 2, which reports results macro-averaged over the six families for the six configurations, and in Table 3 that breaks the same runs down per family.
| Method | Succ. (%) | Mean RT (ms) | P90 RT (ms) | Mean | P90 | S1/S2 Solves |
|---|---|---|---|---|---|---|
| NN | 36.4 | 230 | 305 | 0.601 | 0.817 | 182 / 0 |
| CBF | 82.0 | 657 | 1448 | 0.795 | 0.844 | 0 / 475 |
| MPC | 84.8 | 5708 | 7092 | 0.844 | 0.875 | 0 / 496 |
| DMP-CBF | 84.1 | 983 | 2279 | 0.734 | 0.850 | 204 / 217 |
| DMP-MPC | 85.1 | 4979 | 8086 | 0.780 | 0.877 | 141 / 285 |
| DMP-Warm | 83.6 | 5008 | 8161 | 0.782 | 0.877 | 140 / 278 |
| Method | LS | DC | SW | MB | LSM | BT |
|---|---|---|---|---|---|---|
| Success rate (%) | ||||||
| NN | 82.8 | 74.4 | 12.0 | 19.2 | 13.0 | 17.0 |
| CBF | 96.0 | 94.0 | 59.7 | 65.0 | 98.2 | 78.8 |
| MPC | 98.2 | 94.0 | 77.5 | 69.5 | 76.2 | 93.2 |
| DMP-CBF | 97.0 | 94.8 | 65.6 | 70.0 | 98.2 | 79.2 |
| DMP-MPC | 98.2 | 95.2 | 77.8 | 69.6 | 76.6 | 93.4 |
| DMP-Warm | 98.2 | 95.2 | 77.2 | 67.8 | 74.6 | 88.8 |
| Mean planning runtime (ms) | ||||||
| NN | 112 | 135 | 230 | 255 | 460 | 185 |
| CBF | 121 | 234 | 934 | 961 | 1258 | 437 |
| MPC | 3439 | 3688 | 4886 | 6140 | 12316 | 3781 |
| DMP-CBF | 349 | 449 | 1131 | 1212 | 1974 | 783 |
| DMP-MPC | 1473 | 1670 | 4866 | 5988 | 12490 | 3385 |
| DMP-Warm | 1440 | 1715 | 4909 | 6048 | 12458 | 3476 |
| Mean trajectory quality | ||||||
| NN | 0.781 | 0.750 | 0.580 | 0.484 | 0.425 | 0.586 |
| CBF | 0.817 | 0.817 | 0.778 | 0.772 | 0.801 | 0.787 |
| MPC | 0.860 | 0.861 | 0.836 | 0.826 | 0.827 | 0.853 |
| DMP-CBF | 0.794 | 0.763 | 0.696 | 0.679 | 0.741 | 0.729 |
| DMP-MPC | 0.761 | 0.716 | 0.793 | 0.802 | 0.817 | 0.792 |
| DMP-Warm | 0.762 | 0.709 | 0.814 | 0.790 | 0.817 | 0.796 |
(Q1) Can Dual-MP improve state-of-the-art planning?
Yes. Dual-MP improves the MPC baseline by filtering out easy instances with S1 and invoking MPC only when needed, which reduces the number of expensive optimisation calls while improving overall success. This yields the strongest success-runtime trade-off among the compared methods. The largest gains occur in large sparse and dense clutter environments, where S1 frequently solves instances without escalation. Gains are limited in serial walls, maze branching, and long slalom, where most instances still require MPC. By contrast, DMP-CBF is slower than CBF alone because the cost of an S1 rollout outweighs the savings from avoiding an already inexpensive fallback. Thus, the value of dual-process arbitration is determined primarily by the runtime gap between S1 and S2, rather than by S1 accuracy alone.
(Q2) Does Dual-MP preserve reliability?
Yes. Across the completed families, both Dual-MP variants match or improve the success rate of their respective S2 fallback (Tables 2 and 3). This follows from the verification gate: S1 is accepted only when its rollout is collision-free and reaches the goal, while rejected rollouts are delegated to S2. S1 can also solve some instances missed by the corresponding fallback solver, particularly in constrained environments.
The main trade-off is trajectory quality. MPC produces the highest-quality trajectories because it explicitly optimises path efficiency and obstacle clearance, whereas accepted S1 rollouts prioritise fast, feasible completion. Consequently, Dual-MP may return trajectories that are longer or pass closer to obstacles than solutions produced by MPC alone. Dual-MP is therefore most suitable when reducing online planning time is more important than retaining the trajectory quality of every MPC solution.
(Q3) How does continual learning affect S1?
Table 4 reports the cumulative changes in S1 success rate and trajectory quality, while Figures 1 and 2 show their evolution across training blocks.
Continual learning improves S1, but the type of improvement depends on the S2 teacher. The CBF-taught policy achieves substantial gains in success rate as fallback experience accumulates, increasing the fraction of instances solved directly by S1 and reducing reliance on CBF. The MPC-taught policy primarily improves trajectory quality and increases success in relatively open environments, such as large sparse and dense clutter, but provides limited or inconsistent success gains in more constrained families.
This difference is in the “teacher” capabilities: MPC demonstrations often contain long-horizon, globally route-dependent plans that a reactive policy with local observations cannot reliably reproduce after small closed-loop deviations. CBF, by contrast, provides short, local corrective actions that are easier for S1 to imitate.
| Family | Arm | Base | B0 | B1 | B2 | B3 | B4 | ||
| LS | CBF | 80.4 | 80.8 | 86.0 | 85.0 | 90.0 | 90.0 | ||
| MPC | 54.2 | 64.8 | 71.8 | 73.8 | 75.2 | 72.4 | |||
| DC | CBF | 72.8 | 74.6 | 82.0 | 84.2 | 85.4 | 85.2 | ||
| MPC | 58.2 | 65.6 | 64.6 | 66.8 | 68.8 | 70.8 | |||
| SW | CBF | 11.2 | 15.8 | 20.8 | 24.6 | 27.6 | 30.4 | ||
| MPC | 12.6 | 9.2 | 5.2 | 5.0 | 6.2 | 4.4 | |||
| MB | CBF | 21.2 | 19.4 | 24.0 | 26.8 | 23.8 | 24.2 | ||
| MPC | 7.0 | 6.2 | 4.0 | 5.2 | 5.6 | 4.2 | |||
| LSM | CBF | 13.0 | 16.8 | 16.6 | 18.8 | 19.0 | 24.8 | ||
| MPC | 0.6 | 0.8 | 0.6 | 2.0 | 2.0 | 2.6 | |||
| BT | CBF | 19.2 | 23.0 | 28.8 | 30.6 | 36.4 | 42.4 | ||
| MPC | 22.2 | 19.4 | 23.6 | 25.0 | 24.6 | 27.2 | |||
| Mean | CBF | 36.3 | 38.4 | 43.0 | 45.0 | 47.0 | 49.5 | ||
| MPC | 25.8 | 27.7 | 28.3 | 29.6 | 30.4 | 30.3 |
(Q4) Does warm-starting System-2 from the rejected S1 trajectory help System-2 solving?
Naive warm-starting from a rejected S1 rollout is not useful as shown in Table 5. Initialising MPC with the rejected trajectory does not give a consistent runtime benefit, and it can even steer the optimiser toward the same locally poor region that caused S1 to fail. Although warm-starting occasionally helps, the effect is unreliable and can reduce fallback success on the hardest instances. We therefore do not use rejected S1 rollouts as unconditional warm starts. While Dual-MP does compute an S1 confidence score, we did not tune it for warm-start selection because it was not a sufficiently reliable proxy for feasibility, so we omitted that mechanism. We leave this as future work.
| Family | Cold (s) | Warm (s) | S2 solves | ||
|---|---|---|---|---|---|
| LS | 137 | 3.489 | 3.492 | 128 / 128 | |
| DC | 150 | 3.764 | 3.771 | 126 / 126 | |
| SW | 449 | 4.813 | 4.803 | 339 / 337 | |
| MB | 461 | 5.935 | 6.067 | 310 / 300 | |
| LSM | 477 | 11.096 | 11.057 | 370 / 362 | |
| BT | 364 | 3.870 | 4.001 | 331 / 311 |
Discussion
Although Dual-MP is only an initial step toward dual-process motion planning, it delivers promising results; it outperforms the standalone baselines in the reported success and efficiency metrics. This suggests that verified arbitration is not just a conceptual framework, but a practical way to combine learned policies with optimization-based solvers.
The results show that Dual-MP is most effective when the two systems have distinct costs and complementary strengths. A fast S1 policy is most valuable when it avoids expensive MPC calls, but it offers little benefit when the fallback is already cheap, as with CBF. This means that success rate alone is not sufficient to evaluate a hybrid planner: the runtime balance between S1 and S2 largely determines whether arbitration is worthwhile. Exact rollout verification is central to this behavior. It allows S1 to produce low-latency feasible solutions without letting unsafe or incomplete predictions replace the fallback solver. When S1 fails, S2 takes over, and successful recoveries can then be reused for continual learning. This makes the architecture practical without requiring the neural policy to be perfectly safe in isolation, as long as it is paired with a correct fallback and a verification gate.
Similarly, we note that continual learning is most effective when the teacher’s behavior is compatible with the student’s policy class. CBF provides short, local corrections that improve S1 coverage, whereas MPC often produces longer, route-dependent behaviors that are harder for a reactive policy to reproduce reliably in closed loop. This suggests several promising directions for future work: (i) training S1 on MPC trajectories and pairing it with CBF, (ii) exploring other combinations of fast and slow components, and (iii) using a portfolio metacontroller to select among multiple S2 planners rather than relying on a single fallback. More generally, future work could optimize different objectives depending on the application, including planning time, trajectory quality, success rate, or even accepting slightly suboptimal solutions when that yields a better overall trade-off. In the same spirit, better knowledge transfer across environment families could reduce bootstrap cost, either by sharing an S1 policy across tasks or by making continual learning more selective when offline retraining budget is limited. Rejected S1 rollouts should likewise be used more selectively as MPC initializations; a stronger confidence or feasibility estimate could help determine when a proposal is close enough to serve as a useful warm start and when MPC should instead solve the problem from scratch.
Overall, this work provides an initial but promising step toward optimizing not only the individual tools, but also their composition. It shows that the interaction between learned policies, symbolic solvers, verification, and retraining can itself be designed and improved.
Conclusion
We presented Dual-MP, a dual-process motion-planning architecture that combines a fast neural System-1 policy with a symbolic System-2 fallback. Dual-MP accepts S1 rollouts only when they are collision-free and goal-reaching; otherwise, it delegates planning to either MPC or CBF. Successful S2 recoveries are then reused to improve S1 through continual learning.
Dual-MP with MPC strikes the best balance between success rate and runtime: it achieves the highest success rate while remaining substantially faster than standalone MPC. DMP-CBF is also improved in success rate, but it is slightly slower than CBF because the added S1 rollout overhead outweighs part of the savings. This supports the case for dual-process as it can deliver measurable gains in planning efficiency and performance. At the same time, the results show that continual learning can reduce dependence on System-2. Conversely, naive warm-starting of MPC from rejected S1 trajectories provides no consistent benefit. Overall, Dual-MP offers a promising demonstration that fast-and-slow architecture is a practical mechanism for combining the speed of learned policies with the robustness of optimization-based planning.
References
- Control barrier functions: theory and applications. In 2019 18th European Control Conference (ECC), Vol. , pp. 3420–3431. External Links: Document Cited by: Introduction, Control Barrier Functions.
- Control barrier function based quadratic programs for safety critical systems. IEEE Transactions on Automatic Control 62 (8), pp. 3861–3876. External Links: Document Cited by: Classical Motion Planning..
- A robust and sensitive metric for quantifying movement smoothness. IEEE Transactions on Biomedical Engineering 59 (8), pp. 2126–2136. External Links: Document Cited by: item (iii).
- Offline motion libraries and online mpc for advanced mobility skills. The International Journal of Robotics Research 41, pp. 903–924. External Links: Document Cited by: Introduction, Learning-Augmented Motion Planning..
- Motion planning diffusion: learning and planning of robot motions with diffusion models. External Links: 2308.01557, Link Cited by: Introduction, Learning-Augmented Motion Planning..
- Thinking fast and slow in human and machine intelligence. Commun. ACM 68 (8), pp. 72–79. External Links: Link, Document Cited by: Introduction, Thinking Fast and Slow in AI.
- Motion policy networks. In Proceedings of The 6th Conference on Robot Learning, K. Liu, D. Kulic, and J. Ichnowski (Eds.), Proceedings of Machine Learning Research, Vol. 205, pp. 967–977. External Links: Link Cited by: Classical Motion Planning..
- The coordination of arm movements: an experimentally confirmed mathematical model. In Journal of Neuroscience, External Links: Link Cited by: item (iii).
- Planning, fast and slow: a framework for adaptive real-time safe trajectory planning. External Links: 1710.04731, Link Cited by: Learning-Augmented Motion Planning..
- Thinking fast and slow in ai: the role of metacognition. External Links: 2110.01834, Link Cited by: Thinking Fast and Slow in AI.
- Multi-agent motion planning and coordination in polygonal environments using vector fields and model predictive control. In 2016 European Control Conference (ECC), Vol. , pp. 1856–1861. External Links: Document Cited by: Introduction.
- Learning sampling distributions for robot motion planning. CoRR abs/1709.05448. External Links: Link, 1709.05448 Cited by: Introduction, Classical Motion Planning., Neural S1.
- Motion planning transformers: a motion planning framework for mobile robots. External Links: 2106.02791, Link Cited by: Learning-Augmented Motion Planning..
- Thinking, fast and slow. Farrar, Straus and Giroux. Cited by: Introduction, Thinking Fast and Slow in AI.
- Sampling-based algorithms for optimal motion planning. CoRR abs/1105.1186. External Links: Link, 1105.1186 Cited by: Introduction.
- How to adapt control barrier functions? a learning-based approach with applications to a vtol quadplane. In IEEE Conference on Decision and Control (CDC), Cited by: 1st item.
- Planning algorithms. Cambridge University Press. External Links: Link, Document, ISBN 9780511546877 Cited by: Introduction.
- Using a memory of motion to efficiently warm-start a nonlinear predictive controller. In 2018 IEEE International Conference on Robotics and Automation (ICRA), Vol. , pp. 2986–2993. External Links: Document Cited by: Introduction, Learning-Augmented Motion Planning..
- Constrained model predictive control: stability and optimality. Automatica 36 (6), pp. 789–814. External Links: ISSN 0005-1098, Document, Link Cited by: Introduction, Classical Motion Planning., Nonlinear Model Predictive Control.
- A survey of motion planning and control techniques for self-driving urban vehicles. IEEE Trans. Intell. Veh. 1 (1), pp. 33–55. External Links: Link, Document Cited by: Introduction.
- SOFAI lab: a hands-on guide to building neurosymbolic systems with metacognitive control. In AAAI Conference on Artificial Intelligence, Cited by: Introduction, Thinking Fast and Slow in AI, The Metacognitive Module.
- No-regret reductions for imitation learning and structured prediction. CoRR abs/1011.0686. External Links: Link, 1011.0686 Cited by: Continual learning..
- The open motion planning library. IEEE Robotics and Automation Magazine 19 (4), pp. 72–82. External Links: Document Cited by: item (iii).
- Acados – a modular open-source framework for fast embedded optimal control. Mathematical Programming Computation. Cited by: 2nd item.
- A survey of learning-based robot motion planning. IET Cyber-Systems and Robotics 3 (4), pp. 302–314. External Links: Document, Link, https://ietresearch.onlinelibrary.wiley.com/doi/pdf/10.1049/csy2.12020 Cited by: Classical Motion Planning..
Appendix A Computational Resources and Experimental Configuration
Computational resources
All experiments were run on an Intel Xeon Gold 6248 CPU server with 40 cores and 125 GiB of memory, running Ubuntu 22.04.5 LTS.
Overall, generating the six nonlinear benchmark suites, training the System 1 policies, collecting DAgger recovery demonstrations, and evaluating all System 1/System 2 configurations required several CPU-hours of computation across the experimental batches. A detailed training-time breakdown is provided in Table 6.
| Environment Family | Baseline S1 Training | DAgger Collection | CL S1 Training |
|---|---|---|---|
| Large sparse (LS) | 6 s | 11 s (8–14) | 66 s |
| Dense clutter (DC) | 84 s | 12 s (11–13) | 96 s |
| Serial walls (SW) | 18 s | 32 s (28–35) | 24 s |
| Maze branching (MB) | 27 s | 39 s (37–43) | 18 s |
| Long slalom (LSM) | 90 s | 80 s (75–87) | 102 s |
| Bugtrap (BT) | 2 s | 23 s (21–24) | 36 s |
Parameters and Hyperparameters
Tables 7 and 8 summarize the model parameters and the training, benchmark, and reproducibility hyperparameters used in our experiments.
| Parameter Description | Value |
|---|---|
| Channel dimensions of the three-layer Conv1d context encoder. | |
| Kernel size of the Conv1d context encoder. | 3 |
| Pooling operation applied to the encoded context. | Adaptive average pooling |
| Widths of the situation and dynamics encoders, respectively. | 128 / 128 |
| Width of the goal encoder. | 64 |
| Dimensions of the multilayer perceptron prediction head. | |
| Hidden-layer width of the neural System 1 policy. | 256 |
| Number of channels in the final convolutional layer. | 128 |
| Dropout probability used throughout the policy. | 0.05 |
| Hyperparameter Description | Value |
|---|---|
| Optimizer used to train the neural System 1 policy. | AdamW |
| Learning rate and weight-decay coefficient. | / |
| Training batch size and validation-data fraction. | 64 / 0.10 |
| Maximum gradient norm used for gradient clipping. | 5.0 |
| Supervised training objective and its transition parameter. | Huber loss, |
| Number of continual-learning training epochs. | 12 |
| Number of baseline training epochs. | 35 |
| Random seeds used for training, evaluation, and probing. | 7 / 8 / 700 |
| Continual-learning block size and scenario-ordering seed. | 100 / 42 |
| Numbers of evaluation and DAgger workers. | 16 / 16 |
| System 1 integration step and nominal rollout horizon. | 0.075 s / 900 steps |
| Situation-grid resolution and corridor buffer. | / 2 cells |
| Differentiable rollout horizon and application frequency. | 8 steps / every 8 batches |
| Safety-filter mode used during runtime and differentiable rollouts. | Policy |
| Replay fraction allocated to bootstrap demonstrations. | 0.60 |
| Number of DAgger states sampled from each scenario. | 4 |
| Relative weights assigned to bootstrap and DAgger samples. | 1.0 / 1.0 |
| Planning timeout for each scenario. | 60 s |
Appendix B Visualising Trajectory Quality
The trajectory quality is described by a score that takes the geometric mean of path efficiency, smoothness, and obstacle-clearance score. A higher score corresponds to a shorter, smoother, and/or better-cleared trajectory. Figure 3 shows the neural System 1 (NN), System 2 CBF, and System 2 MPC trajectories and their corresponding scores on the same scenario.
(a) S1-NN. .
(b) S2-CBF. .
(c) S2-MPC. .
Appendix C Visualising the Benchmark Families
Figure 4 provides a visual index of the six nonlinear two-dimensional obstacle-avoidance families and the trajectory computed by System 2 MPC. Large Sparse and Dense Clutter vary obstacle density; Serial Walls and Maze Branching introduce structured barriers and route choices; Long Slalom creates a long alternating corridor; and Bugtrap places the robot in a concave trapping geometry.
(a) Large Sparse (LS).
(b) Dense Clutter (DC).
(c) Serial Walls (SW).
(d) Maze Branching (MB).
(e) Long Slalom (LSM).
(f) Bugtrap (BT).