SRL-MPC: Shape-Aware Reinforcement Learned Model Predictive Control

TL;DR: SRL-MPC combines shape-aware high-order control barrier functions (HOCBFs) with reinforcement learning for online MPC parameter adaptation, enabling safe and efficient navigation of heterogeneous robot shapes in dense crowds without geometry simplification or policy retraining.

Demonstrations

Trained once at 15 robots · evaluated without retraining · two random episodes per scene · 2× speed

Density Sweep

Training distribution

The policy is trained once with 15 random convex-polygon robots in a 10 m × 10 m arena. The same policy is then run, without retraining, at 10, 15, 20, and 25 robots to show how it scales with crowd density.

Random 1
Random 2
10 robots
Random 1
Random 2
15 robotsTraining density
Random 1
Random 2
20 robots
Random 1
Random 2
25 robots

Held-out Scenes

Out-of-distribution · zero-shot

None of these scenes appear in training. The same 15-robot policy is tested without retraining under six distribution shifts: nonconvex footprints, mixed cross geometry, circular polygon exchange, through traffic with moving obstacles, heterogeneous dynamics, and a dense crowd of circular robots.

Random 1
Random 2
Nonconvex union
Random 1
Random 2
Cross geometry
Random 1
Random 2
Circular polygon exchange
Random 1
Random 2
Through traffic
Random 1
Random 2
Heterogeneous dynamics
Random 1
Random 2
Disc crowd20 circular robots

Abstract

Safe and efficient shape-aware navigation in heterogeneous crowds and robot fleets remains challenging. Traditional approaches often assume homogeneous robots, sparse workspaces, simplified geometry, offline computation, or handcrafted parameters to make the problem tractable, which limits their deployment in dense crowd scenarios. Toward this end, we propose Shape-Aware Reinforcement Learned Model Predictive Control (SRL-MPC), a method for safe, efficient, and adaptive navigation in crowds with heterogeneous shapes without geometry simplification.

SRL-MPC formulates high-order control barrier function residuals from geometric separation features based on support-function geometry. A reinforcement learning policy reads these local features and produces real-time MPC parameter updates, while the executed controls remain the solution of an explicit model-predictive optimization problem. Randomized crowd experiments with arbitrarily shaped robot fleets demonstrate strong effectiveness, scalability, and robustness, especially as crowd density increases.

Architecture

SRL-MPC architecture from crowd geometry through geometric separation features, reinforcement learned parameter adaptation, HOCBF-MPC, and robot navigation

Key Advantages and Features

1. Shape-aware geometry: compact geometric separation features represent circular and convex-polygon bodies without reducing every robot to a disc.

2. Adaptive planning: each learned policy updates path tracking, control effort, and safety-distance parameters from local crowd geometry.

3. Model-based execution: reinforcement learning tunes the planner; the action is still produced by an explicit local MPC problem with bounded controls.

4. Dense-scene robustness: independently trained policies use the same 15-robot training distribution and are evaluated in dense 20- and 25-robot scenes.

SRL-MPC reinforcement learning weight adapter architecture

Multi-seed evaluation · 300 episodes per setting

Dense-scene Evaluation

Three independently trained policies, each evaluated for 100 episodes per dense setting. Values report mean ± sample SD across training seeds.

25-robot success 86.7% ±0.6 percentage points across three policies*
Key comparison +65.7 pp

above SARL at 21%, the strongest external baseline in the 25-robot setting.

N = 20 91.0% ±2.6 pp

20-robot success

N = 25 86.7% ±0.6 pp

25-robot success*

Dense macro 88.8% ±1.5 pp

mean over N = 20 and N = 25

Directly labeled slope chart comparing success rates for SRL-MPC and five baselines as robot count increases from 20 to 25

25-robot baseline comparison

Success rate
MethodSuccess
SRL-MPC86.7 ± 0.6%*
SARL21%
AVOCADO11%
VO-polytope8%
ORCA7%
RL-RVO1%

Rebuttal evaluation · no retraining

Out-of-Distribution Evaluation

Three policies trained independently in the 15-robot random-convex-polygon setting are evaluated without retraining for 100 episodes each in five distribution shifts.

77.5%±2.5 pp · OOD macro across five shifts

Heatmap comparing six methods across five out-of-distribution scenarios using SRL-MPC averages over three training seeds

Preliminary deployment

Real-World Demonstration

A controlled indoor platform comparison of SRL-MPC with ORCA, distributed MPC, and manually tuned MPC.

Side-by-side platform comparisonORCA · Dist-MPC · Manual-MPC · SRL-MPC · 2× playback Preliminary demo