SCOPE

Score Curvature for
Online Precision Estimation

Control-Ready Uncertainty for Trajectory Diffusion

CoRL 2026 Spotlight

Zhiwei Xue1,*,† Jia Yue Kam1,* Jinhang Qiu1,* Yifeng Cheng1 Ege Gursoy3 Jiaming Wang1 Vincent Bonnet3 Harold Soh1,2,†

1 National University of Singapore 2 Smart Systems Institute, NUS 3 LAAS-CNRS

* Equal contribution † Corresponding authors

Real-robot manipulation with SCOPE.

Abstract

Diffusion models can represent complex, multimodal trajectory distributions, but extracting uncertainty from them typically requires costly Monte Carlo sampling. This limits their use in real-time control, where robots must rapidly assess risk and maintain safety margins. We introduce Score-Curvature for Online Precision Estimation (SCOPE), a lightweight module that augments diffusion trajectory models with control-ready uncertainty. SCOPE learns a structured precision matrix around each nominal trajectory by distilling score-curvature information and producing calibrated Gaussian tubes with low overhead and without repeated Monte Carlo sampling. These tubes provide per-timestep covariance estimates that can be used both as predicted occupancy for moving agents and as adaptive exploration guides for robot control. We evaluate SCOPE with mode-conditioned multimodal diffusion backbones in pedestrian forecasting, crowd navigation, Maze2D control, and real-world Franka Panda manipulation. Across these settings, SCOPE provides fast uncertainty estimation, which leads to better closed-loop performance.

From Score Curvature to Gaussian Tubes

SCOPE turns diffusion trajectories into Gaussian uncertainty tubes for real-time control.

Pipeline from mode-conditioned diffusion trajectories through the score-curvature precision head to Gaussian uncertainty tubes.
Pipeline. Multimodal diffusion predicts nominal trajectories. The score-curvature head adds structured precision, yielding calibrated per-timestep uncertainty.

SCOPE extracts aleatoric trajectory uncertainty from diffusion models for two uses:

  1. Policy prior: adapt MPPI exploration to the robot's motion flexibility.
  2. World model: predict other agents' occupancy for collision-risk evaluation.

Maze2D Navigation

A diffusion plan guides MPPI through unseen mazes. SCOPE widens exploration in open regions and narrows it near walls.

Vanilla MPPI

Fixed exploration

Log-MPPI

Log-normal exploration

Diff-MPPI

Diffusion guide + fixed exploration

SCOPE

Same diffusion guide + adaptive exploration

Rollouts Executed path 95% covariance tube High-weight rollout Colliding rollouts MPPI plan Collision path / contact
Maze2D success rate and mean path length across rollout budgets. SCOPE achieves the highest success rate and shortest mean path length at every evaluated budget.
Rollout efficiency. SCOPE achieves the highest success rate and shortest average path length across all evaluated rollout budgets. Adaptive exploration helps MPPI escape local traps while limiting collisions near walls.

Y-Corridor Pedestrian Navigation

MID-Ensemble-64

latency-aware

SCOPE

synchronous

Prediction latency delays the robot’s response to P1, leaving too little time to avoid collision.

Success rates for 3, 6, 9, and 12 pedestrians, respectively. No future prediction: 80%, 72%, 70%, 65%. LKF: 96%, 91%, 90%, 85%. MID-UM (no tubes): 100%, 89%, 85%, 86%. MID-MM (no tubes): 100%, 91%, 85%, 88%. MID-Ens-64: 98%, 87%, 82%, 89%. MID-Ens-128: 99%, 90%, 82%, 86%. MID-Ens-256: 99%, 90%, 83%, 87%. SCOPE, no calibration: 100%, 94%, 86%, 88%. SCOPE (full): 100%, 96%, 93%, 92%.
Success rate ↑. 100 episodes per method and density. Success means reaching the goal without collision or timeout. UM / MM: unimodal / multimodal.

Franka Panda Shelf Manipulation

A Franka Panda transfers a tomato to the upper shelf, with static obstacles or a moving human hand. Each method runs 30 trials per condition.

Shelf manipulation setup showing the tomato as the robot target, the Oreo box as the human target, and the surrounding static obstacles.
Task setup. Tomato: robot target. Oreo box: human target in dynamic trials.
A: wider front approach. B: tighter side approach. C: two RGB webcams used to track the human hand.
A: wider front approach. B: tighter side approach. C: two-camera hand tracking.
Static obstacles, 30 trials per method. Success / collision / timeout counts: Diffuser 14 / 16 / 0. Vanilla MPPI 5 / 24 / 1. Diff-MPPI 15 / 7 / 8. SCOPE 25 / 4 / 1. SCOPE success rate: 83%.
Dynamic interaction, 30 trials per method. Success / any collision / timeout counts: SCOPE without world model 11 / 19 / 0. MID-Ens-64 16 / 14 / 0. MID-Ens-128 14 / 16 / 0. SCOPE 21 / 7 / 2. Environment / human collision counts: 8 / 12, 8 / 8, 9 / 10, 4 / 3. Collision types may overlap. SCOPE success rate: 70%.

30 trials per method and condition. Counts derived from the paper’s rounded rates. WM: world model.

Static obstacle avoidance

DiffusionBox collision
Diff-MPPIBox collision
SCOPEObstacle avoidance
DiffusionShelf collision
Diff-MPPIStuck side grasp
SCOPETighter-layout success

Dynamic human interaction

No world modelTop-shelf collision
No world modelHand collision
SCOPEEarly yielding

Limitations

SCOPE distills uncertainty from a learned diffusion score, so its reliability depends on the backbone. Limited training coverage or poorly separated modes can lead to miscalibrated uncertainty tubes. In real-robot trials, inaccurate human-motion forecasts or yielding into constrained configurations can still cause failures that short-horizon control cannot recover from. Current experiments use low-dimensional state and geometry inputs. Extending SCOPE to images and point clouds remains future work.