Legged locomotion · optimal control & estimation

A unified single-rigid-body NMPC + moving-horizon estimator for ANYmal-C

Output-feedback trotting of a 45 kg quadruped in MuJoCo, where a nonlinear model predictive controller and a nonlinear moving-horizon estimator share one reduced-order model, one symbolic source and one design language. The estimator reconstructs the base state and an external disturbance wrench from proprioception alone; the controller plans ground reaction forces inside friction cones at 100 Hz and feeds the disturbance forward.

MuJoCononlinear MPCmoving-horizon estimation real-time iterationdisturbance observer Python · C · SymPy

Closed-loop run at 1080p: stand → trot at 0.4 m/s → 40 N lateral push (red arrow, estimated live in orange) → 0.3 rad/s turn. Green arrows are the commanded ground reaction forces. No ground-truth state anywhere in the loop.

1.4–2.6 mrad
attitude estimation error
from IMU + leg kinematics only
7–10 mm/s
velocity estimation error
proprioceptive leg odometry
0.62 ms
controller cycle, 100 Hz
0.56 ms prep + 56 µs feedback
5 / 5
noise seeds complete
stable to 2× nominal noise
What it does

One reduced-order model, used twice

The single-rigid-body dynamics are written once, symbolically, and instantiated as both the controller and the estimator. The controller decides ground reaction forces over half a second of gait; the estimator reconstructs what the controller needs to know — the base state and the disturbances acting on it — from the sensors a real quadruped actually has.

Controller — nonlinear MPC

50-stage horizon (0.5 s), 100 Hz. Chooses four 3-D contact forces subject to friction pyramids, normal-force limits and attitude bounds, with the gait-scheduled contact sequence baked into the prediction. Feeds the estimated disturbance force forward.

12 states · 12 inputs · 28 constraint rows/stage

Estimator — moving-horizon

20-node window (0.2 s), 100 Hz. Fuses IMU attitude, gyro, and leg odometry from the settled stance feet, with the base state augmented by a 6-D disturbance wrench so the push is observable and can be fed forward.

18 states · 6 disturbance inputs · 21 outputs
Framework

Closed-loop architecture

Every 10 ms the estimator turns the sensors into a state and a disturbance estimate; the planner turns that into references, per-stage parameters and swing targets; the controller returns the ground reaction forces; a 500 Hz leg controller maps them to joint torques.

Gait scheduler + planner footholds · swing curves · references · foot anchoring RTI moving-horizon estimator M = 20 · 100 Hz state + disturbance wrench RTI nonlinear MPC N = 50 · 100 Hz forces in friction cones Leg controller — 500 Hz τ = −Jᵀf + bias (stance) · Cartesian PD (swing) MuJoCo ANYmal-C · 18 DoF · 2 ms x̂, d̂ xʳ, fʳ, π f₀ τ IMU, q, q̇, τ x̂, d̂

Shared model

Dynamics, Jacobians, constraints and outputs are emitted once from a SymPy source with common-subexpression elimination, then compiled into both the controller and the estimator — they agree bitwise for identical inputs.

Real-time iteration

One Newton-type iteration per sample, split into a preparation phase that runs before the measurement and a feedback phase that applies within microseconds of it — the controller's first force lands 0.11 µs after the estimate.

Disturbance feed-forward

The estimated force is fed to the controller; the estimated torque (swing-leg reaction at gait frequency) is deliberately not — the ablation shows why.

Results

Closed-loop performance

All plots are from a single 8-second run with sensor noise and no ground-truth feedback. Shaded bands mark the 40 N lateral push (red) and the turn (blue).

Velocity tracking
Forward speed: command, true and estimate
command true estimate
The body tracks the forward command through the push and into the turn.
Attitude stability
Roll and pitch over the run
roll pitch
Roll stays within ±5° and pitch within ±3.6° across push and turn.
Disturbance estimation
Applied vs estimated lateral force
applied push estimated
A 40 N push is recovered at 39.5 N by the augmented disturbance state.
Push recovery — why feed-forward matters
Lateral drift with and without compensation
no compensation force feed-forward
Feeding the estimated force forward cuts the push-induced drift from 0.74 m to 0.15 m.
Estimation accuracy
RMS estimation errors
RMS error over the walking phase — from proprioception only.
Real-time timing vs. horizon
Per-cycle time vs horizon length
preparation feedback
Cost grows linearly in the horizon; the 50-stage controller uses ~9% of a 10 ms budget.

The full scenario, in numbers

QuantityResultCondition
Attitude estimation error (roll/pitch/yaw)1.4 / 2.6 / 1.7 mradrms, walking phase
Velocity estimation error9 / 10 / 7 mm/sproprioception only
Base height regulation0.452 ± 0.002 m0.45 m commanded
Roll / pitch excursion≤ 0.089 / 0.063 radwhole run
40 N push — estimate & drift39.5 N · 0.15 mvs 0.74 m uncompensated
Controller cycle (N=50)0.56 ms + 56 µsprep + feedback, 100 Hz
Estimator cycle (M=20)0.24 ms + 37 µs1.1 µs instantaneous output
Robustness5 / 5 seeds · to 2× noisefull scenario, no fall
Achievements

What this project demonstrates

On the solver. The NMPC and MHE solvers are my own — a constrained real-time-iteration NMPC solver and its moving-horizon-estimation variant, developed as an extension of my master's thesis on SLQP-MPC. That work is ongoing and unpublished, so the solver algorithms are not part of this project; what is public is the modelling, the simulation framework, the architecture and the results shown here.
Code & report

Explore the project

Stack

MuJoCo 3SymPy code-genC (ctypes) NumPyMatplotlibReal-time iteration SRBD dynamicsFriction-cone constraints Leg odometryDisturbance observer