HuGo: LLMs as Whole-Body Policy Code Designers
for Humanoid Loco-Manipulation

Seoyeon Choi1, Shizhao Ye1,2, Nicholas Bui1, Aayushi Shrivastava1, Kanghyun Ryu1,
Dhruva Tirumala3, Markus Wulfmeier4,†, Negar Mehr1
1ICON Lab at UC Berkeley, 2ShanghaiTech University, 3Google DeepMind, 4Nomagic
Preprint
†Work conducted in part while affiliated with Google DeepMind

TL;DR

Instead of engineering a reward or collecting demonstrations for each new task, an LLM writes the policy code itself. From a task description and the observation and command specifications, HuGo, Humanoid policy code Generation, generates closed-loop high-level policy code that runs on top of a frozen low-level whole-body controller. HuGo then improves that code by analyzing its own rollouts, in simulation and also directly on hardware.

With HuGo, we show that policy code generated in simulation runs directly zero-shot on hardware:

Push Button93.3%
Lift Box66.7%

Method

Overview of HuGo. HuGo generates closed-loop policy code for humanoid loco-manipulation from a task description, without expert demonstrations, reference motions, or reward design.

  1. Closed-loop execution -- The generated code reads the observation at every timestep and outputs a high-level command: planar locomotion, left and right end-effector positions, and body height. A frozen low-level whole-body policy converts this command into joint actions.
  2. Policy code generation -- Given only the task description, robot information, and the observation and command specifications, the LLM writes the task logic itself: where to stand, where to reach, and when to transition between stages. No example policies or skill libraries are provided.
  3. Evaluation and refinement -- The LLM diagnoses the policy's rollouts from numerical trajectories and selected video frames, and proposes a code diff: a localized update to the current policy rather than a full regeneration. Because it only needs rollouts, the same loop runs in simulation and on hardware.

Simulation Results

Rollouts of policy code generated and refined by HuGo on a Unitree G1 in IsaacSim. Every policy is obtained from a task description alone. The initial poses of the humanoid, objects, and goals are randomized, so the generated code must adapt its navigation path, approach pose, and command timing. Select a task to watch its rollouts and read the policy code the LLM wrote.

Five-task suite

The humanoid must pass through two low gates by lowering its body height.

Episode 1
Episode 2
Episode 3
Show generated policy code
Loading...

The humanoid must pass through two narrow gates by turning its body sideways.

Episode 1
Episode 2
Episode 3
Show generated policy code
Loading...

The humanoid must approach a button and press it using one of its end-effectors.

Episode 1
Episode 2
Episode 3
Show generated policy code
Loading...

The humanoid must approach a box, push it to within 10 cm of a target, and hold it there for one second.

Episode 1
Episode 2
Episode 3
Show generated policy code
Loading...

The humanoid must approach a box resting on a podium, grasp it, and lift it above 1.0 m.

Episode 1
Episode 2
Episode 3
Show generated policy code
Loading...

Tasks from HDMI

We also compare against HDMI, which trains a policy per task by imitating reference motions extracted from human videos, on two of its tasks. HuGo generates these policies from a task description without any demonstrations. Each video below is a different generated policy.

The humanoid must approach a door, push it open, and walk through.

Policy 1
Policy 1 code
Loading...
Policy 2
Policy 2 code
Loading...

The humanoid must grasp a bread box, carry it, and place it at a target.

Policy 1
Policy 1 code
Loading...
Policy 2
Policy 2 code
Loading...

Refinement on Hardware

A policy generated entirely in simulation is deployed zero-shot to a Unitree G1 pressing a physical button, and then the same refinement loop is run on the real rollouts. Nothing is retrained and no expert demonstration is collected: each round observes one real-world rollout, diagnoses what went wrong, and proposes a code diff to the current policy code.

20%→90%
In simulation
90%sim policygenerated & refined in IsaacSim
On hardware
20%depth_0deployed zero-shot
10%depth_1_0 · dropped 3 cm closer, much stricter press gate
50%depth_1_1 5 cm closer, forward-reach gate
60%depth_2_0+3 cm press margin, 1.4 s hold
→
90%depth_3_0+4.5 cm press margin, 1.8 s hold

The full refinement trace

Three depths of refinement, which took total of four rounds, ran on hardware. Open a stage to see the artifact it produced.

Depth 1 · round 0 · candidate not accepted
depth_0 20% → depth_1_0 10%
  1. 1Hardware rollout

    15 episodes on the Unitree G1, 3 / 15 successful (20%). One rollout R is the only evidence this round reasons over — no demonstrations, no reward, no retraining.

    depth_0 rollout · 20% success
  2. 2The LLM writes its own analysis code

    The LLM generates an executable analysis function fanalyze, written for this rollout and this policy rather than fixed in advance. Applied to the complete trajectory it returns a numerical summary D and task-relevant keyframe indices kkey.

    analysis_code.py
    Loading…

    D is at most 200 words of quantities and events the LLM judged useful — distance extrema, threshold crossings, tracking errors, internal-state transitions. Here it reports a planned press point 0.44 m forward against a realized reach of 0.409 m.

    analysis_output.txt
    Loading…
  3. 3Keyframes handed for evaluation

    Keyframes alone can miss the broader progression of a rollout, while evenly spaced frames can miss brief events such as a contact, so the two are combined under a fixed frame budget B = 10. When fanalyze returns fewer than B keyframes, the remaining slots are filled with evenly spaced frames, giving the indices k. Episode 2 here.

    kkey from fanalyze (8)0, 36, 51, 101, 106, 229, 346, 371
    Evenly spaced fill (2)167, 287
  4. 4Feedback

    From D and the frames at k, the LLM returns natural-language feedback F describing what occurred and the likely cause. It rests on this single rollout and has no access to the current policy’s scalar success rate.

    Its <summary>, quoted in full:

    The policy’s intended strategy is straightforward: pick the arm corresponding to the button side, walk to a standing point offset in front of the button face, face into the pillar, move the active hand to a pre-press pose, then push inward and retract.

    In this rollout, that high-level plan mostly worked at the navigation level. The robot consistently used the left arm, approached the correct side of the pillar, reduced stand-point distance from 0.80 m to as low as 0.021 m, and achieved near-perfect yaw alignment. It entered the align and press phases multiple times, and the left hand came very close to the commanded press target, with minimum actual EE-to-target error of 0.018 m. Visually, the hand repeatedly reached the button area.

    But the task was not completed robustly. The controller cycled through approach/align/press/retract three times, ending still in press phase. That means it got close repeatedly without producing a decisive button actuation.

    The root cause is that the press geometry is too marginal. The policy declares “ready” too early and commands a press from a stance that is only barely reachable, often with the hand target near the forward action limit. Small pose errors then turn into contact errors, so the hand reaches near the button but not with reliable inward normal push.

    Full diagnosis (reasoning + summary)
    Loading…
  5. 5Code diff applied, then re-run

    Rather than regenerating the policy, the LLM proposes a localized diff Δ conditioned on the current policy code, preserving the logic that works and modifying only what F identifies as the fault. The candidate is then evaluated on the robot.

    Reasoning
    Loading…
    Diff
    Loading…
    Resulting policy code
    Loading…
    depth_1_0 rollout · 10% success
Depth 1 · round 1 · candidate accepted
depth_1_0 10% → depth_1_1 50%
  1. 1Hardware rollout

    10 episodes on the Unitree G1, 1 / 10 successful (10%). One rollout R is the only evidence this round reasons over — no demonstrations, no reward, no retraining.

    depth_1_0 rollout · 10% success
  2. 2The LLM writes its own analysis code

    The LLM generates an executable analysis function fanalyze, written for this rollout and this policy rather than fixed in advance. Applied to the complete trajectory it returns a numerical summary D and task-relevant keyframe indices kkey.

    analysis_code.py
    Loading…

    D is at most 200 words of quantities and events the LLM judged useful — distance extrema, threshold crossings, tracking errors, internal-state transitions. Here it reports a planned press point 0.44 m forward against a realized reach of 0.409 m.

    analysis_output.txt
    Loading…
  3. 3Keyframes handed for evaluation

    Keyframes alone can miss the broader progression of a rollout, while evenly spaced frames can miss brief events such as a contact, so the two are combined under a fixed frame budget B = 10. When fanalyze returns fewer than B keyframes, the remaining slots are filled with evenly spaced frames, giving the indices k. Episode 6 here.

    kkey from fanalyze (8)0, 15, 27, 69, 136, 172, 211, 242
    Evenly spaced fill (2)48, 102
  4. 4Feedback

    From D and the frames at k, the LLM returns natural-language feedback F describing what occurred and the likely cause. It rests on this single rollout and has no access to the current policy’s scalar success rate.

    Its <summary>, quoted in full:

    The policy’s intended strategy is a four-phase routine: choose the nearer/more appropriate arm, walk to an offset standing pose in front of the button, align the torso to face into the button, move the active hand to a pre-press pose, then advance to a press pose and retract.

    In this rollout, the navigation part worked well. The robot started already facing almost perfectly toward the button face, walked from 0.414 m away to within 0.010 m of its planned stand target, and reduced lateral hand alignment error from 0.320 to essentially zero. It transitioned through approach, align, press, retract, then attempted a second press. The left arm was chosen and repeatedly extended toward the button.

    The task still failed because the final press geometry was too marginal. At the best stance, the planned press point was about 0.44 m forward in the pelvis frame, while the observed left-hand forward reach only reached 0.409 m at best. Commanded hand targets matched the intended pre-press/press points closely, but the actual hand did not fully realize them. Since the button only responds to a true inward push on its face, this small reach/contact shortfall prevented successful activation.

    Full diagnosis (reasoning + summary)
    Loading…
  5. 5Code diff applied, then re-run

    Rather than regenerating the policy, the LLM proposes a localized diff Δ conditioned on the current policy code, preserving the logic that works and modifying only what F identifies as the fault. The candidate is then evaluated on the robot.

    Reasoning
    Loading…
    Diff
    Loading…
    Resulting policy code
    Loading…
    depth_1_1 rollout · 50% success
Depth 2 · round 0 · candidate accepted
depth_1_1 50% → depth_2_0 60%
  1. 1Hardware rollout

    10 episodes on the Unitree G1, 5 / 10 successful (50%). One rollout R is the only evidence this round reasons over — no demonstrations, no reward, no retraining.

    depth_1_1 rollout · 50% success
  2. 2The LLM writes its own analysis code

    The LLM generates an executable analysis function fanalyze, written for this rollout and this policy rather than fixed in advance. Applied to the complete trajectory it returns a numerical summary D and task-relevant keyframe indices kkey.

    analysis_code.py
    Loading…

    D is at most 200 words of quantities and events the LLM judged useful — distance extrema, threshold crossings, tracking errors, internal-state transitions. Here it reports a hand XY error to the press target of 0.041 m at best, 0.048 m at the end.

    analysis_output.txt
    Loading…
  3. 3Keyframes handed for evaluation

    Keyframes alone can miss the broader progression of a rollout, while evenly spaced frames can miss brief events such as a contact, so the two are combined under a fixed frame budget B = 10. When fanalyze returns fewer than B keyframes, the remaining slots are filled with evenly spaced frames, giving the indices k. Episode 7 here.

    kkey from fanalyze (10)0, 4, 14, 16, 36, 65, 75, 86, 103, 147
    Evenly spaced fillnone — fanalyze already returned B keyframes
  4. 4Feedback

    From D and the frames at k, the LLM returns natural-language feedback F describing what occurred and the likely cause. It rests on this single rollout and has no access to the current policy’s scalar success rate.

    Its <summary>, quoted in full:

    The policy’s strategy is a four-phase sequence: walk to a stand point offset from the button, align body and hand laterally, move the active hand to a pre-press pose, then press for a fixed duration and retract. In this rollout it chose the right hand, walked from (0.380, 0.695) to a good stance near the pillar, reduced stand-target distance from 0.466 m to 0.046 m, and achieved good yaw and lateral alignment. It entered align at step 16 and press at step 36, then held a pressing posture before retracting at step 86.

    What actually happened is that navigation and gross positioning worked, but the button press itself did not complete. The commanded hand target got close to the intended press point, but the observed hand never matched it closely enough: best observed XY error to the press target was 0.041 m, while it returned almost exactly to the pre-press target by the end. The robot therefore hovered or lightly contacted near the button instead of driving it inward by the required 0.02 m.

    The root cause of failure is interaction geometry and execution margin: the press target is too marginal, the actual hand undershoots it by several centimeters, and the policy retracts on a timer without verifying button displacement.

    Full diagnosis (reasoning + summary)
    Loading…
  5. 5Code diff applied, then re-run

    Rather than regenerating the policy, the LLM proposes a localized diff Δ conditioned on the current policy code, preserving the logic that works and modifying only what F identifies as the fault. The candidate is then evaluated on the robot.

    Reasoning
    Loading…
    Diff
    Loading…
    Resulting policy code
    Loading…
    depth_2_0 rollout · 60% success
Depth 3 · round 0 · candidate accepted
depth_2_0 60% → depth_3_0 90%
  1. 1Hardware rollout

    10 episodes on the Unitree G1, 6 / 10 successful (60%). One rollout R is the only evidence this round reasons over — no demonstrations, no reward, no retraining.

    depth_2_0 rollout · 60% success
  2. 2The LLM writes its own analysis code

    The LLM generates an executable analysis function fanalyze, written for this rollout and this policy rather than fixed in advance. Applied to the complete trajectory it returns a numerical summary D and task-relevant keyframe indices kkey.

    analysis_code.py
    Loading…

    D is at most 200 words of quantities and events the LLM judged useful — distance extrema, threshold crossings, tracking errors, internal-state transitions. Here it reports a minimum error to the press target of 0.023 m, against the 0.02 m inward push required.

    analysis_output.txt
    Loading…
  3. 3Keyframes handed for evaluation

    Keyframes alone can miss the broader progression of a rollout, while evenly spaced frames can miss brief events such as a contact, so the two are combined under a fixed frame budget B = 10. When fanalyze returns fewer than B keyframes, the remaining slots are filled with evenly spaced frames, giving the indices k. Episode 2 here.

    kkey from fanalyze (8)0, 107, 277, 281, 302, 346, 412, 539
    Evenly spaced fill (2)192, 475
  4. 4Feedback

    From D and the frames at k, the LLM returns natural-language feedback F describing what occurred and the likely cause. It rests on this single rollout and has no access to the current policy’s scalar success rate.

    Its <summary>, quoted in full:

    The policy’s intended strategy is a four-phase finite-state controller: walk to an offset “stand” point in front of the button, align body yaw and lateral body position, move the chosen arm to a prepress pose, then advance to a press pose and finally retract/hold. In this rollout it chose the right arm, approached the correct side of the pillar, and reached the stand region successfully.

    What actually happened is that navigation worked, but alignment took a long time. The robot first got near the stand target around step 107, then spent 170 steps in align making small stance corrections and recovering from a large temporary yaw disturbance. It finally became ready to press around step 277 and entered press at step 281. During press, the hand got close to the button, but not close enough: the minimum error to the intended press target was 0.023 m. Since the task requires 0.02 m inward displacement, this is a near miss.

    The root cause of failure is therefore manipulation underreach, not navigation. The controller can accurately reach the prepress pose, but its realized press pose never penetrates far enough inward before the fixed press timer expires and the policy retracts instead of retrying.

    Full diagnosis (reasoning + summary)
    Loading…
  5. 5Code diff applied, then re-run

    Rather than regenerating the policy, the LLM proposes a localized diff Δ conditioned on the current policy code, preserving the logic that works and modifying only what F identifies as the fault. The candidate is then evaluated on the robot.

    Reasoning
    Loading…
    Diff
    Loading…
    Resulting policy code
    Loading…
    depth_3_0 rollout · 90% success

BibTeX

@article{choi2026hugo,
  title={HuGo: LLMs as Whole-Body Policy Code Designers for Humanoid Loco-Manipulation},
  author={Choi, Seoyeon and Ye, Shizhao and Bui, Nicholas and Shrivastava, Aayushi and Ryu, Kanghyun and Tirumala, Dhruva and Wulfmeier, Markus and Mehr, Negar},
  journal={arXiv preprint arXiv:2609.30594},
  year={2026}
}