ALTER: Residual Denoising Enables Sample-Efficient
Multi-Agent Coordination on Demand

Dayi Dong, Maulik Bhatt, Aayushi Shrivastava, Lasse Peters, Negar Mehr

University of California, Berkeley

Coordination on demand with single-agent skill retention.

Overview

ALTER: Adaptation from Limited demonstrations for Team coordination with Existing-skill Retention.

We adapt pretrained single-arm diffusion policies to multi-agent tasks using limited collaborative demonstrations, without access to the original training data. The adapted policy coordinates when deployed in a team while retaining the ability to act independently. Execution is decentralized: each robot acts only on its own visual observations, without explicit inter-agent communication.

Residual coordination adapter

We freeze the pretrained base policy and jointly train a residual coordination adapter on multi-agent demonstrations and policy-distilled single-arm replay.
Dϕ=Dbase+Δϕ

Frozen base policyKeep the pretrained denoiser fixed during adaptation.

Residual adaptationLearn corrections for multi-agent coordination.

Policy-distilled replayRetain source-task behavior using rollouts of the frozen policy.

Technical details
The residual coordination adapter corrects the frozen base prediction at every denoising step.

Residual denoising. The adapter receives the noised action chunk, noise level, local observation, and base prediction. Its visual conditioning combines frozen base-encoder features with a lightweight convolutional network. We zero-initialize its output pathway, so adaptation starts from the base policy.

Training objective. We jointly train the adapter on collaborative demonstrations and policy-distilled replay. The coordination loss applies denoising regression to the adapted denoiser; the replay loss penalizes the residual on source-task inputs:

ℒ=ℒcoord+λretℒreplay

The retention weight λret trades coordination gain against source-task retention. A zero residual recovers the base policy; gradients update only the adapter.

Decentralized deployment. Each robot runs a copy of the same adapted policy using only its own RGB images. The adapter remains active on both source and collaborative tasks. Robots exchange no information and receive no external signal indicating whether a partner is present.

Two-arm coordination and source-task retention

In TwoArmPlaceWipe, one robot moves a tray to a waiting area while the other wipes the dirt beneath it. The robots must complete the wipe before returning the tray, and return the sponge to its pad.

Each robot uses only its own shoulder-camera RGB images.
ALTER · coordinated tray removal, wiping, and return

The adapted policy retains place-and-return and wipe-and-return behaviors with the coordination head active.

Adaptation budgets pair 20 / 20, 40 / 40, or 60 / 60 multi-agent / distilled single-arm demonstrations for ALTER, FS, and FT-mixed. FT-multi uses the same multi-agent counts and zero single-arm demonstrations. Combined source success averages place-and-return and wipe-and-return success. Tables I–II. Download data · Figure PDF

ALTER achieves the highest reported coordination success at each budget while retaining 97.0–97.5% combined source success. With 40 multi-agent and 40 distilled single-arm demonstrations, ALTER achieves 71.5% coordination success, exceeding FT-mixed (63.5%) and FS (59.0%) with 60 demonstrations of each type.

Hardware experiments

Two xArm7 robots coordinate to place a plush bird inside a lidded box. One robot removes the lid, the other places the bird, and the first replaces the lid. Each robot uses its own shoulder-view RGB camera.

Hardware setup and local visual observations.
ALTER · open, place, close

ALTER and FT-mixed each succeed in 18/20 coordination trials, compared with 13/20 for FS. Across the three source tasks, ALTER succeeds in 59/60 trials, compared with 47/60 for FT-mixed and 40/60 for FS.

Each method uses 69 multi-agent and 68 policy-distilled single-arm demonstrations. Source success combines 20 trials each of lid removal, lid replacement, and bird pick-and-place. Table IV.

Failures and near misses

Selected FT-mixed and FS trials illustrate failures and successful trials with nearly missed grasps or placements.

FT-mixed · failure
FS · failure
FS · success with near-missed grab
FT-mixed · success with near-missed placement

These qualitative examples are separate from the aggregate success rates in Table IV.

Citation

If you find ALTER useful in your research, please cite:

@article{dong2026alter,
  title={Residual Denoising Enables Sample-Efficient Multi-Agent Coordination on Demand},
  author={Dong, Dayi and Bhatt, Maulik and Shrivastava, Aayushi and Peters, Lasse and Mehr, Negar},
  journal={arXiv preprint arXiv:2609.32129},
  year={2026},
  url={https://arxiv.org/abs/2609.32129}
}