UC Berkeley
ICON Lab

SAMBAR: Selective Anchoring via Method of Multipliers for Balanced Knowledge Acquisition and Retention in Vision-Language-Action Models

Overview

SAMBAR overview: importance-weighted anchoring vs. SFT across sequential finetuning
An overview of SAMBAR: a replay-free continual learning algorithm that prevents catastrophic forgetting during VLA finetuning without requiring access to the demonstrations of any previously learned task. A policy learns tasks $T_1, T_2, T_3$ sequentially using only the current task's demonstrations. Orange, blue, and green indicate parameter importance for $T_1$, $T_2$, and $T_3$, with darker shades indicating greater importance. Standard supervised finetuning (SFT) overwrites parameters important to earlier tasks, causing forgetting. SAMBAR limits changes to those parameters using method of multipliers along with importance weighting while allowing other parameters to adapt to new tasks.

Hardware Experiments

SFT and SAMBAR performance on an xArm7 manipulator. Videos play at 10× speed.

Both policies are finetuned on cake-on-cake-tray and then on remove-lid-and-place-bird, in sequence. SAMBAR remembers the pretraining task wheras SFT forgets. Both have similar performance on the both finetuning tasks.
eraser-in-bowl
SFT
SAMBAR (ours)
cake-on-cake-tray
SFT
SAMBAR (ours)
remove-lid-and-place-bird
SFT
SAMBAR (ours)

Sequential continual learning on hardware

SAMBAR is the only method that remembers all four pretraining tasks while acquiring both new tasks at 93.3% Weighted Avg SR against 86.7% for the best baseline.
Pretrain SRPretrain SR90SFT90RETAIN100EWC100SAMBAR(Ours)0255075100Success Rate (%)
cake-on-cake-traycake-on-cake-tray80SFT100RETAIN80EWC80SAMBAR(Ours)0255075100Success Rate (%)
remove-lid-and-place-birdremove-lid-and-place-bird60SFT0RETAIN40EWC80SAMBAR(Ours)0255075100Success Rate (%)
Weighted Avg SRWeighted Avg SR83.3SFT76.7RETAIN86.7EWC93.3SAMBAR(Ours)0255075100Success Rate (%)
Show the exact numbers
Method Pretraining tasks ↑ Pretrain SR ↑ Finetune tasks ↑ Weighted Avg SR ↑
bird-in-bintwo-birds-in-bineraser-in-mugeraser-in-bowl cake-on-cake-trayremove-lid-and-place-bird
SFT5/55/55/53/590.04/53/583.3
RETAIN ($\alpha=0.5$)5/55/54/54/590.05/50/576.7
EWC5/55/55/55/5100.04/52/586.7
SAMBAR (ours)5/55/55/55/5100.04/54/593.3
Sequential continual learning on hardware. Individual tasks are reported as the number of successful rollouts out of 5. Pretrain SR is the percentage of the 20 pretraining rollouts; Weighted Avg SR is the percentage of all 30 rollouts.

Finetuning on a single task

SAMBAR is the only method that have high performance on every task at 100% Weighted Avg SR against 88.0% for the best baseline.
Pretrain SRPretrain SR60SFT85RETAIN80EWC100SAMBAR(Ours)0255075100Success Rate (%)
Finetune SRFinetune SR100SFT100RETAIN80EWC100SAMBAR(Ours)0255075100Success Rate (%)
Weighted Avg SRWeighted Avg SR68SFT88RETAIN80EWC100SAMBAR(Ours)0255075100Success Rate (%)
Show the exact numbers
Method Pretraining tasks ↑ Pretrain SR ↑ Finetune SR ↑ Weighted Avg SR ↑
bird-in-bintwo-birds-in-bineraser-in-mugeraser-in-bowl
Finetune task: cake-on-cake-tray
SFT4/51/55/52/560.05/568.0
RETAIN ($\alpha=0.5$)5/54/53/55/585.05/588.0
EWC5/54/55/52/580.04/580.0
SAMBAR (ours)5/55/55/55/5100.05/5100.0
Finetune task: remove-lid-and-place-bird
SFT5/54/54/53/580.05/584.0
RETAIN ($\alpha=0.5$)5/55/53/55/590.00/572.0
EWC5/55/53/55/590.00/572.0
SAMBAR (ours)5/55/55/55/5100.04/596.0
Learning one new task on hardware. We pretrain $\pi_{0.5}$ on four tasks and then finetune it on each of the two finetuning tasks separately. Every task is evaluated with 5 rollouts. Pretrain SR is the percentage of the 20 pretraining rollouts that succeed; Weighted Avg SR is the percentage of all 25 rollouts.

Simulation Experiments

The three tasks we learn in sequence

pot-on-stove initial scene
pot-on-stove $T_1$
"turn on the stove and put the moka pot on it"
bowl-in-drawer initial scene
bowl-in-drawer $T_2$
"put the black bowl in the bottom drawer of the cabinet and close it"
bottle-on-rack initial scene
bottle-on-rack $T_3$
"put the wine bottle on the rack"
Starting from a policy pretrained on 32 LIBERO tasks, we finetune on pot-on-stove, then bowl-in-drawer, then bottle-on-rack, in the order shown. The first two are long-horizon tasks; the third is a goal task.

After learning all three tasks

SAMBAR performs every one of them, while SFT has has completely forgotten first task.
SAMBAR (Ours)
pot-on-stove ($T_1$)
bowl-in-drawer ($T_2$)
bottle-on-rack ($T_3$)
SFT
pot-on-stove ($T_1$) fully forgotten
bowl-in-drawer ($T_2$) forgotten
bottle-on-rack ($T_3$)
Rollouts of the final policies, played at 2× speed. SFT succeeds on pot-on-stove 0.0% of the time after learning the two later tasks, while SAMBAR still performs it 84.0% of the time.

Sequential Continual Learning

Every baseline falls to 0.0% on pot-on-stove after finetuning on two more tasks; SAMBAR holds 84.0% and leads Weighted Avg SR at 87.4% against 56.7% for the best baseline.
Pretrain SRPretrain SR12.9SFT17.4RETAIN45.4LoRA0EWC58.3L244.4SimpleRecipe88.2SAMBAR(Ours)0255075100Success Rate (%)
pot-on-stovepot-on-stove0SFT0RETAIN0LoRA0EWC0L20SimpleRecipe84SAMBAR(Ours)0255075100Success Rate (%)
bowl-in-drawerbowl-in-drawer60SFT68RETAIN18LoRA0EWC42L236SimpleRecipe70SAMBAR(Ours)0255075100Success Rate (%)
bottle-on-rackbottle-on-rack84SFT96RETAIN92LoRA100EWC80L290SimpleRecipe82SAMBAR(Ours)0255075100Success Rate (%)
Weighted Avg SRWeighted Avg SR15.9SFT20.6RETAIN44.6LoRA2.9EWC56.7L244.2SimpleRecipe87.4SAMBAR(Ours)0255075100Success Rate (%)
Show the exact numbers
Method Pretrain SR ↑ Success on each newly learned task ↑ Weighted Avg SR ↑
pot-on-stove bowl-in-drawer bottle-on-rack
Co-Training95.4 ±1.0100.0 ±0.098.0 ±5.688.0 ±13.695.4
SFT12.9 ±1.90.0 ±0.060.0 ±21.584.0 ±11.115.9
RETAIN17.4 ±0.70.0 ±0.068.0 ±16.296.0 ±6.820.6
LoRA45.4 ±1.30.0 ±0.018.0 ±13.692.0 ±10.444.6
EWC0.0 ±0.00.0 ±0.00.0 ±0.0100.0 ±0.02.9
L258.3 ±2.60.0 ±0.042.0 ±34.580.0 ±24.856.7
Simple Recipe Works44.4 ±3.20.0 ±0.036.0 ±14.290.0 ±8.844.2
SAMBAR w/o importance89.8 ±0.80.0 ±0.00.0 ±0.04.0 ±6.882.2
SAMBAR (ours)88.2 ±2.384.0 ±6.870.0 ±15.282.0 ±16.287.4
Sequential continual learning on LIBERO. Success rates in percent, averaged over 5 seeds with 10 episodes/seed; error bars are 95% confidence intervals, clipped at 100%. Higher is better (↑) everywhere. Weighted Avg SR is the task-count-weighted mean over all 35 tasks.

Single-Task Finetuning

Narrowly pretrained (32 tasks)

SAMBAR retains 89.7% of the pretraining tasks while still learning the new one at 86.0%; the best baseline retains only 84.0%.
Pretrain SRPretrain SR56.1SFT66.5RETAIN75.4LoRA47.1EWC71.5L284SimpleRecipe89.7SAMBAR(Ours)0255075100Success Rate (%)
Finetune SRFinetune SR96SFT94RETAIN100LoRA96EWC74L286SimpleRecipe86SAMBAR(Ours)0255075100Success Rate (%)
Weighted Avg SRWeighted Avg SR57.3SFT67.3RETAIN76.2LoRA48.6EWC71.6L284SimpleRecipe89.6SAMBAR(Ours)0255075100Success Rate (%)
Show the exact numbers
MethodPretrain SR ↑Finetune SR ↑Weighted Avg SR ↑
Finetune task: pot-on-stove
Pretrained (no finetuning)93.9 ±1.46.0 ±16.7–
Co-Training93.8 ±1.394.0 ±6.893.8
SFT56.1 ±2.696.0 ±6.857.3
RETAIN66.5 ±2.194.0 ±11.167.3
LoRA75.4 ±1.8100.0 ±0.076.2
EWC47.1 ±1.496.0 ±11.148.6
L271.5 ±2.274.0 ±16.771.6
Simple Recipe Works84.0 ±1.386.0 ±16.784.0
SAMBAR w/o importance91.8 ±1.44.0 ±6.889.1
SAMBAR (ours)89.7 ±1.786.0 ±18.889.6
Finetune task: bowl-in-drawer
Pretrained (no finetuning)93.9 ±1.40.0 ±0.0–
Co-Training93.1 ±1.7100.0 ±0.093.3
SFT23.1 ±1.4100.0 ±0.025.4
RETAIN27.4 ±1.896.0 ±6.829.5
LoRA74.0 ±3.0100.0 ±0.074.8
EWC19.6 ±0.990.0 ±12.421.7
L264.9 ±2.378.0 ±16.265.3
Simple Recipe Works74.1 ±1.292.0 ±10.474.5
SAMBAR w/o importance90.6 ±1.938.0 ±30.989.0
SAMBAR (ours)91.8 ±1.296.0 ±6.891.9
Learning one new task from a narrowly pretrained policy. $\pi_{0.5}$ pretrained on 32 LIBERO tasks (8 from each suite), finetuned on two LIBERO-Long tasks. 5 seeds × 10 episodes/seed; error bars are 95% confidence intervals, clipped at 100%. Weighted Avg SR is over all 33 tasks. Forgetting is severe here: every baseline forgets the pretraining tasks, whereas SAMBAR remembers them.

Broadly pretrained (117 tasks)

Forgetting is mild from a broadly pretrained policy, but SAMBAR still has the highest Weighted Avg SR at 91.9%.
Pretrain SRPretrain SR89SFT89.7RETAIN88.1LoRA89.5EWC84.8L291.9SAMBAR(Ours)0255075100Success Rate (%)
Finetune SRFinetune SR100SFT100RETAIN100LoRA100EWC84L292SAMBAR(Ours)0255075100Success Rate (%)
Weighted Avg SRWeighted Avg SR89.1SFT89.8RETAIN88.2LoRA89.6EWC84.8L291.9SAMBAR(Ours)0255075100Success Rate (%)
Show the exact numbers
MethodPretrain SR ↑Finetune SR ↑Weighted Avg SR ↑
Finetune task: pot-on-stove
Pretrained (no finetuning)93.2 ±1.24.0 ±11.1–
SFT89.0 ±0.9100.0 ±0.089.1
RETAIN89.7 ±1.2100.0 ±0.089.8
LoRA88.1 ±0.6100.0 ±0.088.2
EWC89.5 ±1.4100.0 ±0.089.6
L284.8 ±0.784.0 ±11.184.8
SAMBAR w/o importance89.7 ±1.456.0 ±32.489.5
SAMBAR (ours)91.9 ±0.992.0 ±13.691.9
Finetune task: items-into-basket
Pretrained (no finetuning)93.2 ±1.20.0 ±0.0–
SFT93.2 ±1.5100.0 ±0.093.3
RETAIN93.3 ±2.196.0 ±11.193.4
LoRA93.4 ±0.9100.0 ±0.093.5
EWC93.5 ±1.2100.0 ±0.093.6
L293.8 ±1.292.0 ±13.693.8
SAMBAR w/o importance94.3 ±0.536.0 ±11.193.8
SAMBAR (ours)94.3 ±1.3100.0 ±0.094.4
Finetune task: mugs-on-plates
Pretrained (no finetuning)93.2 ±1.20.0 ±0.0–
SFT80.7 ±2.096.0 ±11.180.8
RETAIN85.4 ±1.188.0 ±22.285.5
LoRA91.9 ±1.3100.0 ±0.091.9
EWC78.8 ±0.896.0 ±11.178.9
L286.1 ±0.680.0 ±30.486.0
SAMBAR w/o importance93.2 ±1.276.0 ±20.893.1
SAMBAR (ours)90.7 ±1.192.0 ±13.690.8
Learning one new task from a broadly pretrained policy. $\pi_{0.5}$ pretrained on 117 LIBERO tasks. 5 seeds × 5 episodes/seed; error bars are 95% confidence intervals, clipped at 100%. Weighted Avg SR is over all 118 tasks. Forgetting is mild in this regime, and SAMBAR learns each new task without significantly forgetting the 117 pretrained tasks.

BibTeX

@article{shrivastava2026sambar,
  title   = {SAMBAR: Selective Anchoring via Method of Multipliers for Balanced
             Knowledge Acquisition and Retention in Vision-Language-Action Models},
  author  = {Shrivastava, Aayushi and Zhou, Xunlan and Zhao, Hongrui and
             Chen, Ziyu and Mehr, Negar},
  journal = {arXiv preprint arXiv:2609.32108},
  year    = {2026}
}