RESEARCHTimelyDAgger Paper on arXiv
ROBOT LEARNING · arXiv:2609.33157

TimelyDAgger

Timing-Aware Expert Querying
for VLA Policy Improvement

Zhixuan Zhao1,6,†, Peiyan Li1,6,†, Enhao Zhang2,6, Yueran Tao1,6,
Hao Wang3,6, Chenghao Yue1,6, Lei Lv4,6, Wentao Zhao1,6, Jiahao Chen5,6,
Xin Liu1,6, Kangyao Huang1,6, Yu Luo1,6,*, Huaping Liu1,6,*

1Tsinghua University2Imperial College London3Dalian University of Technology4Tongji University5Peking University6SEEN·E Robotics

† Equal contribution    * Corresponding authors

Better takeover timing. More useful supervision.

START WITH THE FILM

When should a robot ask for help?

85 sec · English narration · CC

01 / THE PROBLEM

Same expert budget. Different demonstrations.

The takeover time changes which behavior the policy learns from.

Too early

Repeat familiar behavior.

Too late

Miss part of the target behavior.

Timely

Focus on the behavior still to learn.

Figure 1 and the DAgger learning loop
Figure 1: expert-querying loop and early, timely, and late takeover

Robot execution → expert supervision → aggregate demonstrations → update the policy.

Human-gated DAgger: a person chooses when to intervene. Robot-gated DAgger: the robot requests help. Both learn from expert demonstrations.

Video excerpts retain the source playback speeds. Equal-budget composition in the motivation film is schematic.

02 / THE METHOD

Monitor the policy. Learn when to ask.

Internal features trigger assistance; expert feedback adjusts future takeover timing.

TimelyDAgger framework: build a gate, collect expert data with feedback, update the policy

Build the gate → collect demonstrations → update the policy.

Technical details

Bridge-PCA measures departure from a PCA reference fitted to in-distribution policy features. A residual above the calibrated threshold requests expert assistance.

s(oₜ) = ‖(I − VVᵀ)(zₜ − μ)‖₂

γᵢ₊₁ = γᵢ exp(−βyᵢ)

After a completed intervention, FTA lowers the threshold for a corrective reversal, raises it when expert and policy agree without substantial correction, and otherwise keeps it. Updates apply to subsequent episodes.

The policy is fine-tuned on retained successful expert suffixes mixed with the original demonstrations, then evaluated without assistance.

03 / THE TEST

Shift the task. Keep the expert budget fixed.

ID and OOD task configurations: changes to position, pose, or object identity
Evaluation protocol and metrics

Within each task and backbone, methods share the starting checkpoint, retained expert-action budget, original-data mixing ratio, and policy-training settings. Updated policies run without assistance.

AUPRC: failure detection on shared autonomous rollouts. TASR: the fraction of retained expert actions aligned with the shifted operation. OOD success: post-training task success.

TASR uses successful OOD references, manually annotated target intervals, sequence alignment, and a calibrated state tolerance. Simulation expert actions come from a task oracle; the Human-Gated baseline uses a predefined failure-triggered takeover rule.

04 / THE EVIDENCE

Better supervision. Stronger policies.

Post-training success for four acquisition methods in all 15 evaluated settings

Matched retained expert-action budgets · Table I.

Exact results and analysis
Matched retained expert-action budgets and training settings within each task and backbone. Bold numbers mark the highest reported success in each row.
BackboneTaskOffline BCHG-DAggerDiff-DAggerTimelyDAgger
π₀.₅StackCube–Green49293481
π₀.₅StackCube–Red48395576
π₀.₅OpenDrawer–Handle46415568
π₀.₅OpenDrawer–Place43375264
π₀.₅StackPyramid–Green50474653
π₀.₅StackPyramid–Red60574579
π₀.₅StackPyramid–Blue75675480
π₀.₅PickPlane0638481
π₀.₅OpenDrawer–Pose42355166
π₀.₅YCB–Object48494252
π₀.₅Eggplant47405569
OpenVLAPickPlane30261450
OpenVLAStackCube–Green71636482
X-VLAPickPlane72355884
X-VLAStackCube–Green58766467

TimelyDAgger has the highest reported AUPRC, TASR, and post-training success in 13 of 15 settings for each metric. For success, Diff-DAgger leads on π₀.₅ / PickPlane and HG-DAgger leads on X-VLA / StackCube–Green.

The timing study compares six fixed takeover steps on OpenDrawer Grasp-OOD. The FTA ablation reports higher TASR and success on both tested tasks. TASR–success correlations compare four methods within each setting and describe association, rather than a causal estimate.

05 / ON THE REAL ROBOT

Same task. A shifted insertion target.

Cable insertion training configuration and two shifted target-board positions

Offline BC

20.0% 3 / 15 successes

Human-Gated DAgger

33.3% 5 / 15 successes

TimelyDAgger

60.0% 9 / 15 successes

15 attempts per clip · selected trial sequences · 10× · silent.

Real-world results and recording notes
Post-training cable insertion success: BC, HG-DAgger, and TimelyDAgger under OOD1, OOD2, and overall

Overall combines both OOD conditions. The setup figure reports the base policy’s success; the chart and video counters report post-training results.

Recordings include failures and resets. “Robot-gated DAgger” in the original clips refers to TimelyDAgger; its first middle-position attempt starts partway through insertion.

Watch expert data collection

Robot-Gated DAgger · OOD1 · full collection recording

COLLECTION RECORDINGS

THE COMPLETE STUDY

The story in three minutes.

3 min · with audio

Citation

Download .bib
@misc{zhao2026timelydagger,
  title = {TimelyDAgger: Timing-Aware Expert Querying for VLA Policy Improvement},
  author = {Zhixuan Zhao and Peiyan Li and Enhao Zhang and Yueran Tao and Hao Wang and Chenghao Yue and Lei Lv and Wentao Zhao and Jiahao Chen and Xin Liu and Kangyao Huang and Yu Luo and Huaping Liu},
  year = {2026},
  eprint = {2609.33157},
  archivePrefix = {arXiv},
  primaryClass = {cs.RO},
  url = {https://arxiv.org/abs/2609.33157}
}