Too early
Repeat familiar behavior.
Timing-Aware Expert Querying
for VLA Policy Improvement
1Tsinghua University2Imperial College London3Dalian University of Technology4Tongji University5Peking University6SEEN·E Robotics
Better takeover timing. More useful supervision.
85 sec · English narration · CC
The takeover time changes which behavior the policy learns from.
Repeat familiar behavior.
Miss part of the target behavior.
Focus on the behavior still to learn.

Robot execution → expert supervision → aggregate demonstrations → update the policy.
Human-gated DAgger: a person chooses when to intervene. Robot-gated DAgger: the robot requests help. Both learn from expert demonstrations.
Video excerpts retain the source playback speeds. Equal-budget composition in the motivation film is schematic.
Internal features trigger assistance; expert feedback adjusts future takeover timing.
Bridge-PCA measures departure from a PCA reference fitted to in-distribution policy features. A residual above the calibrated threshold requests expert assistance.
s(oₜ) = ‖(I − VVᵀ)(zₜ − μ)‖₂
γᵢ₊₁ = γᵢ exp(−βyᵢ)
After a completed intervention, FTA lowers the threshold for a corrective reversal, raises it when expert and policy agree without substantial correction, and otherwise keeps it. Updates apply to subsequent episodes.
The policy is fine-tuned on retained successful expert suffixes mixed with the original demonstrations, then evaluated without assistance.
Within each task and backbone, methods share the starting checkpoint, retained expert-action budget, original-data mixing ratio, and policy-training settings. Updated policies run without assistance.
AUPRC: failure detection on shared autonomous rollouts. TASR: the fraction of retained expert actions aligned with the shifted operation. OOD success: post-training task success.
TASR uses successful OOD references, manually annotated target intervals, sequence alignment, and a calibrated state tolerance. Simulation expert actions come from a task oracle; the Human-Gated baseline uses a predefined failure-triggered takeover rule.
| Backbone | Task | Offline BC | HG-DAgger | Diff-DAgger | TimelyDAgger |
|---|---|---|---|---|---|
| π₀.₅ | StackCube–Green | 49 | 29 | 34 | 81 |
| π₀.₅ | StackCube–Red | 48 | 39 | 55 | 76 |
| π₀.₅ | OpenDrawer–Handle | 46 | 41 | 55 | 68 |
| π₀.₅ | OpenDrawer–Place | 43 | 37 | 52 | 64 |
| π₀.₅ | StackPyramid–Green | 50 | 47 | 46 | 53 |
| π₀.₅ | StackPyramid–Red | 60 | 57 | 45 | 79 |
| π₀.₅ | StackPyramid–Blue | 75 | 67 | 54 | 80 |
| π₀.₅ | PickPlane | 0 | 63 | 84 | 81 |
| π₀.₅ | OpenDrawer–Pose | 42 | 35 | 51 | 66 |
| π₀.₅ | YCB–Object | 48 | 49 | 42 | 52 |
| π₀.₅ | Eggplant | 47 | 40 | 55 | 69 |
| OpenVLA | PickPlane | 30 | 26 | 14 | 50 |
| OpenVLA | StackCube–Green | 71 | 63 | 64 | 82 |
| X-VLA | PickPlane | 72 | 35 | 58 | 84 |
| X-VLA | StackCube–Green | 58 | 76 | 64 | 67 |
TimelyDAgger has the highest reported AUPRC, TASR, and post-training success in 13 of 15 settings for each metric. For success, Diff-DAgger leads on π₀.₅ / PickPlane and HG-DAgger leads on X-VLA / StackCube–Green.
The timing study compares six fixed takeover steps on OpenDrawer Grasp-OOD. The FTA ablation reports higher TASR and success on both tested tasks. TASR–success correlations compare four methods within each setting and describe association, rather than a causal estimate.

15 attempts per clip · selected trial sequences · 10× · silent.

Overall combines both OOD conditions. The setup figure reports the base policy’s success; the chart and video counters report post-training results.
Recordings include failures and resets. “Robot-gated DAgger” in the original clips refers to TimelyDAgger; its first middle-position attempt starts partway through insertion.
Robot-Gated DAgger · OOD1 · full collection recording
3 min · with audio
@misc{zhao2026timelydagger,
title = {TimelyDAgger: Timing-Aware Expert Querying for VLA Policy Improvement},
author = {Zhixuan Zhao and Peiyan Li and Enhao Zhang and Yueran Tao and Hao Wang and Chenghao Yue and Lei Lv and Wentao Zhao and Jiahao Chen and Xin Liu and Kangyao Huang and Yu Luo and Huaping Liu},
year = {2026},
eprint = {2609.33157},
archivePrefix = {arXiv},
primaryClass = {cs.RO},
url = {https://arxiv.org/abs/2609.33157}
}