The value depends on the interaction.
Command supervision helps most in the evaluated insertion tasks, where motion can remain constrained.
Task and phase evidence →Exploiting Command–State Discrepancy
for Robot Imitation Learning
1Tsinghua University2Imperial College London3Dalian University of Technology4Tongji University5Peking University6SEEN·E Robotics
Measured motion records what happened.
Commands preserve what was requested.
CSDW uses their temporal relationship to improve imitation learning.
EXPLORE THE VIDEO
Command supervision helps most in the evaluated insertion tasks, where motion can remain constrained.
Task and phase evidence →Hybrid tests the local value of command information while keeping State targets elsewhere.
Selective retention →CSDW keeps every Command target and weights its loss, with the same policy at inference.
How CSDW works →State-as-Action constructs action targets from measured robot motion. Under interaction constraints, a robot can move little while its operator continues to command movement. Reconstructing targets from motion alone can discard part of that continuing control request.
We study when this distinction matters, whether selectively retaining command information preserves its benefit, and how to use temporal discrepancy to guide training.
An insertion example: free motion, constrained interaction, and the command–state trajectories across the demonstration. A gap is a cue to examine, not a direct measurement of force or intent.
Two similar initial gaps can evolve differently. CSDW examines whether a command persists, whether measured state later progresses, and whether the request is reduced. It computes weights offline from paired training trajectories.

Calibrate channel ranges, polarity, and response times. Combine gap magnitude, command persistence, unresolved progress, and a discount for rapid responses into unfulfilled effort.
Correct earlier effort using later progress and persistent residual gaps. Reducing a command can weaken support for the earlier request; the release itself remains a training target.
Pool channels and combine delayed progress, sustained effort, and effort changes. Spread evidence over the estimated response time to form continuous frame weights.
wₜ = 1 + γ · clip(½Cᵣ(t) + ¼Cₗ(t) + ¼Cₑ(t), 0, 1)
ℒ = mean(wₜ · ℓₜCommand)
All command targets stay. Each target carries its own frame weight, including across overlapping action chunks.During training
Fixed offline weights scale the original flow-matching loss.
At inference
The same observations, policy architecture, and inference process.
For scoring
No task-phase annotations, image semantics, or success labels.
The same paired recordings and policy architecture within each task. Change the supervision source, selectively retain commands, or weight the same targets.
Paper Table I · rates use all trials per method
State · Command · Hybrid · CSDW. Edited evaluation footage accompanies the paper’s reported results.
Sustained control against resistance. Reaching alignment is not the same as completing insertion.
Purple follows the paper’s CSDW color convention. All bars use a shared 0–100% scale.
35 trials per method · final full insertion
| Method | Cable alignment | Cable insertion | Connector alignment | Connector insertion | Package pickup | Package placement |
|---|---|---|---|---|---|---|
| State | 27/35 · 77.1% | 9/35 · 25.7% | 8/30 · 26.7% | 0/30 · 0.0% | 29/30 · 96.7% | 24/30 · 80.0% |
| Command | 27/35 · 77.1% | 21/35 · 60.0% | 11/30 · 36.7% | 10/30 · 33.3% | 26/30 · 86.7% | 24/30 · 80.0% |
| Hybrid | 25/35 · 71.4% | 25/35 · 71.4% | 11/30 · 36.7% | 11/30 · 36.7% | 28/30 · 93.3% | 23/30 · 76.7% |
| CSDW | 33/35 · 94.3% | 33/35 · 94.3% | 22/30 · 73.3% | 22/30 · 73.3% | 30/30 · 100.0% | 25/30 · 83.3% |
In cable insertion, State and Command both reach 77.1% alignment, but full insertion rises from 25.7% to 60.0%. The Command–State completion gap is also 33.3 percentage points for connector mating. Package placement stays at 80.0% for both.
Alignment alone does not capture successful engagement.
Hybrid keeps Command arm targets in score-selected frames and State targets elsewhere, using the ordinary loss. It reaches 71.4% cable insertion and 36.7% connector mating success, compared with 25.7% and 0.0% for State.
This supports local value in selected commands; it does not establish optimal frame selection.
CSDW reaches 94.3% cable insertion and 73.3% connector mating success, improving over uniform Command by 34.3 and 40.0 percentage points. Package placement rises from 24/30 to 25/30, a smaller observed change.
The evaluated weighting scheme improves both insertion tasks without changing inference.
Hybrid changes the supervision source: selected frames use Command arm targets, and other frames use future measured states. It uses the ordinary loss. CSDW keeps all Command targets and changes their weights. These answer different questions.

The paper’s illustrative Hybrid timeline. A dataset-level threshold does not imply the same retained fraction in every demonstration.
A large gap can reflect tracking delay or a request that is subsequently reduced. Temporal persistence, later progress, and residual support help interpret the initial request. The paper does not establish optimal frame selection.
No. Future states are used only while computing weights offline from training demonstrations, within the same demonstration. The deployed policy receives no future observations.
Unfulfilled and corrected effort are discrepancy-derived scores. They are not measured forces or ground-truth human intent. Limited directional progress is not necessarily zero robot motion.
Within each task, supervision settings use the same paired recordings and π₀.₅ policy architecture. The datasets contain 482 cable-insertion demonstrations, 1,044 connector-mating demonstrations, and 151 package-transfer demonstrations. Gripper targets remain recorded commands across settings.
They demonstrate benefits under the evaluated robot, controller, and task conditions. Package placement improves only from 24/30 to 25/30 with CSDW. This small observed change should not be presented as a broad or statistically established advantage.
Weights are not normalized to unit mean: CSDW changes both temporal allocation and overall loss scale. The comparison does not isolate each component’s effect or prove that weighting location alone causes the gain.
The main film covers the complete narrative; task videos illustrate recorded behavior. Reported success rates and trial counts follow Table I. The clips are edited for presentation and are not complete per-trial logs.
When building robot datasets, preserve clearly documented command information alongside measured states. The difference between what was requested and what happened may contain useful control information that processed action labels cannot recover.
Retain original controller commands and measured states alongside processed training labels. Preserve timestamps and the pairing needed to recover their relationship.
Specify the control interface, units, coordinate frames, temporal alignment, and label transformations. State clearly whether targets are commands or reconstructed motion.
Study when and where command–state differences matter. Simulation benchmarks should expose both streams and vary controller dynamics and contact constraints.
An actionable checklist derived from the paper’s data-preservation recommendation.
@misc{li2026csdw,
title = {Beyond State-as-Action: Exploiting Command-State Discrepancy for Robot Imitation Learning},
author = {Peiyan Li and Yueran Tao and Enhao Zhang and Zhixuan Zhao and Chenghao Yue and Hao Wang and Lei Lv and Wentao Zhao and Jiahao Chen and Xin Liu and Kangyao Huang and Yu Luo and Huaping Liu},
year = {2026},
eprint = {2609.33145},
archivePrefix = {arXiv},
primaryClass = {cs.RO},
url = {https://arxiv.org/abs/2609.33145}
}