RESEARCHCSDWPaper ↗
ROBOT LEARNING · arXiv:2609.33145

Beyond State-as-Action

Exploiting Command–State Discrepancy
for Robot Imitation Learning

Peiyan Li1,6,†, Yueran Tao1,6,†, Enhao Zhang2,6,†, Zhixuan Zhao1,6,
Chenghao Yue1,6, Hao Wang3,6, Lei Lv4,6, Wentao Zhao1,6, Jiahao Chen5,6,
Xin Liu1,6, Kangyao Huang1,6, Yu Luo1,6,*, Huaping Liu1,6,*

1Tsinghua University2Imperial College London3Dalian University of Technology4Tongji University5Peking University6SEEN·E Robotics

† Equal contribution    * Corresponding authors

CSDW 5.5 code is available. Installation & usage ↗

Measured motion records what happened.
Commands preserve what was requested.
CSDW uses their temporal relationship to improve imitation learning.

RESEARCH OVERVIEW

The complete story in three minutes.

3 tasks · 4 settings · 380 trials

EXPLORE THE VIDEO

01 / WHEN

The value depends on the interaction.

Command supervision helps most in the evaluated insertion tasks, where motion can remain constrained.

Task and phase evidence →
02 / WHERE

Selected commands preserve the benefit.

Hybrid tests the local value of command information while keeping State targets elsewhere.

Selective retention →
03 / HOW

Temporal evidence guides training.

CSDW keeps every Command target and weights its loss, with the same policy at inference.

How CSDW works →
01 / MOTIVATION

What gets lost when
motion becomes the label?

State-as-Action constructs action targets from measured robot motion. Under interaction constraints, a robot can move little while its operator continues to command movement. Reconstructing targets from motion alone can discard part of that continuing control request.

We study when this distinction matters, whether selectively retaining command information preserves its benefit, and how to use temporal discrepancy to guide training.

An insertion example: free motion, constrained interaction, and the command–state trajectories across the demonstration. A gap is a cue to examine, not a direct measurement of force or intent.

02 / METHOD

A gap is a beginning.
Time adds context.

Two similar initial gaps can evolve differently. CSDW examines whether a command persists, whether measured state later progresses, and whether the request is reduced. It computes weights offline from paired training trajectories.

Paper method overview: discrepancy modeling, importance estimation, and weighted policy fine-tuning
The framework from the paper: paired trajectories → temporal evidence → frame weights → policy training.
01

Interpret the request

Calibrate channel ranges, polarity, and response times. Combine gap magnitude, command persistence, unresolved progress, and a discount for rapid responses into unfulfilled effort.

02

Use later evidence

Correct earlier effort using later progress and persistent residual gaps. Reducing a command can weaken support for the earlier request; the release itself remains a training target.

03

Allocate training emphasis

Pool channels and combine delayed progress, sustained effort, and effort changes. Spread evidence over the estimated response time to form continuous frame weights.

THE TRAINING RULE

wₜ = 1 + γ · clip(½Cᵣ(t) + ¼Cₗ(t) + ¼Cₑ(t), 0, 1)

ℒ = mean(wₜ · ℓₜCommand)

All command targets stay. Each target carries its own frame weight, including across overlapping action chunks.

During training
Fixed offline weights scale the original flow-matching loss.

At inference
The same observations, policy architecture, and inference process.

For scoring
No task-phase annotations, image semantics, or success labels.

03 / REAL-ROBOT EVALUATION

Three tasks. Three controlled comparisons.

The same paired recordings and policy architecture within each task. Change the supervision source, selectively retain commands, or weight the same targets.

380real-robot trials
3manipulation tasks
4supervision settings
π₀.₅shared policy architecture

Completion across three tasks

Paper Table I · rates use all trials per method

State · Command · Hybrid · CSDW. Edited evaluation footage accompanies the paper’s reported results.

Cable insertion

Sustained control against resistance. Reaching alignment is not the same as completing insertion.

Purple follows the paper’s CSDW color convention. All bars use a shared 0–100% scale.

35 trials per method · final full insertion

See every evaluated stage and success count
Stage success, reported in the paper. Counts correspond to the reported rounded percentages.
MethodCable alignmentCable insertionConnector alignmentConnector insertionPackage pickupPackage placement
State27/35 · 77.1%9/35 · 25.7%8/30 · 26.7%0/30 · 0.0%29/30 · 96.7%24/30 · 80.0%
Command27/35 · 77.1%21/35 · 60.0%11/30 · 36.7%10/30 · 33.3%26/30 · 86.7%24/30 · 80.0%
Hybrid25/35 · 71.4%25/35 · 71.4%11/30 · 36.7%11/30 · 36.7%28/30 · 93.3%23/30 · 76.7%
CSDW33/35 · 94.3%33/35 · 94.3%22/30 · 73.3%22/30 · 73.3%30/30 · 100.0%25/30 · 83.3%
H1 / WHEN

Command benefits depend on task and phase.

In cable insertion, State and Command both reach 77.1% alignment, but full insertion rises from 25.7% to 60.0%. The Command–State completion gap is also 33.3 percentage points for connector mating. Package placement stays at 80.0% for both.

Alignment alone does not capture successful engagement.

H2 / WHERE

Selective retention preserves the observed advantage.

Hybrid keeps Command arm targets in score-selected frames and State targets elsewhere, using the ordinary loss. It reaches 71.4% cable insertion and 36.7% connector mating success, compared with 25.7% and 0.0% for State.

This supports local value in selected commands; it does not establish optimal frame selection.

H3 / WEIGHTING

The same Command targets, with more useful emphasis.

CSDW reaches 94.3% cable insertion and 73.3% connector mating success, improving over uniform Command by 34.3 and 40.0 percentage points. Package placement rises from 24/30 to 25/30, a smaller observed change.

The evaluated weighting scheme improves both insertion tasks without changing inference.

04 / A CLOSER LOOK

The details that matter.

Hybrid is an evidence test.

Hybrid changes the supervision source: selected frames use Command arm targets, and other frames use future measured states. It uses the ordinary loss. CSDW keeps all Command targets and changes their weights. These answer different questions.

Hybrid demonstration with Command and State target intervals

The paper’s illustrative Hybrid timeline. A dataset-level threshold does not imply the same retained fraction in every demonstration.

Why not just rank the raw gap?

A large gap can reflect tracking delay or a request that is subsequently reduced. Temporal persistence, later progress, and residual support help interpret the initial request. The paper does not establish optimal frame selection.

Does looking ahead change deployment?

No. Future states are used only while computing weights offline from training demonstrations, within the same demonstration. The deployed policy receives no future observations.

What does “effort” mean here?

Unfulfilled and corrected effort are discrepancy-derived scores. They are not measured forces or ground-truth human intent. Limited directional progress is not necessarily zero robot motion.

What is actually controlled in the comparisons?

Within each task, supervision settings use the same paired recordings and π₀.₅ policy architecture. The datasets contain 482 cable-insertion demonstrations, 1,044 connector-mating demonstrations, and 151 package-transfer demonstrations. Gripper targets remain recorded commands across settings.

What can the experiments establish?

They demonstrate benefits under the evaluated robot, controller, and task conditions. Package placement improves only from 24/30 to 25/30 with CSDW. This small observed change should not be presented as a broad or statistically established advantage.

Weights are not normalized to unit mean: CSDW changes both temporal allocation and overall loss scale. The comparison does not isolate each component’s effect or prove that weighting location alone causes the gain.

How should the videos be read?

The main film covers the complete narrative; task videos illustrate recorded behavior. Reported success rates and trial counts follow Table I. The clips are edited for presentation and are not complete per-trial logs.

05 / FOR DATASET BUILDERS

Keep the command.
Study the gap.

When building robot datasets, preserve clearly documented command information alongside measured states. The difference between what was requested and what happened may contain useful control information that processed action labels cannot recover.

01 / PRESERVE

Keep both streams.

Retain original controller commands and measured states alongside processed training labels. Preserve timestamps and the pairing needed to recover their relationship.

02 / DOCUMENT

Make “action” unambiguous.

Specify the control interface, units, coordinate frames, temporal alignment, and label transformations. State clearly whether targets are commands or reconstructed motion.

03 / INVESTIGATE

Treat the gap as evidence.

Study when and where command–state differences matter. Simulation benchmarks should expose both streams and vary controller dynamics and contact constraints.

Download the data-recording checklist ↓

An actionable checklist derived from the paper’s data-preservation recommendation.

Citation

@misc{li2026csdw,
  title = {Beyond State-as-Action: Exploiting Command-State Discrepancy for Robot Imitation Learning},
  author = {Peiyan Li and Yueran Tao and Enhao Zhang and Zhixuan Zhao and Chenghao Yue and Hao Wang and Lei Lv and Wentao Zhao and Jiahao Chen and Xin Liu and Kangyao Huang and Yu Luo and Huaping Liu},
  year = {2026},
  eprint = {2609.33145},
  archivePrefix = {arXiv},
  primaryClass = {cs.RO},
  url = {https://arxiv.org/abs/2609.33145}
}