---
title: traceflow-realworld
canonical_url: "https://www.modelscope.ai/datasets/JoeyXuan/traceflow-realworld"
md_url: "https://www.modelscope.ai/datasets/JoeyXuan/traceflow-realworld.md"
repository: JoeyXuan/traceflow-realworld
last_updated: 2026-09-20
license: cc-by-4.0
storage_size: "209 GB"
downloads: 1051
stars: 0
---

# traceflow-realworld

> traceflow-realworld - An open-source dataset by JoeyXuan on ModelScope. The realworld train and tracebank source data collected in ARX ACone biomanual with quest3s remote teleoperation.

JoeyXuan/traceflow-realworld is a dataset on ModelScope. totalling 209 GB. licensed under cc-by-4.0.

- **Repository**: JoeyXuan/traceflow-realworld
- **License**: cc-by-4.0
- **Storage size**: 209 GB
- **Downloads**: 1051
- **Stars**: 0
- **Last updated**: 2026-09-20

Source: https://www.modelscope.ai/datasets/JoeyXuan/traceflow-realworld

---

# TraceFlow Real-World

[![Project Page](https://img.shields.io/badge/Project-Page-2ea44f?logo=githubpages&logoColor=white)](https://zhangjiaxuan-xuan.github.io/TraceFlow/)
[![Paper](https://img.shields.io/badge/arXiv-2609.20646-b31b1b?logo=arxiv&logoColor=white)](https://arxiv.org/abs/2609.20646)
[![Code](https://img.shields.io/badge/GitHub-Code-181717?logo=github&logoColor=white)](https://github.com/zhangjiaxuan-Xuan/TraceFlow)
[![Model](https://img.shields.io/badge/ModelScope-Models-624aff?logoColor=white)](https://modelscope.ai/models/JoeyXuan/traceflow-models)
[![Real-world Data](https://img.shields.io/badge/ModelScope-Real--world_Data-ff6a00?logoColor=white)](https://www.modelscope.ai/datasets/JoeyXuan/traceflow-realworld)

![TraceFlow method overview](assets/traceflow-method.png)

TraceFlow Real-World is a real-world bimanual manipulation dataset collected with an ARX AC One robot and Meta Quest 3 teleoperation. It is intended for imitation-learning and vision-language-action experiments, including conversion to PI0.5-compatible training formats.

## Current release

- 200 accepted episodes (approximately 209 GB)
- 2 collection-level language instructions, with 100 source episodes per instruction
- Robot state and action sampled at 50 Hz
- Three asynchronous RGB camera streams recorded at approximately 30 FPS
- Left wrist, right wrist, and third-person camera views
- Two 7-dimensional arm records per timestep (14 dimensions total, including grippers)
- Binary gripper command labels derived from measured gripper motion and filtered teleoperation input

The two task instructions used during data collection are:

1. Clean the table and place the cable and the single-sided tape on the table.
2. Place the tape in the upper drawer, then place the cable in the lower drawer.

These are full-session collection labels, not the final task labels used to
train the released policy. In particular, each episode of the first long task
contains three sequential behaviors that are split by annotation for training.

## Collection labels and derived training tasks

The first 100-episode collection task is non-destructively divided at its
annotated boundaries into three smaller training tasks per episode. The second
100-episode drawer task remains one training segment per episode. The resulting
training index therefore contains 400 segments across four task categories:
three tasks cut from the first long task, plus the drawer task. This segmentation
does not modify, overwrite, or duplicate the 200 source episodes. The segment
records are provided under
[`annotations/four_task_dataset_motion_trimmed/`](annotations/four_task_dataset_motion_trimmed/),
and arm usage is explicitly recorded in
[`ARM_USAGE_METADATA.json`](ARM_USAGE_METADATA.json):

1. Shelf objects: left arm only.
2. Fruits to box: right arm only.
3. Box, tape, and hammer: sequential bimanual execution, using the left arm for
   the box and then the right arm for the tape and hammer.
4. Drawer task: right arm only.

There is no simultaneous dual-arm motion in the current dataset. Training
pipelines should retain the complete bimanual action/state vector while using
`active_arm = left | right | none` as an Upper-VLM routing label. The inactive
arm should remain in its recorded hold/wait state rather than being discarded.

## Episode layout

Each directory under `data/perfect/` is one accepted episode:

```text
data/perfect/<episode_id>/
├── raw_demo.npz
├── time_alignment.npz
├── gripper_labels.npz
├── left_arm_camera.nut
├── right_arm_camera.nut
├── third_person_camera.nut
├── candidate.json
├── decision.json
├── quality_report.json
├── gripper_filter_report.json
└── timing.json
```

`raw_demo.npz` contains robot timestamps, observations, velocities, efforts,
actions, the original collection-level task text, and Quest controller
telemetry. For training, use the segment-level task label and frame boundaries
from the derived annotation index rather than assigning the first long-task
instruction unchanged to all three segments. `time_alignment.npz` maps every
50 Hz robot timestep to the nearest frame in each independent camera stream and
records frame age/lag. `gripper_labels.npz` contains filtered binary left/right
gripper commands and confidence masks.

The `.nut` files preserve the captured camera streams. Camera clocks are independent; consumers should use `time_alignment.npz` instead of aligning streams only by frame index.

## Data quality

Only episodes accepted by the collection and audit workflow are included in `data/perfect/`. Rejected, pending, automatically deleted, temporary, and FFmpeg diagnostic files are excluded from this release. The original source recordings are retained without resizing or destructive conversion.

## PI0.5 use

This repository stores lossless source episodes rather than a model-specific cache. A training conversion should:

1. select a common robot timeline from `time_alignment.npz`;
2. retrieve the aligned image from each camera stream;
3. convert the 14-dimensional robot state/action representation to the target PI0.5 embodiment schema;
4. use the derived segment-level task string as the language instruction for
   the four-task training index, while retaining the episode task string as
   collection provenance; and
5. preserve raw episodes and write converted shards separately.

## Derived end-effector kinematics

Every accepted episode now includes an immutable `eef_kinematics.npz` sidecar
derived from measured joint states with the official ARX X5 forward-kinematics
solver. The original `raw_demo.npz` files are unchanged.

The sidecar provides, on the original 50 Hz robot timeline:

- left/right 6D end-effector poses (`xyz` plus rotation vector);
- left/right 6D twists and 3D linear-speed magnitudes for kinematic keyframes;
- proper SE(3) transitions from timestep `t` to `t+1`; and
- the complete 14D bimanual PI0.5 action: right 6D delta + gripper 0/1,
  followed by left 6D delta + gripper 0/1.

Rotational deltas use rotation composition rather than Euler-angle subtraction.
Poses, twists and deltas are expressed in each arm's local base frame. See
[`metadata/EEF_KINEMATICS_DATA.md`](metadata/EEF_KINEMATICS_DATA.md) and
`registry/eef_kinematics_registry.jsonl` for the complete schema and per-episode
hashes.

## Privacy and responsible use

The videos were captured in a real laboratory environment and may contain incidental background information. Users must review local privacy, consent, and safety requirements before redistributing the data or deploying policies trained from it. Robot policies can produce unsafe motion and must be validated with independent limits, emergency-stop procedures, and human supervision.

## Citation

When using this dataset in a research publication, please cite:

```bibtex
@article{zhang2026traceflow,
  title   = {TraceFlow: Guiding Frozen Flow-Matching Robot Policies with Success and Failure Traces},
  author  = {Jiaxuan Zhang and Ruizhe Liu and Yu Zhang and Yanchao Yang},
  journal = {arXiv preprint arXiv:2609.20646},
  year    = {2026},
  url     = {https://arxiv.org/abs/2609.20646}
}
```

## License

The original TraceFlow real-world dataset is licensed under the
Creative Commons Attribution 4.0 International license (CC BY 4.0):
[https://creativecommons.org/licenses/by/4.0/](https://creativecommons.org/licenses/by/4.0/)

You may use, copy, modify, and redistribute the dataset, including for
commercial purposes. You must give appropriate credit to the creators,
provide a link to the license, and indicate whether changes were made.
Do not imply endorsement by the authors or their institutions.

When using this dataset in a research publication, please cite the
TraceFlow paper above. Include the dataset URL and license in any
redistribution; citation alone does not replace the other license conditions.

This license applies to original material the contributors have the
right to license. Separately identified third-party material retains its
own terms. Privacy, publicity, patent, and trademark rights are not
granted by this license. See [`DATA_USE_NOTICE.md`](DATA_USE_NOTICE.md) and
[`LICENSE.md`](LICENSE.md) for the open-use notice and license reference.
