SAVRN
Search Contact SAVRN

SAVRN Model Hub · Datasets by Task

Robotics Datasets

32 datasets in the SAVRN Model Hub for robotics, from publishers including IPEC at Shanghai AI Laboratory, Jiu, Kun Zhao, Yifanwin.

32 datasets.

We provide a set of datasets used for post-training of GR00T N1. Each dataset is a collection of trajectories from different robot embodiments and tasks. Users can download a specific subset of data by specifying the dataset name. 1. Option 1: with huggingface-cli Replace gr1armsonly.CanSort/ with the dataset name you want to download. 2. Option 2: Github LFS Replace the "gr1armswaist.CupToDrawer" with the intended dataset folder

Publicly accessible cc-by-4.0

Dataset · Robotics

10Kh-RealOmin-OpenData

Genrobot.ai

Boasting over 13,000 hours of cumulative data and 5 million+ clips, it ranks as the largest open-source embodied intelligence dataset in the industry. Update Notes:Stage 3 data upload completed. 13,000+ hours of pure dual-hand data with frame-level alignment latency < 1ms Full high-precision trajectory reconstruction, breaking the limit of superficial open source, fully ready-to-use 3,000+ contributors and 10,000+ real household scenarios with exceptional diversity

Access requested at publisher cc-by-sa-4.0 n>1T

Dataset · Robotics

AgiBotWorld2026

AgiBot World

Real-World Embodied Intelligence Dataset As robotics research advances into real-world scenarios, the demand for authentic, high-quality data has become increasingly urgent. Following AGIBOT WORLD's "ImageNet moment," we now release the AGIBOT WORLD 2026 dataset. Built upon massive real-world scenes, it systematically spans pivotal research directions in embodied intelligence, designed to power the next generation of embodied agents. The AGIBOT WORLD 2026 dataset is collected from 100% real-world environments, covering commercial spaces, home, and other general-purpose scenarios. Collected on the AGIBOT G2 robot platform through a free-form collection mode, the dataset provides developers…

Publicly accessible cc-by-nc-sa-4.0 1K<n<10K

Dataset · Robotics

stereo-550

FPV Labs

FPV Labs Open-Source Stereo Hardware A first-person calibrated stereo RGB video dataset capturing everyday human manipulation across objects, materials, tools, and multi-step activities. Every session is recorded as a synchronized left/right camera pair with per-session stereo calibration, giving the visual geometry of hands, object interaction, state

Access requested at publisher other 1K<n<10K

Dataset · Robotics

droid

Remi Cadene

This dataset was created using LeRobot. One of the biggest open-source dataset for robotics with 27.044,326 frames, 92,223 episodes, 31,308 unique task description in natural language. Ported from Tensorflow Dataset format (2TB) to LeRobotDataset format (400GB) with the help from IPEC-COMMUNITY.

Publicly accessible apache-2.0 10M<n<100M

Dataset · Robotics

ABC-130k

XDOF

ABC-130k ABC-130k is the largest open-source robot teleoperation dataset. It contains bimanual manipulation trajectories collected on two-arm YAM stations. Episodes are distributed as MCAP files, with subtask annotations kept as separate artifacts so they can be revised or extended independently of the underlying episode data. For details on the accompanying paper, see abc.bot. Please see the GitHub repo here for code to train and deploy with this dataset. Dataset

Access requested at publisher apache-2.0 n>1T

Dataset · Robotics

HiFi-UMI-2K

Simple World Lab

HiFi-UMI-2K: High-Fidelity Robot-Free Manipulation Data 2,000 hours released · 6 accuracy · <40 µs synchronization Project Website | Dataset | Paper: arXiv:2607.25895 Examples from the HiFi-UMI corpus. Click the image to play the video. Introduction HiFi-UMI is a portable, high-fidelity bimanual capture system for collecting robot-free manipulation demonstrations.

Publicly accessible cc-by-4.0

Dataset · Robotics

ACE-Data-0

ACE Robotics

ACE-Data-0 Human-Centric Ambient Capture as Embodied Data Engine S-Lab, Nanyang Technological University, Singapore · ACE Robotics ACE turns real home environments into spatially calibrated, temporally synchronized recording studios for embodied AI. ▶ Demo video · Full story, figures, and interactive examples on the blog What this is Learning to act in the physical

Publicly accessible other 10K<n<100K

Retargeted AMASS for Robotics Project Overview This project aims to retarget motion data from the AMASS dataset to various robot models and open-source the retargeted data to facilitate research and applications in robotics and human-robot interaction. AMASS (Archive of Motion Capture as Surface Shapes) is a high-quality human motion capture dataset, and the SMPL-X model is a powerful tool for generating realistic human motion data. By adapting the motion data from AMASS

Publicly accessible cc-by-4.0 10K<n<100K

Retargeted AMASS for Robotics Project Overview This project aims to retarget motion data from the AMASS dataset to various robot models and open-source the retargeted data to facilitate research and applications in robotics and human-robot interaction. AMASS (Archive of Motion Capture as Surface Shapes) is a high-quality human motion capture dataset, and the SMPL-X model is a powerful tool for generating realistic human motion data. By adapting the motion data from AMASS

Publicly accessible cc-by-4.0 10K<n<100K

This project aims to retarget motion data from the AMASS dataset to various robot models and open-source the retargeted data to facilitate research and applications in robotics and human-robot interaction. AMASS (Archive of Motion Capture as Surface Shapes) is a high-quality human motion capture dataset, and the SMPL-X model is a powerful tool for generating realistic human motion data. By adapting the motion data from AMASS to different robot models, we hope to provide a more diverse and accessible motion dataset for robot training and human-robot interaction. This open-source project includes the following: 1. Retargeted Motions: Motion files retargeted from AMASS to various robot models.…

Publicly accessible cc-by-4.0 10K<n<100K

Dataset · Robotics

Hy-Embodied-0.5-VLA-Data

Tencent

We introduce Hy-Embodied-0.5-VLA (Hy-VLA) — an end-to-end Vision-Language-Action system that spans the full robot learning stack: data collection, model design, pre-training, supervised fine-tuning, RL post-training, and real-world deployment. Built on the Hy-Embodied-0.5 MoT backbone, Hy-VLA integrates a flow-matching action expert, a compact memory encoder for multi-frame history, and a delta-chunk action representation decoupled from embodiment-specific kinematics. Powered by 10,000+ hours of high-fidelity UMI demonstrations collected via a custom fingertip interface with optical motion-capture, Hy-VLA achieves state-of-the-art results on the RoboTwin 2.0 benchmark (90.9% / 90.1% on…

Publicly accessible cc-by-4.0 1M<n<10M

Dataset · Robotics

pico-robotics-basic

Jiu

Builds on the Advanced edition by adding coarse action segmentation. Every sequence is divided into labelled temporal segments, so the data can be used directly for action recognition, temporal segmentation, and behaviour-understanding tasks without an annotation pass of your own. This is a gated dataset. Access requests are reviewed manually; submit one from the dataset page. Everything except segments.json is documented in the Segments within a sequence are contiguous and non-overlapping; frames that fit no class are labelled other. TBD — list the action classes here, with a one-line definition and the frame count for each, Labels are deliberately coarse: boundaries are approximate and…

Publicly accessible cc-by-4.0 10K<n<100K

Dataset · Robotics

FDM-gcodes

Kieran O'Connor

A massive, highly-diverse dataset of Fused Deposition Modeling (FDM) G-codes generated directly from Printables. This dataset is specifically designed for training machine learning models on raw 3D printing manufacturing instructions (G-code). It can be used for tasks like G-code generation, print failure prediction, semantic analysis of toolpaths, and printer-agnostic slice classification. To ensure a highly robust and diverse set of training data, every original 3D model from the source dataset has been iteratively sliced into multiple variants (default: 10 per model) using PrusaSlicer. The pipeline supports and natively slices.stl,.obj,.3mf,.step, and.amf files. For each variant, a…

Publicly accessible other n>1T

Dataset · Robotics

pico-robotics-advanced

Jiu

Builds on the Advanced edition by adding coarse action segmentation. Every sequence is divided into labelled temporal segments, so the data can be used directly for action recognition, temporal segmentation, and behaviour-understanding tasks without an annotation pass of your own. This is a gated dataset. Access requests are reviewed manually; submit one from the dataset page. Everything except segments.json is documented in the Segments within a sequence are contiguous and non-overlapping; frames that fit no class are labelled other. TBD — list the action classes here, with a one-line definition and the frame count for each, Labels are deliberately coarse: boundaries are approximate and…

Access requested at publisher cc-by-4.0 10K<n<100K

Dataset · Robotics

pico-robotics-annotated

Jiu

Builds on the Advanced edition by adding coarse action segmentation. Every sequence is divided into labelled temporal segments, so the data can be used directly for action recognition, temporal segmentation, and behaviour-understanding tasks without an annotation pass of your own. This is a gated dataset. Access requests are reviewed manually; submit one from the dataset page. Everything except segments.json is documented in the Segments within a sequence are contiguous and non-overlapping; frames that fit no class are labelled other. TBD — list the action classes here, with a one-line definition and the frame count for each, Labels are deliberately coarse: boundaries are approximate and…

Access requested at publisher cc-by-4.0 10K<n<100K

GR00T-N1.7 LIBERO-X backbone features — 90-task fine-tune (LEVEL1-3) Aligned rollouts of a LIBERO-X fine-tune of GR00T-N1.7 (rohansiva/gr00t-libero-x-90task) on the LIBERO-X simulator, over the exact 90 tasks that checkpoint was fine-tuned on (30 tasks × 3 difficulty levels, LEVEL1–LEVEL3). Train and eval task sets are identical by design, so this is an in-distribution dataset for the checkpoint. 90 tasks × 20 rollouts = 1,800 episodes (600 per level). Every GR00T inference

Access requested at publisher 1K<n<10K

Dataset · Robotics

MesaTask-CTRC-100-shuffled

Yifanwin

This dataset holds the starting state of a scene restoration task: compared with the target state, 1-3 objects in each scene have been moved to random positions and need to be put back. source.scene on each object identifies its scene. - position is the model origin (not the bounding-box center), in centimeters - size is the object's target extent; the model is already scaled to match it. (x, y) are the horizontal extents and z is the height. A model's own local up-axis may not align with z, so orient it using size as the reference - rotation is the complete absolute orientation as a quaternion in [x, y, z, w] order. Apply it directly, per object, and do not compose any additional rotation…

Publicly accessible cc-by-nc-4.0

Dataset · Robotics

MesaTask-CTRC-100-target

Yifanwin

This dataset holds the target state of a scene restoration task: each scene should be restored to the arrangement given here. source.scene on each object identifies its scene and can be used to pair with other data from the same batch. The dataset is self-contained: every referenced model lives under assets/, there are no absolute paths, and no external asset library is required. - position is the model origin (not the bounding-box center), in centimeters - size is the object's target extent; the model is already scaled to match it. (x, y) are the horizontal extents and z is the height. A model's own local up-axis may not align with z, so orient it using size as the reference - rotation is…

Publicly accessible cc-by-nc-4.0

Teleoperated demonstration dataset for training a grasping policy on an SO-ARM101 arm with an AmazingHand dexterous hand (imitation learning). Pick up the cube with the dexterous hand — grasp a cube on the table using the dexterous hand. Collection flow: an operator moves the leader arm → the follower tracks it while its gripper proportionally drives the hand's open/close → joint angles and both camera streams are recorded synchronously. Per-episode flow: start from a fixed initial pose → perform one complete grasp (approach → grasp → lift → move → place) → end. Position coverage: the cube was placed at 4 different table positions, 5 episodes each, to teach positional generalization.…

Publicly accessible apache-2.0

Who Publishes These Datasets

Other tasks

See all