news_article.exe
📰

Building a Streaming Robotics Learning Pipeline Using NVIDIA Cosmos3-DROID

2026年10月5日1 次浏览来源:MarkTechPost 阅读原文

In this tutorial, we design an end-to-end streaming robotics learning pipeline around the NVIDIA Cosmos3-DROID dataset without downloading its 707 GB repository locally. We first introspect the LeRobotDataset v3.0 structure and construct a metadata graph from info.json, task metadata, episode tables, and dataset statistics, then use HTTP byte-range access with PyArrow to selectively read Parquet row groups and columns. We convert individual episodes into state-action trajectories and analyze joint motion, gripper events, Cartesian end-effector paths, and action-frequency spectra before decoding only the required AV1 video windows through seek-based PyAV/FFmpeg access. We then normalize observations and actions using dataset statistics, construct an ACT-style chunked PyTorch dataset with...

In this tutorial, we design an end-to-end streaming robotics learning pipeline around the NVIDIA Cosmos3-DROID dataset without downloading its 707 GB repository locally. We first introspect the LeRobotDataset v3.0 structure and construct a metadata graph from info.json, task metadata, episode tables, and dataset statistics, then use HTTP byte-range access with PyArrow to selectively read Parquet row groups and columns. We convert individual episodes into state-action trajectories and analyze joint motion, gripper events, Cartesian end-effector paths, and action-frequency spectra before decoding only the required AV1 video windows through seek-based PyAV/FFmpeg access. We then normalize observations and actions using dataset statistics, construct an ACT-style chunked PyTorch dataset with optional visual conditioning, and train a multimodal behavior-cloning policy. Finally, we evaluate the learned policy through open-loop rollout with temporally ensembled action chunks, report per-joint MSE and R^2 against a mean-action baseline, visualize predicted versus ground-truth actions, and save the complete policy checkpoint for downstream use. Copy CodeCopiedUse a different Browser We initialize the Colab environment, install the required libraries, and configure the Cosmos3-DROID dataset, episode, video, and training parameters. We inspect the repository structure and identify the available data, video, and metadata shards without downloading the complete dataset. We then load the core metadata and task descriptions to understand the dataset schema, available state/action features, and episode organization. Copy CodeCopiedUse a different Browser We implement a byte-range Parquet reader that accesses only the required row groups and columns directly through the Hugging Face filesystem. We identify episode boundaries within a data shard and convert selected state and action fields into NumPy trajectories. We then visualize joint positions, gripper activity, Cartesian motion, action distributions, frequency spectra, and episode-duration statistics. Copy CodeCopiedUse a different Browser We build a seek-based video pipeline that retrieves only the required temporal window from an episode, rather than downloading an entire video shard. We support both PyAV and FFmpeg decoding paths to handle AV1 video efficiently and resize selected frames for lightweight processing. We also load dataset-level normalization statistics from stats.json, with an empirical fallback when those statistics are unavailable. Copy CodeCopiedUse a different Browser We load a configurable collection of episodes and optionally cache synchronized visual observations for a small subset to keep training computationally manageable. We construct an ACT-style PyTorch dataset that combines observation history and optional images with normalized future action chunks. We then define a chunked policy architecture that combines an MLP state encoder with an optional CNN vision encoder and predicts a sequence of future actions. Copy CodeCopiedUse a different Browser We initialize the chunked policy and optimize it with AdamW, OneCycle learning-rate scheduling, mixed-precision execution, gradient scaling, and gradient clipping. We use a Smooth L1 loss to make behavior cloning more robust to noisy or variable teleoperation actions. We train the policy for the configured number of epochs while tracking both training and validation losses to monitor learning behavior. Copy CodeCopiedUse a different Browser We evaluate the trained policy through open-loop rollout and combine overlapping action predictions using exponentially weighted temporal ensembling. We compute per-joint MSE, baseline error, and R^2, and plot predicted actions against ground-truth trajectories, along with the training/validation loss curves. Finally, we save the trained model with normalization statistics and configuration metadata so we can reuse the policy in subsequent experiments. In conclusion, we showed how to turn a massive real-world robot dataset into a learning pipeline while keeping storage and data-transfer requirements extremely low. We used metadata-driven episode discovery, column- and row-group-level Parquet projection, and seek-based video decoding to access only the information required for analysis and training rather than materializing the full dataset. We combined proprioceptive state history with optional visual observations to train a chunked behavior-cloning policy and used temporal ensembling to obtain smoother action predictions during open-loop evaluation. The resulting workflow gives us a compact but extensible foundation that we can scale across additional shards, failure demonstrations, camera views, language instructions, or alternative action representations for more sophisticated robotics and vision-language-action experiments. Check out the FULL CODES here. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us The post Building a Streaming Robotics Learning Pipeline Using NVIDIA Cosmos3-DROID appeared first on MarkTechPost.

> 分享: