RoboDojo is a unified sim-and-real benchmark for evaluating generalist robot manipulation policies across 42 simulation tasks and 18 real-world tasks using three robot embodiments.

Stars

311

7-day growth

No data

Forks

27

Open issues

6

License

MIT

Last updated

2026-07-28

AI repository intelligence
FR-AI / ANALYSIS

Why it is worth attention

It provides a comprehensive, challenging, and reproducible evaluation framework with five distinct capability dimensions, heterogeneous parallel simulation, and a publicly maintained leaderboard.

Who it is for

  • Robot learning researchers
  • Policy developers and engineers
  • Robotics benchmarking community
  • Academic labs evaluating manipulation policies

Use cases

  • Benchmarking policy generalization and memory skills
  • Testing long-horizon manipulation capabilities
  • Comparing sim-to-real transfer performance
  • Submitting results to a community leaderboard

Strengths

  • Unified simulation and real-world evaluation across 60 tasks
  • Five capability dimensions (generalization, memory, precision, long-horizon, open) that probe diverse skills
  • Heterogeneous parallel simulation for fast, scalable feedback
  • Seed-controlled reproducibility and one-command leaderboard aggregation

Considerations

  • Eval-only release; policy training and integration must be handled by XPolicyLab
  • Requires Isaac Sim 5.1 and NVIDIA GPU hardware
  • Non-commercial license restricts commercial usage
  • Real-world evaluation depends on access to specific robot platforms

README quick start

RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies

Webpage | Document | Paper | Community | Leaderboard

https://private-user-images.githubusercontent.com/88101805/619409345-cc074c5d-4567-4418-8a29-1385aaba9d5b.mp4

✨ Highlights

Overview of RoboDojo. RoboDojo unifies efficient simulation evaluation and reproducible real-world testing for generalist robot manipulation, covering 42 simulation tasks, 18 real-world tasks, heterogeneous parallel simulation, RoboDojo-RealEval, XPolicyLab, and a continuously updated leaderboard.

RoboDojo is eval-only in this release: it provides the simulator client, benchmark tasks, asset/config validation, and result artifacts. Policy integration and policy servers are owned by XPolicyLab.

  • 🌐 Unified sim-and-real benchmark — 42 simulation tasks and 18 real-world tasks across 3 robot embodiments for generalist robot manipulation.
  • 🧭 Five capability dimensions — Generalization, Memory, Precision, Long-Horizon, and Open, designed to probe different skills rather than simple object or layout reskins.
  • 🧗 Challenging by design — intentionally hard, diverse, long-horizon tasks that expose failures hidden by simpler benchmarks.
  • Heterogeneous parallel simulation — runs different tasks, scenes, and processes concurrently on Isaac Sim for fast, scalable feedback.
  • 🧱 Physically grounded assets — rigid, articulated, and deformable objects in a single configuration-driven scene.
  • 🤖 Integrate once, evaluate everywhereXPolicyLab unifies 40+ policies behind one interface for both simulation and real-world runs.
  • 📊 Reproducible & leaderboard-ready — seed-controlled layouts and one-command summarize aggregation into a leaderboard table.

📚 Documentation

The RoboDojo documentation is the canonical reference. Key sections:

SectionDescription
Usage OverviewEnd-to-end walkthrough of the evaluation workflow.
Installation & Downloading (Assets and Data)Environment setup and downloading robot/object/layout assets/

Related repositories

Similar projects matched by category, topics, and programming language.

lopopolo
Featured
lopopolo GitHub avatar

harness-engineering

Harness Engineering is a methodology for improving coding agent outputs by carefully crafting the environment around them—providing curated context, tools, and executable constraints that encode an organization’s nonfunctional requirements and cumulative lessons.

AI & Machine LearningAI Agents
2,390
slvDev
Featured
slvDev GitHub avatar

esp32-ai

A 28.9 million parameter language model runs on an $8 ESP32-S3 microcontroller entirely on-device, generating simple stories at about 9.5 tokens per second.

AI & Machine LearningLarge Language Models
1,960
littledivy
Featured
littledivy GitHub avatar

mimic

mimic captures traffic from any iOS or web app and automatically generates a Python client library that lets you call the app's API like a regular library.

AI & Machine Learning
1,482