QwenLM GitHub avatar

Qwen-RobotManip

QwenLM

Qwen-RobotManip is a vision-language-action foundation model that aligns heterogeneous robot manipulation data via representation, motion, and behavior alignment to achieve strong out-of-distribution generalization and cross-embodiment transfer.

Stars

116

7-day growth

No data

Forks

3

Open issues

11

License

No data

Last updated

2026-06-29

AI repository intelligence
FR-AI / ANALYSIS

Why it is worth attention

It introduces a three-dimensional alignment framework that makes diverse open-source manipulation data trainable together, achieves state-of-the-art OOD benchmark results, and ranks #1 on the RoboChallenge Table30 v1 generalist track using only public datasets.

Who it is for

  • Robotics researchers studying manipulation foundation models
  • AI engineers developing generalist robot policies
  • Academic labs working on imitation learning and transfer
  • Industry practitioners in automated manufacturing or service robotics

Use cases

  • Open-ended instruction following in real-robot manipulation with no predefined task list
  • Cross-embodiment transfer of manipulation skills across single-arm, dual-arm, dexterous hand, and mobile platforms
  • Reactive recovery from execution errors without explicit recovery scripts
  • OOD generalization to novel scenes, objects, and language instructions

Strengths

  • Three-dimensional alignment (representation, motion, behavior) unifies heterogeneous multi-source data
  • 80D canonical action space with masks accommodates diverse embodiments in one model
  • Outperforms prior methods on LIBERO-Plus, RoboTwin-Clean2Rand, and EBench; #1 in RoboChallenge Table30
  • Uses 38,100 hours of open-source data including 24,808 hours of synthetic human-to-robot demonstrations

Considerations

  • Model weights are not publicly released and no plan to release them
  • Training requires large-scale data and compute, potentially limiting reproducibility
  • Primarily focused on tabletop manipulation; applicability to other domains not demonstrated

README quick start

Qwen-RobotManip

Alignment Unlocks Scale for Robotic Manipulation Foundation Models

Qwen Team

📑 Technical Report |
📖 Blog |
🖥️ Demo

Welcome to the official repository of Qwen-RobotManip. Here, you can find official information about Qwen-RobotManip and post your questions (Issues).

Note: There is currently no plan to release the model weights for Qwen-RobotManip or Qwen-RobotNav. We will continue adding report resources that can be publicly released to this repository.

🎬 Demo

Qwen-Omni × Qwen-RobotManip — Qwen-Omni observes the scene, randomly proposes manipulation tasks via speech, and judges execution in real time. The two demos below are English-subtitled versions. Qwen-RobotManip completes tasks on the fly with no pre-defined task list, demonstrating open-ended instruction following and generalization.

Qwen-RobotManip is validated across real-robot platforms and tasks, demonstrating generalization to novel scenes, unseen language instructions, and cross-embodiment transfer.

If the videos do not render in your browser, open the direct links: Omni Demo 1, Omni Demo 2, Real Robot 1, Real Robot 2, and Real Robot 3.

💡 Introduction

Qwen-RobotManip is a generalizable vision-language-action foundation model built upon Qwen-VL / Qwen3.5-4B. It couples a vision-language backbone with a flow-matching Diffusion Transformer action expert, enabling continuous action generation while preserving the perception and language grounding needed for robotic manipulation.

The central principle is alignment before scale. Robot manipulation data is naturally heterogeneous: robot embodiments, action spaces, camera systems, coordinate frames, collection pipelines, and task distributions vary widely. Qwen-RobotManip introduces a unified alig

Description

Official Repo for Qwen-RobotManip

Related repositories

Similar projects matched by category, topics, and programming language.

MoonshotAI
Featured
MoonshotAI GitHub avatar

Kimi-K3

Kimi K3 is an open-weight, 2.8T-parameter native multimodal agentic model with a 1M-token context window, designed for frontier coding, knowledge work, and reasoning tasks.

AI & Machine LearningAI Agents
3,348
xuchonglang
Featured
xuchonglang GitHub avatar

investing-for-beginners

A structured investing guide for Chinese beginners covering US stocks, options, and cryptocurrency, with focus on foundational concepts and risk awareness.

Blockchain & Web3
2,739
Krishnagangwal
Featured
Krishnagangwal GitHub avatar

CS-Fundamentals

A curated collection of Computer Science fundamentals (PDFs, notes, cheatsheets, interview question banks) for placement preparation, covering seven core subjects plus general resources.

Data & DatabasesDatabases & Storage
2,326