Qwen-RobotManip 是一个视觉-语言-动作基础模型,通过表示对齐、运动对齐和行为对齐使异构的机器人操作数据能够统一训练,从而实现强大的分布外泛化和跨本体迁移。

Stars

116

7 天增长

暂无数据

Fork 数

3

开放 Issue

11

开源协议

暂无数据

最近更新

2026-06-29

AI 仓库情报摘要
FR-AI / ANALYSIS

为什么值得关注

该工作提出了三维对齐框架,使多种开源操作数据可协同训练,在多个 OOD 基准上达到最优,并在 RoboChallenge Table30 v1 通用赛道中排名第一,且仅使用公开数据集。

适合谁使用

  • 机器人操作基础模型研究者
  • 开发通用机器人策略的 AI 工程师
  • 从事模仿学习与迁移学习的学术实验室
  • 自动制造业或服务机器人行业的从业者

典型使用场景

  • 真实机器人上无需预定义任务列表的开放式指令跟随
  • 跨本体操作技能迁移(单臂、双臂、灵巧手、移动平台)
  • 无需显式恢复脚本的错误后主动恢复
  • 对新场景、物体和语言指令的分布外泛化

项目优势

  • 三维对齐(表示、运动、行为)统一异构多源数据
  • 80 维规范动作空间配合掩码支持多本体共享一个模型
  • 在 LIBERO-Plus、RoboTwin-Clean2Rand 和 EBench 上超越先前方法,RoboChallenge Table30 第一
  • 利用 38100 小时开源数据,其中 24808 小时为合成的示教数据

使用前须知

  • 模型权重未公开且无发布计划
  • 训练需要大规模数据和计算资源,复现门槛较高
  • 主要针对桌面操作任务,其他领域的适用性尚未验证

README 快速开始

Qwen-RobotManip

Alignment Unlocks Scale for Robotic Manipulation Foundation Models

Qwen Team

📑 Technical Report |
📖 Blog |
🖥️ Demo

Welcome to the official repository of Qwen-RobotManip. Here, you can find official information about Qwen-RobotManip and post your questions (Issues).

Note: There is currently no plan to release the model weights for Qwen-RobotManip or Qwen-RobotNav. We will continue adding report resources that can be publicly released to this repository.

🎬 Demo

Qwen-Omni × Qwen-RobotManip — Qwen-Omni observes the scene, randomly proposes manipulation tasks via speech, and judges execution in real time. The two demos below are English-subtitled versions. Qwen-RobotManip completes tasks on the fly with no pre-defined task list, demonstrating open-ended instruction following and generalization.

Qwen-RobotManip is validated across real-robot platforms and tasks, demonstrating generalization to novel scenes, unseen language instructions, and cross-embodiment transfer.

If the videos do not render in your browser, open the direct links: Omni Demo 1, Omni Demo 2, Real Robot 1, Real Robot 2, and Real Robot 3.

💡 Introduction

Qwen-RobotManip is a generalizable vision-language-action foundation model built upon Qwen-VL / Qwen3.5-4B. It couples a vision-language backbone with a flow-matching Diffusion Transformer action expert, enabling continuous action generation while preserving the perception and language grounding needed for robotic manipulation.

The central principle is alignment before scale. Robot manipulation data is naturally heterogeneous: robot embodiments, action spaces, camera systems, coordinate frames, collection pipelines, and task distributions vary widely. Qwen-RobotManip introduces a unified alig

项目描述

Official Repo for Qwen-RobotManip

相关仓库与替代方案

根据分类、Topic 和编程语言匹配的相似项目。

MoonshotAI
精选
MoonshotAI GitHub avatar

Kimi-K3

Kimi K3 is an open-weight, 2.8T-parameter native multimodal agentic model with a 1M-token context window, designed for frontier coding, knowledge work, and reasoning tasks.

AI 与机器学习AI 智能体
3,348
xuchonglang
精选
xuchonglang GitHub avatar

investing-for-beginners

A structured investing guide for Chinese beginners covering US stocks, options, and cryptocurrency, with focus on foundational concepts and risk awareness.

区块链与 Web3
2,739
Krishnagangwal
精选
Krishnagangwal GitHub avatar

CS-Fundamentals

A curated collection of Computer Science fundamentals (PDFs, notes, cheatsheets, interview question banks) for placement preparation, covering seven core subjects plus general resources.

数据与数据库数据库与存储
2,326