LingBot-Video 是首个面向具身智能的开源大规模混合专家(MoE)视频生成模型,具备高效的推理能力并在机器人相关基准测试上达到领先水平。

Stars

878

7 天增长

暂无数据

Fork 数

39

开放 Issue

9

开源协议

Apache-2.0

最近更新

2026-07-10

AI 仓库情报摘要
FR-AI / ANALYSIS

为什么值得关注

它是首个专门为具身智能设计的开源 MoE 视频生成模型,在 RBench 排行榜上排名第一,并且完全以 Apache 2.0 许可证开源。

适合谁使用

  • 具身智能与机器人领域的研究人员
  • 视频生成模型的开发者
  • 机器人仿真与训练的从业者
  • 开源人工智能爱好者

典型使用场景

  • 生成具有物理合理性的机器人任务规划与仿真视频
  • 制作高质量的文字转视频和图像条件视频
  • 在具身基准(操作、导航等)上评估视频合成模型
  • 开发自定义提示重写和自动负提示的视频生成流水线

项目优势

  • 高效的 MoE 架构,推理速度比同等容量稠密模型快约 3 倍
  • 使用超过 70,000 小时的具身数据与海量网络视频训练,具备强大的物理理解能力
  • 在 RBench 排行榜上多个具身类别中取得开源模型最佳平均得分(0.620)
  • 提供完整的推理工作流:提示重写、自动负提示、多 GPU 支持(FSDP、CP8、SGLang)

使用前须知

  • 需要特定的运行时版本(Python ≥3.10、自定义 PyTorch 构建)和环境配置
  • 30B-A3B 的 MoE 模型加载时需要大量 GPU 显存和系统内存(即使使用 FSDP)
  • 推理期望结构化 JSON 描述而非自然语言提示,必须配合重写器流水线使用

README 快速开始

LingBot-Video

🌐 Project Page | 🤗 Hugging Face | 🤖 ModelScope | 📄 Paper | ⚖️ License| 💬 WeChat 微信 Group

📘 English Usage: English Documentation
📕 中文使用文档: 中文文档

We are excited to introduce LingBot-Video, the first open-source large-scale MoE (Mixture-of-Experts) video generation model dedicated to embodied intelligence. As a top-tier video model, LingBot-Video is designed to bridge the gap between video synthesis and physical world understanding.

🔥 Key Highlights

  • 🚀 Efficient MoE Architecture: Scaled from scratch; balanced between capacity and cost with ~3x faster inference.
  • 📦 Data Engine: Trained on massive web videos integrated with 70,000+ hours of embodied data.
  • ⚖️ Multi Reward System: Rewarded for high aesthetics, physical rationality, and task completion.

🎬 Video Demos

🔥 Latest News

  • July 9, 2026: 🎉 We release the technical report, code, models, rewriters for LingBot-Video.

📦 Model Download

Model NameComponentsTasksDownload
⚡ LingBot-Video-DenseDense (1.3B)T2I, T2V, TI2V🤗 Huggingface   🤖 ModelScope
💪 LingBot-Video-MoEMoE (30B-A3B) + RefinerT2I, T2V, TI2V, Refinement🤗 Huggingface   🤖 ModelScope
📝 LingBot-Video-Rewriter-BaseQwen3.6-27B officialPrompt rewriter (Expand)🤗 Huggingface   🤖 ModelScope
📝 LingBot-Video-Rewriter-AdapterQwen3.6-27B LoRAPrompt rewriter (Json)🤗 Huggingface   🤖 ModelScope

🚀 Quick Start

🛠️ Installation

The root `requireme

项目描述

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence

相关仓库与替代方案

根据分类、Topic 和编程语言匹配的相似项目。

Vincentwei1021
Vincentwei1021 GitHub avatar

video-shotcraft

An AI agent skill that turns Claude Code or Codex into a motion-design studio for crafting cinematic product videos with Remotion, offering 106 shot recipe cards, 162 styles, 161 motion previews, and a production-ready template.

TypeScript
2,432
lopopolo
精选
lopopolo GitHub avatar

harness-engineering

Harness Engineering is a methodology for improving coding agent outputs by carefully crafting the environment around them—providing curated context, tools, and executable constraints that encode an organization’s nonfunctional requirements and cumulative lessons.

AI 与机器学习AI 智能体
2,390
slvDev
精选
slvDev GitHub avatar

esp32-ai

A 28.9 million parameter language model runs on an $8 ESP32-S3 microcontroller entirely on-device, generating simple stories at about 9.5 tokens per second.

AI 与机器学习大语言模型
1,960