INTACT-JEPA
INTACT is an end-to-end JEPA world model that learns an isomorphic intent-to-action interface, enabling search-free direct control for goal-conditioned robot tasks with high success rates after just one epoch of training.

LingBot-Video is the first open-source large-scale Mixture-of-Experts (MoE) video generation model designed for embodied intelligence, offering efficient inference and state-of-the-art performance on robotics-related benchmarks.
878
No data
39
9
Apache-2.0
2026-07-10
It is the first open-source MoE video generation model specifically targeting embodied intelligence, achieving top rankings on the RBench leaderboard while being fully open-sourced under Apache 2.0.
🌐 Project Page | 🤗 Hugging Face | 🤖 ModelScope | 📄 Paper | ⚖️ License| 💬 WeChat 微信 Group
📘 English Usage: English Documentation
📕 中文使用文档: 中文文档
We are excited to introduce LingBot-Video, the first open-source large-scale MoE (Mixture-of-Experts) video generation model dedicated to embodied intelligence. As a top-tier video model, LingBot-Video is designed to bridge the gap between video synthesis and physical world understanding.
| Model Name | Components | Tasks | Download |
|---|---|---|---|
| ⚡ LingBot-Video-Dense | Dense (1.3B) | T2I, T2V, TI2V | 🤗 Huggingface 🤖 ModelScope |
| 💪 LingBot-Video-MoE | MoE (30B-A3B) + Refiner | T2I, T2V, TI2V, Refinement | 🤗 Huggingface 🤖 ModelScope |
| 📝 LingBot-Video-Rewriter-Base | Qwen3.6-27B official | Prompt rewriter (Expand) | 🤗 Huggingface 🤖 ModelScope |
| 📝 LingBot-Video-Rewriter-Adapter | Qwen3.6-27B LoRA | Prompt rewriter (Json) | 🤗 Huggingface 🤖 ModelScope |
The root `requireme
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
Similar projects matched by category, topics, and programming language.
INTACT is an end-to-end JEPA world model that learns an isomorphic intent-to-action interface, enabling search-free direct control for goal-conditioned robot tasks with high success rates after just one epoch of training.
Harness Engineering is a methodology for improving coding agent outputs by carefully crafting the environment around them—providing curated context, tools, and executable constraints that encode an organization’s nonfunctional requirements and cumulative lessons.
A 28.9 million parameter language model runs on an $8 ESP32-S3 microcontroller entirely on-device, generating simple stories at about 9.5 tokens per second.