LingBot-VLA 2.0 is a Vision-Language-Action foundation model that improves real-world robot generalization through 60,000 hours of pre-training data, a unified 55-dimensional action space, MoE action experts, and dual-query distillation.

Stars

666

7-day growth

No data

Forks

50

Open issues

10

License

Apache-2.0

Last updated

2026-07-25

AI repository intelligence
FR-AI / ANALYSIS

Why it is worth attention

It bridges large-scale pre-training to practical robot deployment, significantly outperforms prior models (e.g., π0, GR00T N1.7) on both bimanual and mobile manipulation benchmarks, and releases open-source weights and training code.

Who it is for

  • Robotics researchers studying VLA models
  • AI engineers deploying manipulation policies on real robots
  • Developers working on cross-embodiment generalization
  • Simulation-to-real transfer practitioners

Use cases

  • Bimanual manipulation on platforms like AgileX Cobot Magic and Galaxea R1Pro
  • Long-horizon mobile manipulation (e.g., refrigerator sorting, stove cleaning)
  • Post-training on custom datasets (e.g., RoboTwin 2.0) for new tasks
  • Real-time deployment on NVIDIA RTX 4090D with ~130ms inference

Strengths

  • 60,000 hours of diverse pre-training data covering 20 robot configurations and egocentric human video
  • Unified 55-dimensional action representation supporting arms, grippers, dexterous hands, waist, head, and mobile base
  • MoE action expert with fine-grained segmentation for better cross-embodiment scaling
  • Achieves state-of-the-art results on GM-100 bimanual and RoboTwin 2.0 benchmarks

Considerations

  • Requires specific dependencies (Python 3.12, PyTorch 2.8.0, flash-attn) and careful environment setup
  • Post-training demands custom data preparation in LeRobot format with robot config files
  • Real-robot deployment may need additional hardware tuning (GPU memory, communication load)

README quick start

Installation

Description

From Foundation to Application

Related repositories

Similar projects matched by category, topics, and programming language.

lopopolo
Featured
lopopolo GitHub avatar

harness-engineering

Harness Engineering is a methodology for improving coding agent outputs by carefully crafting the environment around them—providing curated context, tools, and executable constraints that encode an organization’s nonfunctional requirements and cumulative lessons.

AI & Machine LearningAI Agents
2,390
slvDev
Featured
slvDev GitHub avatar

esp32-ai

A 28.9 million parameter language model runs on an $8 ESP32-S3 microcontroller entirely on-device, generating simple stories at about 9.5 tokens per second.

AI & Machine LearningLarge Language Models
1,960
littledivy
Featured
littledivy GitHub avatar

mimic

mimic captures traffic from any iOS or web app and automatically generates a Python client library that lets you call the app's API like a regular library.

AI & Machine Learning
1,482