MiniCPM-Robot is a family of compact vision-language-action models for generalist robot manipulation and on-device target tracking, achieving strong performance with 1.5B and 0.9B parameters.

Stars

262

7-day growth

+60

Forks

19

Open issues

2

License

Apache-2.0

Last updated

2026-07-20

AI repository intelligence
FR-AI / ANALYSIS

Why it is worth attention

It introduces the first fully on-device embodied target tracker (0.9B), runs 5+ FPS on a Unitree Go2, and surpasses far larger models (e.g., π₀.₅, Qwen-VLA) on manipulation benchmarks while being open-source.

Who it is for

  • Robotics researchers building generalist manipulation policies
  • Developers deploying AI on edge/embedded devices
  • Hobbyists with Unitree Go2 or similar legged robots
  • Embodied AI practitioners seeking efficient VLAs

Use cases

  • Generalist pick-and-place and tool use with a single set of weights
  • Real-world target tracking in cluttered or adversarial environments
  • Deploying vision-only language-commanded tracking on low-power hardware
  • Streaming memory for long-horizon manipulation without recomputation overhead

Strengths

  • 1.5B generalist VLA beats 3B–5B+ models on multiple manipulation evals
  • Streaming inference reduces per-step TFLOPs from 125 to 3.3 for 60-frame history
  • First open-source on-device tracker with SOTA on EVT-Bench across static/dynamic/adversarial targets
  • Detailed one-command deployment guide for Unitree Go2 with 5+ FPS end-to-end latency

Considerations

  • Real-robot deployment currently validated only on Unitree Go2 EDU with Jetson Orin NX
  • Environment setup requires many third-party dependencies (Habitat, PyTorch, TensorRT) and manual asset downloads
  • Only two models released so far (Manipulation and Tracking); broader task coverage is future work

README quick start

Quick Start

Description

A Smarter and Faster On-Device AI Brain for Robots

Related repositories

Similar projects matched by category, topics, and programming language.

lopopolo
Featured
lopopolo GitHub avatar

harness-engineering

Harness Engineering is a methodology for improving coding agent outputs by carefully crafting the environment around them—providing curated context, tools, and executable constraints that encode an organization’s nonfunctional requirements and cumulative lessons.

AI & Machine LearningAI Agents
2,390
slvDev
Featured
slvDev GitHub avatar

esp32-ai

A 28.9 million parameter language model runs on an $8 ESP32-S3 microcontroller entirely on-device, generating simple stories at about 9.5 tokens per second.

AI & Machine LearningLarge Language Models
1,960
littledivy
Featured
littledivy GitHub avatar

mimic

mimic captures traffic from any iOS or web app and automatically generates a Python client library that lets you call the app's API like a regular library.

AI & Machine Learning
1,482