esp32-ai
A 28.9 million parameter language model runs on an $8 ESP32-S3 microcontroller entirely on-device, generating simple stories at about 9.5 tokens per second.
A 49-million-parameter language model that implements key architectural ideas from Moonshot AI's Kimi K3 (including Kimi Delta Attention, Gated MLA, Stable LatentMoE, and per-head Muon) and can be trained from scratch on TinyStories using a single 8GB GPU.
46
No data
2
0
No data
2026-07-28
It makes complex frontier-model components—like linear attention with delta rule, sparse MoE with quantile balancing, and block attention residuals—accessible for hands-on study and modification on consumer hardware, without requiring massive compute.
A 49M-parameter Kimi K3-inspired language model trainable on one 8GB GPU
Similar projects matched by category, topics, and programming language.
A 28.9 million parameter language model runs on an $8 ESP32-S3 microcontroller entirely on-device, generating simple stories at about 9.5 tokens per second.

A comprehensive guide for building and configuring a high-end local machine to run state-of-the-art LLMs, with detailed hardware choices, BIOS tuning, and Docker-based model serving.
Cindy is an open-source AI agent that runs locally on your machine, integrates multiple AI harnesses and models, and provides memory, skills, and automation to perform real work in your projects and apps.