esp32-ai
A 28.9 million parameter language model runs on an $8 ESP32-S3 microcontroller entirely on-device, generating simple stories at about 9.5 tokens per second.
Online KL Shampoo (OKLS) is a zero-staleness Kronecker-factored optimizer that approximates full-matrix AdaGrad for large-scale language model training, achieving 1.45× parameter efficiency over Muon with comparable throughput.
26
No data
2
0
No data
2026-07-28
It combines KL-optimal Kronecker preconditioning with a novel scaled Chebyshev inverse-root method and zero-staleness covariance updates, demonstrating significant parameter efficiency gains in scaling experiments.
Similar projects matched by category, topics, and programming language.
A 28.9 million parameter language model runs on an $8 ESP32-S3 microcontroller entirely on-device, generating simple stories at about 9.5 tokens per second.
Deltafin is a research project that runs the 2.8-trillion-parameter Mixture-of-Experts model Kimi K3 on a single Apple Silicon Mac (e.g., M1 Max with 64 GB) at about 16 seconds per token, using exact, reproducible inference with local or streaming expert loading.

A comprehensive guide for building and configuring a high-end local machine to run state-of-the-art LLMs, with detailed hardware choices, BIOS tuning, and Docker-based model serving.