
local-llm
A comprehensive guide for building and configuring a high-end local machine to run state-of-the-art LLMs, with detailed hardware choices, BIOS tuning, and Docker-based model serving.

A recipe to serve the full-quality, unpruned GLM-5.2-Int4-Int8Mix checkpoint (256 experts) on a 4× NVIDIA DGX Spark cluster, achieving 36 tok/s peak single-stream and 200K context with MTP speculative decoding and CUDA graphs.
72
No data
8
2
Apache-2.0
2026-07-25
Delivers near-lossless inference of a large Mixture-of-Experts model on modest hardware (4× GB10) with detailed, battle-tested setup instructions and open-source contributions back to the community.
Recipe: GLM-5.2 (unpruned QuantTrio Int4-Int8Mix) at 200K ctx with MTP spec decode on a 4x NVIDIA DGX Spark (GB10) cluster
Similar projects matched by category, topics, and programming language.

A comprehensive guide for building and configuring a high-end local machine to run state-of-the-art LLMs, with detailed hardware choices, BIOS tuning, and Docker-based model serving.
Claudemux is a lightweight, dependency-minimal message bus that lets multiple Claude Code sessions communicate in real-time over tmux, enabling one session to ask another a question and receive the answer without manual intervention.
A set of bash scripts using ffmpeg to simulate vintage cassette tape audio profiles with noise, wow/flutter, bandwidth limits, and EQ.