deer-workflow
An open-source Dynamic Workflow runtime that combines deterministic TypeScript orchestration with replaceable Agent runtimes.
CosyVoice3Pro is a production-ready voice cloning service built on NVIDIA Triton Inference Server and TensorRT-LLM, featuring a decoupled speaker registry that lets you register a voice once and then generate speech with just a speaker ID and text.
19
No data
0
0
Apache-2.0
2026-07-29
It provides a complete, high-performance serving stack for voice cloning with real-time benchmarks (e.g., 0.148 RTF on A100), a developer-friendly HTTP API, a web management console, and audio post-processing, all while decoupling audio feature extraction from inference to reduce per-request overhead.
Production-ready CosyVoice serving with Triton, reusable Speaker Registry, Public HTTP API and Web console.
Similar projects matched by category, topics, and programming language.
An open-source Dynamic Workflow runtime that combines deterministic TypeScript orchestration with replaceable Agent runtimes.
A 28.9 million parameter language model runs on an $8 ESP32-S3 microcontroller entirely on-device, generating simple stories at about 9.5 tokens per second.

A comprehensive guide for building and configuring a high-end local machine to run state-of-the-art LLMs, with detailed hardware choices, BIOS tuning, and Docker-based model serving.