From the creator of Redis; run LLM locally with ds4
Points and comments are a snapshot, not live.
DwarfStar 4 is a narrow C inference engine for running frontier open-weight LLMs locally on high-memory Mac, CUDA, and ROCm machines.
DwarfStar 4 (ds4) is a C inference engine for high-memory Mac, CUDA, and ROCm machines. It supports DeepSeek V4/V4.1 Flash, GLM 5.x, and Qwen3.8 Flash Next. The engine uses asymmetric 2-bit quantization on routed experts while keeping critical paths precise. It provides a CLI, HTTP APIs, and a native agent. KV cache can be saved to SSD and resumed by prompt hash. Benchmarks on M5 Max 128GB show 790.2 t/s prefill and 39.4 t/s generation at 2,048 context, and 398.5 t/s prefill and 27.6 t/s generation at 65,536 context. Supports tensor parallelism, session batching, and vision input. Uses project GGUFs, not generic GGUF files.
What commenters are saying
Commenters largely praised ds4's performance and focused approach, with one noting the quantized models beat Unsloth quants. Several discussed the choice of C over Rust, with antirez cited as preferring C for LLM-assisted coding due to the quality of C in training data. Some pointed to the project's GitHub page as a better introduction than the landing page. A commenter shared a related inference engine for Intel Xe-LP hardware. One critique noted the landing page is generic-looking despite the project's quality. Hardware costs were discussed, with a 128GB M5 Max MacBook Pro noted at 7800€.