DS4 · LOCAL FRONTIER INFERENCE Run frontier open weights locally with ds4. Dwarf Star 4 is a narrow C inference engine for high-memory Mac, CUDA and ROCm machines. It supports Deep Seek V4 and V4.1 Flash, GLM 5.x and Qwen3.8 Flash Next, with text and vision models, local APIs, a CLI and a native agent in one stack.
SUPPORTED: DEEPSEEK V4 / V4.1 + GLM 5.x + QWEN3.8 · MIT LICENSE· C / MET AL / CUDA / ROCM· QWEN ON 64GB PRINCIPLE. The 284-billion-parameter Deep Seek V4 Flash is a large mixture-of-experts model. ds4 starts from the opposite constraint of remote serving. Asymmetric quantization targets the routed experts while preserving critical paths, making the model practical on high-memory machines.
The local engine exposes a CLI, HTTP APIs and a native agent, all sharing the same model state and cache. ds4 is not a generic GGUF runner; it follows a small, opportunistic set of model families and validates each supported layout end to end. Core features include asymmetric 2-bit quantization, KV cache as a disk citizen, and one engine with three interfaces: ./ds4 for chat, ./ds4-server for local APIs and ./ds4-agent for persistent coding sessions.
Run ds4 in three steps: download the project GGUF, build for your backend, then start the CLI or server. Hardware fit includes platforms like M5 Max, 128 GB achieving 34.4 T/S generation and 557 T/S prefill. ds4-server speaks OpenAI and Anthropic-style APIs, enabling local coding agents to connect to your own machine.
Source: Hacker News · Summarized by HeadlinesBriefing