Janus is a single Go binary that runs .gguf models on your machine (GPU or CPU) and exposes an Open AI-compatible API. No Python, no Docker, no Ollama required. Use it your way: call it from the command line (curl, Power Shell, scripts), wire it into Cursor / Cline / any Open AI client — same local models, whatever workflow fits you.\n\nWhat you get includes local inference via llama.cpp through Vulkan (AMD / Intel / NVIDIA) or CPU fallback, an Open AI-compatible API with endpoints like /v1/chat/completions and /v1/models, hot-swap models without restarting, thinking model support with reasoning split, chat template auto-detection from GGUF metadata, and zero dependencies — one .exe on Windows.\n\nRequirements: Windows 10/11 with Go 1.22+ and Vulkan-capable GPU recommended, Linux with Go 1.22+ and Vulkan or CPU, or macOS with Go 1.22+ and CPU backend.
Disk space includes model size (often 2–8 GB per model) plus ~50 MB for Janus + llama.dll. Quick start involves cloning the repo, building with build.ps1, downloading a .gguf file, configuring the .env file (e.g., INFERENCE_BACKEND=vulkan, JANUS_MODEL_PATH), and running the server on 127.0.0.1:8990.\n\nThe project layout includes cmd/janus/ for the main server, cmd/modelget/ for model downloading, and internal/engine/ for the llama.cpp Vulkan runner. Pitfalls include ensuring no leftover processes hold port 8990, correct JANUS_MODEL_PATH, and updated GPU drivers.
Source: Hacker News · Summarized by HeadlinesBriefing