🦙 Llama Server GUI
The featherweight, high-performance native desktop cockpit and flag configurator for llama-server and llama.cpp. Zero terminal friction. Pure silicon inference speed.
Everything You Need to Run Local Models
Eliminate confusing CLI arguments, batching flags, and CUDA errors.
⚡ Auto Hardware Tuning
Instantly detects NVIDIA CUDA GPUs, offloads all layers (-ngl 99), and enables Flash Attention (-fa on) automatically for peak tokens/sec.
🔍 Local Model Scanner
Automatically crawls your Downloads, Documents, and HuggingFace directories for .gguf model files. 1-click select and launch.
💬 API Chat Playground
Test and converse with your local models right inside the GUI with real-time streaming tokens, time-to-first-token (TTFT), and latency performance metrics.
🔌 OpenAI Compatible API
Serves a drop-in local endpoint (http://127.0.0.1:8080/v1) compatible with Open WebUI, AnythingLLM, Continue.dev, Cursor, and custom code.
🪶 Ultra Lightweight (2.8 MB)
Built with Tauri v2 and Rust. Uses negligible RAM and disk space, completely ditching bloated Electron runtimes for pure native desktop speed.
🎨 Retro Noir & OLED Themes
Styled with amber terminal accents, OLED midnight blacks, and responsive shortcuts for power users (Ctrl+K, F1-F3).
Llama Server GUI v2.0.0
Windows 10 / 11 (x64) • NSIS Clean Installer • 2.8 MB • Open Source under Apache 2.0