Desktop Native Tauri v2 + Rust Production Ready v2.0.0

🦙 Llama Server GUI

The featherweight, high-performance native desktop cockpit and flag configurator for llama-server and llama.cpp. Zero terminal friction. Pure silicon inference speed.

⬇️ Download for Windows (v2.0.0) View on GitHub →
Llama Server GUI Cockpit Interface
// Engineered For Local AI

Everything You Need to Run Local Models

Eliminate confusing CLI arguments, batching flags, and CUDA errors.

⚡ Auto Hardware Tuning

Instantly detects NVIDIA CUDA GPUs, offloads all layers (-ngl 99), and enables Flash Attention (-fa on) automatically for peak tokens/sec.

🔍 Local Model Scanner

Automatically crawls your Downloads, Documents, and HuggingFace directories for .gguf model files. 1-click select and launch.

💬 API Chat Playground

Test and converse with your local models right inside the GUI with real-time streaming tokens, time-to-first-token (TTFT), and latency performance metrics.

🔌 OpenAI Compatible API

Serves a drop-in local endpoint (http://127.0.0.1:8080/v1) compatible with Open WebUI, AnythingLLM, Continue.dev, Cursor, and custom code.

🪶 Ultra Lightweight (2.8 MB)

Built with Tauri v2 and Rust. Uses negligible RAM and disk space, completely ditching bloated Electron runtimes for pure native desktop speed.

🎨 Retro Noir & OLED Themes

Styled with amber terminal accents, OLED midnight blacks, and responsive shortcuts for power users (Ctrl+K, F1-F3).

Standard Windows Installer

Llama Server GUI v2.0.0

Windows 10 / 11 (x64) • NSIS Clean Installer • 2.8 MB • Open Source under Apache 2.0

⬇️ Download Installer (.exe)