Input text from speech in any Linux window, the lean, fast and accurate way, using whisper.cpp OFFLINE. Speak with local LLMs via llama.cpp.
-
Updated
Jul 25, 2025 - Shell
Input text from speech in any Linux window, the lean, fast and accurate way, using whisper.cpp OFFLINE. Speak with local LLMs via llama.cpp.
Bare Steam Deck motherboards as diskless, PXE-booted LLM nodes, plus the GPU clock pin that keeps them fast.
Community benchmark database for running LLMs on Apple Silicon Macs
Run Hermes Agent + Claude Code locally on llama.cpp — zero API costs. A 4h / 7M-token session that would have cost $94 on Claude Opus 4.7
Deploy Qwen3.6-35B-A3B (Q4_K_XL) + MTP speculative decoding on a single NVIDIA L4 24GB — GCP g2-standard-8 — via the official llama.cpp Docker image. Decode-optimized to ~91–99 tok/s (min ~91 chat, max ~99 math), lossless (full GPU residency + ECC-off).
A shell script to automatically update or build llama.cpp with optimal GPU support. Supports cross-platform, smart architecture detection, safe updates, and easy installation. Choose fast binary downloads or custom source builds for maximum performance.
Evidence-backed structural validation of Kimi K3 UD-IQ1_M and UD-Q4_K_XL split GGUF releases using OMIV.
AI-assisted rescue boot USB — a Ventoy multi-boot drive on Ubuntu 24.04 with offline LLM inference (Ollama + llama.cpp), pre-installed agentic CLIs (Claude Code, Codex, OpenCode, Copilot, OpenClaw, omnius, Aider), rescue ISOs, and a one-command build wizard.
Run LLMs on AMD Vega8 / Vega10 APUs and GPUs (Ryzen 5700G / gfx90c and Vega 56/64): ROCm 7 gfx900 backport + Vulkan/RADV launchers, Docker images, benchmarks for llama.cpp and LM Studio
qwen3.8-flash-next on strix halo — 91g quant, 56 tok/s, 262k context, receipts included
🧠 A comprehensive toolkit for benchmarking, optimizing, and deploying local Large Language Models. Includes performance testing tools, optimized configurations for CPU/GPU/hybrid setups, and detailed guides to maximize LLM performance on your hardware.
NVIDIA DGX Spark workstation — self-hosted LLM stack (vLLM, llama.cpp, Ollama + Open WebUI) behind Traefik, with Cloudflare Tunnel + Tailscale ingress and Netdata observability.
Runpod-LLM provides ready-to-use container scripts for running large language models (LLMs) easily on RunPod.
To associate your repository with the llama-cpp topic, visit your repo's landing page and select "manage topics."