Skip to content
#

llama-cpp

Here are 177 public repositories matching this topic...

Deploy Qwen3.6-35B-A3B (Q4_K_XL) + MTP speculative decoding on a single NVIDIA L4 24GB — GCP g2-standard-8 — via the official llama.cpp Docker image. Decode-optimized to ~91–99 tok/s (min ~91 chat, max ~99 math), lossless (full GPU residency + ECC-off).

  • Updated Aug 24, 2026
  • Shell

AI-assisted rescue boot USB — a Ventoy multi-boot drive on Ubuntu 24.04 with offline LLM inference (Ollama + llama.cpp), pre-installed agentic CLIs (Claude Code, Codex, OpenCode, Copilot, OpenClaw, omnius, Aider), rescue ISOs, and a one-command build wizard.

  • Updated Jul 31, 2026
  • Shell

🧠 A comprehensive toolkit for benchmarking, optimizing, and deploying local Large Language Models. Includes performance testing tools, optimized configurations for CPU/GPU/hybrid setups, and detailed guides to maximize LLM performance on your hardware.

  • Updated Mar 27, 2025
  • Shell

Add this topic to your repo

To associate your repository with the llama-cpp topic, visit your repo's landing page and select "manage topics."

Learn more