| name | minicpm5-deploy-llama-cpp |
|---|---|
| description | Run MiniCPM5-1B or MiniCPM5-2B with llama.cpp using the released GGUF artifacts (F16 / Q8_0 / Q4_K_M). Use when the user wants CPU-only / consumer-GPU / cross-platform native deployment, asks for "llama.cpp", "llama-cli", "llama-server", "GGUF", or has no Python available. |
CPU / edge / consumer-GPU deployment via the released GGUF artifacts. The artifacts work directly with vanilla llama.cpp and every downstream runtime (Ollama / LM Studio / llama-cpp-python).
| Var | Example | Default |
|---|---|---|
GGUF_REPO |
openbmb/MiniCPM5-2B-GGUF |
required; openbmb/MiniCPM5-1B-GGUF also works |
QUANT |
Q4_K_M (1.56 GB, recommended) / Q8_0 (2.68 GB) / F16 (5.04 GB) |
Q4_K_M |
NGL |
99 (all layers on GPU) / 0 (CPU only) |
99 if NVIDIA GPU, else 0 |
CTX |
8192 (default) up to 131072 (128 K) |
8192 |
# macOS
brew install llama.cpp
# Linux / cross-platform: pre-built binary
curl -fsSL https://github.com/ggerganov/llama.cpp/releases/latest/download/llama-cli-linux.tar.gz | tar -xz
# OR build from source:
git clone --depth=1 https://github.com/ggerganov/llama.cpp.git && cd llama.cpp
mkdir build && cd build
cmake .. -DGGML_CUDA=ON -DCMAKE_BUILD_TYPE=Release # CPU-only: omit GGML_CUDA=ON
cmake --build . --config Release -j $(nproc) --target llama-cli llama-servermkdir -p ~/minicpm5 && cd ~/minicpm5
huggingface-cli download ${GGUF_REPO} MiniCPM5-2B-${QUANT}.gguf --local-dir .llama-cli -m MiniCPM5-2B-${QUANT}.gguf \
-n 2048 --temp 1.0 --top-p 0.95 --min-p 0.0 -ngl ${NGL} -c ${CTX}llama-server -m MiniCPM5-2B-${QUANT}.gguf \
--port 8080 -ngl ${NGL} -c ${CTX} --jinjacurl http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "MiniCPM5-2B",
"messages": [{"role":"user","content":"1+1=?"}],
"temperature": 1.0, "top_p": 0.95, "min_p": 0.0, "max_tokens": 64
}'Expected: "2" in the reply.
In llama.cpp, the default min_p=0.05 can lead to repetitive output: it filters out tokens whose probability is below 5% of the highest-probability token, potentially discarding the exact tokens needed to break out of a repetition loop. To prevent this, we set min_p=0.0.
| Mode | --temp |
--top-p |
--min-p |
|---|---|---|---|
| MiniCPM5-2B Think | 1.0 | 0.95 | 0.0 |
| MiniCPM5-1B Think | 0.9 | 0.95 | 0.0 |
| MiniCPM5-1B No-think | 0.7 | 0.95 | 0.0 |
| Quant | Disk | Quality |
|---|---|---|
| F16 | 5.04 GB | reference |
| Q8_0 | 2.68 GB | ~indistinguishable from F16 |
| Q4_K_M | 1.56 GB | small drop, ideal for laptops |
- Slow on CPU + large context: drop
-c 131072to-c 8192if you don't need 128 K.
If you've trained your own MiniCPM5-1B or MiniCPM5-2B variant, build a GGUF with:
python convert_hf_to_gguf.py /path/to/your-fp16-hf --outfile out/F16.gguf --outtype f16
llama-quantize out/F16.gguf out/Q4_K_M.gguf Q4_K_MTrained a LoRA adapter (not a full model) and want to apply it at runtime with --lora instead of baking it in? Convert it to a GGUF adapter — see minicpm5-finetune-gguf-lora.
- NVIDIA GPU + want OpenAI-compatible serving →
minicpm5-deploy-vllm - Apple Silicon native →
minicpm5-deploy-mlxis faster - Just want one-line desktop run →
minicpm5-deploy-ollama - Want a desktop GUI →
minicpm5-deploy-lmstudio