Llama-server 16GB VRAM Configs
3 minutes
This page documents llama-server presets for one system. The system has an RTX 5060 Ti with 16GB of VRAM and 96GB of DDR5 memory. The display output runs on the iGPU. This leaves the full 16GB of VRAM free for llama-server.
Set two environment variables for the best results. Set GGML_CUDA_DISABLE_GRAPHS=1 if spurious out-of-memory errors occur. Set GGML_CUDA_ENABLE_UNIFIED_MEMORY=1 to allow a small VRAM overshoot. This keeps the speed the same. The system then reports the full VRAM size as RAM in use. Other running programs may be affected.
Each preset is an ini section for llama-server. Copy the section into your models-preset.ini. Replace the filenames with your local paths.
Coding and Long Context
Qwen3.8-27B with a pruned ASCII vocabulary. The prune keeps 129,006 of 248,320 vocabulary rows. The freed VRAM becomes KV cache. The preset sets a 200K-token context window.
This preset needs troed’s fork of llama.cpp . The fork is based on Raymond’s adaptive-KV work . The fork keeps the full KV history in pinned system RAM. It holds only a bounded working set on the GPU. A bounded working set is what fits a 200K context in 16GB. A PCIe 5.0 link keeps the streaming fast.
The model uses Multi-Token Prediction (MTP) for speculative decoding. The MTP head sits inside the model file. It drafts up to five tokens per step. The target model checks them in one pass. This preset needs no separate draft file. A DFlash2 draft file exists for standard llama.cpp. This preset does not use it.
[Qwen3.8-27B]
spec-type = draft-mtp
spec-draft-n-max = 5
spec-draft-p-min = 0.8
device-draft = CUDA0
n-gpu-layers-draft = all
shared-device-memory-mib = 3904
log-file = probe-server.log
chat-template-file = chat_template_qwen3.8.jinja
# Created from https://huggingface.co/byteshape/Qwen3.8-27B-GGUF
# using the tools at https://github.com/bsaleh03/ASCII-Condensed-prune-tools
m = Qwen3.8-27B-ASCII-Condensed-IQ4_XS-3.84bpw.gguf
ctx-size = 200192
n-gpu-layers = 99
batch-size = 256
ubatch-size = 256
cache-type-k = q8_0
cache-type-v = q4_0
fit = off
parallel = 1
temp = 1.0
top-p = 0.95
top-k = 20
min-p = 0.0
presence-penalty = 0.0
repeat-penalty = 1.0
reasoning = on
reasoning-preserve = on
mmproj = Qwen3.8-mmproj-BF16.gguf
# mmap slows down PP
load-mode = none
flash-attn = on
Model: Qwen3.8-27B-ASCII-Condensed-IQ4_XS-3.84bpw.gguf
mmproj: mmproj-bf16.gguf
Prune tools: bsaleh03/ASCII-Condensed-prune-tools
Engine: troed/llama.cpp-adaptive-kv-streaming
Performance: PP ~500-1000 t/s, TG ~25-50 t/s on the 5060 Ti.
Prose
Gemma 4 26B A4B is a Google Mixture-of-Experts model. It has 25.2B total and 3.8B active parameters. It has 30 layers and a 256K context. The preset targets prose, not code.
The model is 23GB at Q6_K. It does not fit in 16GB of VRAM. The preset runs on standard llama.cpp. n-cpu-moe = 21 moves the MoE expert weights of the first 21 layers to the CPU. Attention stays on the GPU. The 96GB of RAM holds the offloaded experts.
The preset runs two speculative methods at once. MTP (Multi-Token Prediction) reads a separate draft file. It predicts the next few tokens; the target checks them. ngram-mod needs no draft model. It keeps a small hash pool of n-grams from the context. It proposes continuations that already appear in the text. The two methods run independently.
[gemma-4-26B-A4B]
m = gemma-4-26B-A4B-it-UD-Q6_K_XL.gguf
mm = gemma4-26-mmproj-BF16.gguf
md = mtp-gemma-4-26B-A4B-it.gguf
ctx-size = 140000
fit = 0
np = 1
load-mode = none
ub = 2048
n-cpu-moe = 21
spec-type = draft-mtp,ngram-mod
spec-draft-type-k = q8_0
spec-draft-type-v = q8_0
spec-draft-n-max = 4
chat-template-file = chat_template_gemma4_26b.jinja
reasoning = on
reasoning-preserve = on
temp = 1.0
top-p = 0.95
top-k = 64
repeat-penalty = 1.0
jinja = on
Model: gemma-4-26B-A4B-it-UD-Q6_K_XL.gguf
MTP draft: mtp-gemma-4-26B-A4B-it.gguf
mmproj: mmproj-BF16.gguf
Engine: standard llama.cpp
Performance: PP ~2000 t/s, drops to ~1000 as the context fills.