Running Qwen 3.8 27B on a 16 GB GPU
A dense model has to fit entirely in VRAM, since every weight is read for every token and spilling to system RAM collapses throughput. Before this I ran Qwen3.6-35B-A3B, a Mixture-of-Experts (MoE) model that uses only a small part of its weights for each token, so llama.cpp could keep some of them in system RAM for a small speed hit. Qwen3.8-27B is dense. Unsloth publishes quantized versions of open models, and its 4-bit build (UD-IQ4_XS) gets the 27.8B parameters down to 13.27 GB, leaving about 3 GB for the KV cache (the attention state that grows with context) and llama.cpp’s buffers.
That’s tight for 64K of context, but most of Qwen3.8’s layers don’t keep a KV cache (the model card has the details), so 64K fits with the cache quantized to 4-bit (q4_0). The whole thing uses 15769 MiB of the card’s 16311 MiB.
It decodes at about 24.9 tokens/second, and on a 5000-token prompt it processes 888 tokens/second, then decodes at 21.9. Serious work still goes to a hosted frontier model, but the local one has moved past redacting and cleaning text, and with the GitHub MCP tools wired in it found a real bug in one of my repos. The Docker Compose file below runs it on an RTX 5060 Ti with llama.cpp (why not Ollama) serving it and Open WebUI as the chat front end, along with the settings that tripped me up.
The compose file
It needs llama.cpp b10450 or later. Older builds load the model and then produce fluent garbage on CUDA. llama-server reads each --flag from an LLAMA_ARG_FLAG environment variable, so most settings live under environment:.
services:
llama-server:
image: ghcr.io/ggml-org/llama.cpp:server-cuda-v0.5.0
restart: unless-stopped
# Five-minute idle unload. This one has no LLAMA_ARG_ variable, so it only
# works as an argument.
command: --sleep-idle-seconds 300
environment:
# Pin downloads to the mounted volume, or a container recreate re-fetches 13 GB.
- LLAMA_CACHE=/root/.cache/llama.cpp
- LLAMA_ARG_HF_REPO=unsloth/Qwen3.8-27B-GGUF:UD-IQ4_XS
- LLAMA_ARG_N_GPU_LAYERS=99 # all layers on GPU; dense, so no --n-cpu-moe to fall back on
- LLAMA_ARG_CTX_SIZE=65536
- LLAMA_ARG_N_PARALLEL=1 # single user, one KV cache slot
- LLAMA_ARG_FLASH_ATTN=1 # big KV cache VRAM savings
# q4_0 rather than q8_0: resident dense weights leave less room than the MoE,
# which kept ~2 GB of its weights in system RAM. Qwen tolerates it well.
- LLAMA_ARG_CACHE_TYPE_K=q4_0
- LLAMA_ARG_CACHE_TYPE_V=q4_0
- LLAMA_ARG_MMPROJ_AUTO=0 # it's a vision-language model; skip the ~0.93 GB vision projector
- LLAMA_ARG_BATCH=2048 # not LLAMA_ARG_BATCH_SIZE, which nothing reads
- LLAMA_ARG_UBATCH=512
- LLAMA_ARG_TEMPERATURE=0.7 # Qwen's sampling for thinking-off mode
- LLAMA_ARG_TOP_K=20
- LLAMA_ARG_TOP_P=0.8
- LLAMA_ARG_MIN_P=0
- LLAMA_ARG_JINJA=1 # correct chat template and tool calling
- LLAMA_ARG_REASONING=off # thinking off, see below
- LLAMA_ARG_PORT=11434
- LLAMA_ARG_HOST=0.0.0.0
volumes:
- ./models:/root/.cache/llama.cpp
ports:
- "11434:11434"
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities: [gpu]
open-webui:
image: ghcr.io/open-webui/open-webui:v0.11.4
restart: unless-stopped
ports:
- "3000:8080"
environment:
- ENABLE_OLLAMA_API=false
- OPENAI_API_BASE_URLS=http://llama-server:11434/v1
- OPENAI_API_KEYS=no-key
- ENABLE_CONTEXT_COMPACTION=true
# Under CTX_SIZE, or llama.cpp refuses the request before compaction fires.
- CONTEXT_COMPACTION_TOKEN_THRESHOLD=48000
- CONTEXT_COMPACTION_RETENTION_PERCENTAGE=40 # share of recent messages kept verbatim
volumes:
- ./open-webui:/app/backend/data
depends_on:
- llama-server
Qwen also recommends presence_penalty=1.5 for thinking-off mode to curb repetition. That’s LLAMA_ARG_PRESENCE_PENALTY=1.5, or Open WebUI’s per-model Advanced Params.
The thinking switch
Qwen 3.8 defaults to reasoning effort xhigh and will spend tens of thousands of tokens on a trivial question. --reasoning off (LLAMA_ARG_REASONING=off above) turns thinking off for the whole server. For a single API request, "reasoning_effort": "none" in the body does the same, since llama-server turns it into thinking off before the chat template sees it.
Don’t reach for the server flag --reasoning-effort none, even though Unsloth’s model page lists none as a level. llama.cpp passes that flag’s value straight into the model’s chat template, and every request fails:
Error: Jinja Exception: Unexpected reasoning effort none.
Supported types are xhigh (default), medium, and low.
Context compaction
64K runs out in agent sessions, because every request carries the schema for every enabled tool plus everything already read:
request (71461 tokens) exceeds the available context size (65536 tokens)
Open WebUI’s context compaction summarises earlier turns once a chat passes a token threshold. It’s off by default, and the default threshold of 80000 sits above the 64K served here, so the compose sets both. Expect it to fire before the threshold, since the backend double-counts cached tokens.
Leave CONTEXT_COMPACTION_MODEL unset. Pointed at a cloud model, it sends the earlier turns of every local conversation out to be summarised, which defeats the point of running locally. TASK_MODEL, which Open WebUI uses for side jobs like chat titles, stays unset for the same reason.