Qwen3.8-27B on 12GB: How Far Can an RTX 5070 Really Push It?
A complete Qwen3.8-27B campaign on WumboJetsII covering quant selection, context limits, templates, reasoning, speculative decoding, coding, reliability, and final readiness.
WUMBOLABS / INDEX
Lab Records are the dated evidence archive and working history for WumboLabs.
Projects describe what exists. Lab Records document what was tested, observed, changed, or learned on real hardware. Results stay tied to their hardware, runtime, model, context, suite, scoring status, and known limitations.
These records are practical evidence, not leaderboard rankings or universal model judgments.
This is a compact orientation for current WumboLabs local-model work. It is not a leaderboard, schedule, or claim that queued models have been tested.
Read the Qwen3.8-27B Lab Record. This is the current 12GB reference on WumboJetsII, not a universal best-model claim.
| Family / experiment | Status | High-level finding |
|---|---|---|
| Empero Qwen3.8 distilled 9B / 4B / 2B | Tested — complete campaign | 9B Q4_K_M strongest broad candidate from the family; 4B Q8 strongest bounded coding/efficiency candidate; 2B extremely fast but materially weaker as a primary coding/reasoning assistant. Huge context fit did not guarantee synthesis quality. Reasoning off preferred for strict interfaces. |
| Bonsai 27B Q1 | Partially tested — reliability concerns | Loaded; compression technically impressive; practical Linux/system reliability failures, hallucinated/unsupported commands, and degeneration/repetition. Not recommended as a systems/coding assistant. |
| Ternary Bonsai 27B | Blocked — runtime support | Loader/runtime failed admission. No quality claim. |
| Gemma 4 12B NVFP4 | Blocked — full-GPU memory fit | Native backend/kernel admission succeeded; complete RTX 5070 12GB full-GPU loading failed. No inference-quality verdict. |
| Gemma 4 12B GGUF | Tested — historical strong baseline | Historical performance/efficiency reference. See the existing Gemma report. |
| Mellum2 12B | Tested — historical practical baseline | Historical practical/speed reference. |
| Grug-12B | Tested — historical baseline | Historical comparison entry. |
| Qwen3 / Qwen3.6 | Tested — historical baselines | Earlier Qwen-family evidence, not the current 12GB reference. |
| SGLang / Qwen runtime research | Runtime research — blocked | Environment/toolchain failure, not a model-quality failure. |
Producer claims are not WumboLabs findings. Use language such as advertised, we tested, reproduced, not reproduced, blocked, physically untestable, or bounded result. Example: Liquid AI's advertised approximately 97% BF16-retention result for QAD remains a producer claim until independently reproduced here.
This is a living research queue, not a fixed schedule. These models have not been tested by this snapshot.
High priority
Medium / high
Revisit
Specialized
External control
Future higher-VRAM hardware
Current Qwen3.8 findings include: native MTP works; approximately 1.4–1.5x generation improvement in matched bounded tests; MTP has meaningful VRAM cost; DFlash2 Q4 could not physically coexist with the final Qwen target under strict 12GB full-GPU constraints; n-gram speculation gave little useful benefit.
Future methods to monitor may include DFlash, DFlash2, DSpark, MTP improvements, and EAGLE-like approaches where supported. A runtime exposing a flag is not evidence that the method is worth using.
A complete Qwen3.8-27B campaign on WumboJetsII covering quant selection, context limits, templates, reasoning, speculative decoding, coding, reliability, and final readiness.
Gemmable 4 12B, Gemma 4 12B QAT Q4, and Gemma 4 12B UD-Q5 tested through LLMGauge on WumboJetsII.
JetBrains Mellum2 Instruct and Thinking Q4_K_M tested through LLMGauge on WumboJetsII.
Notes on building the WumboLabs website as a public lab record.