WUMBOLABS / RECORD
12B Gemma-Based Practical Use Test
Gemmable 4 12B, Gemma 4 12B QAT Q4, and Gemma 4 12B UD-Q5 tested through LLMGauge on WumboJetsII.
REPORT / 2026-06-21
Gemmable 4 12B MTP Q4_K_M, Gemma 4 12B IT QAT UD-Q4_K_XL, and Gemma 4 12B IT UD-Q5_K_XL were tested through LLMGauge on WumboJetsII.
This is a practical local usefulness test, not a leaderboard ranking.
2026-06-21 / LLMGAUGE PRACTICAL USE TEST
Test Context
Practical Use Summary
Gemma 4 12B IT QAT UD-Q4_K_XL
- Score: 259.2 / 300.0
- Average: 4.32 / 5.00
- Generation: 73.45 tok/s avg
- Prompt eval: 1867.27 tok/s avg
- Peak VRAM: 7539 MiB
- Headroom: 4688 MiB
- Verdict: best practical balance
Gemma 4 12B IT UD-Q5_K_XL
- Score: 254.0 / 300.0
- Average: 4.23 / 5.00
- Generation: 59.15 tok/s avg
- Prompt eval: 1522.53 tok/s avg
- Peak VRAM: 9341 MiB
- Headroom: 2886 MiB
- Verdict: strong, but heavier
Mia-AiLab Gemmable 4 12B MTP Q4_K_M
- Score: 119.8 / 300.0
- Average: 2.00 / 5.00
- Generation: 69.10 tok/s avg
- Prompt eval: 1675.03 tok/s avg
- Peak VRAM: 8173 MiB
- Headroom: 4054 MiB
- Verdict: structurally viable, practically mixed
What Was Tested
The WumboLabs Practical Use Suite tested normal local-assistant usefulness across Linux, coding, Docker, honesty, summarization, and local LLM advice.
Scored prompt results
Linux update guidance
- Gemmable Q4_K_M: 1.50
- Gemma QAT Q4: 4.44
- Gemma UD-Q5: 4.39
Python log parser
- Gemmable Q4_K_M: 1.50
- Gemma QAT Q4: 4.35
- Gemma UD-Q5: 4.39
Docker Compose review
- Gemmable Q4_K_M: 1.50
- Gemma QAT Q4: 4.19
- Gemma UD-Q5: 3.40
Unknown package honesty
- Gemmable Q4_K_M: 1.50
- Gemma QAT Q4: 4.17
- Gemma UD-Q5: 4.30
Technical run summary
- Gemmable Q4_K_M: 4.48
- Gemma QAT Q4: 4.79
- Gemma UD-Q5: 4.79
12GB local LLM advice
- Gemmable Q4_K_M: 1.50
- Gemma QAT Q4: 3.98
- Gemma UD-Q5: 4.13
The scores are manual local-context judgments. They are review aids, not universal model rankings.
Runtime settings
| Field | Value |
|---|---|
| Backend | llama.cpp |
| Context | 8192 |
| Max tokens | 1200 |
| Temperature | 0.2 |
| Top-p | 0.95 |
| Batch | 256 |
| UBatch | 64 |
| GPU layers | 999 |
All three models completed the six-prompt suite with zero runtime failures.
Initial Read
Gemma 4 12B IT QAT UD-Q4_K_XL was the strongest practical-use result in this comparison. It produced useful answers across the suite while also using the least VRAM and delivering the fastest average generation speed.
Gemma 4 12B IT UD-Q5_K_XL was also strong, but it used substantially more VRAM and ran slower. In this run, the heavier quant did not clearly justify its extra 12GB hardware cost.
The result also showed that quantization weight alone did not determine practical value: the heavier UD-Q5 was slower, used more VRAM, and scored slightly lower in this suite.
Mia-AiLab Gemmable 4 12B MTP Q4_K_M fit well and produced a good summarization answer, but most practical prompts drifted into action-oriented or incomplete responses instead of directly answering from the provided context. It remains interesting for prompt-template or MTP runtime investigation, but it was not the best practical-use model in this comparison.
The no-hype takeaway: for WumboJetsII / RTX 5070 12GB practical local use, Gemma 4 12B IT QAT UD-Q4_K_XL is currently the strongest 12B Gemma-based result from this test.