WUMBOLABS / RECORD
Mellum2 Agent Backend Test
JetBrains Mellum2 Instruct and Thinking Q4_K_M tested through LLMGauge on WumboJetsII.
FIT TEST / 2026-06-17
JetBrains Mellum2 12B-A2.5B Instruct Q4_K_M and Thinking Q4_K_M were tested through LLMGauge on WumboJetsII.
This is a practical local agent-backend fit test, not a leaderboard ranking.
2026-06-17 / LLMGAUGE MODEL TEST
Test Context
64k Summary
Mellum2 Instruct Q4_K_M
- Run: 5/5 complete, 0 failed
- Generation: 251.0-257.2 tok/s
- Prompt eval: 1603.1-2187.6 tok/s
- Peak VRAM: 9203 MiB
- Headroom: 3024 MiB
Mellum2 Thinking Q4_K_M
- Run: 5/5 complete, 0 failed
- Generation: 254.9-259.2 tok/s
- Prompt eval: 1681.2-2274.9 tok/s
- Peak VRAM: 9203 MiB
- Headroom: 3024 MiB
Context Ladder
Both variants completed the 8k / 16k / 32k agent-backend context ladder.
Additional fit checks
The Instruct model completed 64k fake-tool and synthetic-agent-preload checks.
The Thinking model completed the same fit/performance path.
These checks confirm that both variants fit and run through the selected LLMGauge paths on WumboJetsII. They do not automatically prove that either model is safe or preferable for daily use.
Initial Read
Both Mellum2 variants look like strong fit/performance candidates for 64k local agent-backend testing on WumboJetsII.
The result is still a lab note, not a final quality verdict. Fit, speed, and artifact validity are only part of the decision. Shell safety, tool honesty, long-context retention, and practical output quality still need manual review before treating either model as a preferred daily driver.
Early qualitative read: Mellum2 Instruct is the safer candidate to prefer unless manual scoring shows otherwise. Mellum2 Thinking is also fast and stable, but it needs stricter review around shell safety, tool-use restraint, and whether its extra reasoning behavior actually improves outputs.