WUMBOLABS / RECORD READ-ONLY / DOCUMENT

← Back to Lab Records

WUMBOLABS / RECORD

Mellum2 Agent Backend Test

JetBrains Mellum2 Instruct and Thinking Q4_K_M tested through LLMGauge on WumboJetsII.

FIT TEST / 2026-06-17

JetBrains Mellum2 12B-A2.5B Instruct Q4_K_M and Thinking Q4_K_M were tested through LLMGauge on WumboJetsII.

This is a practical local agent-backend fit test, not a leaderboard ranking.

2026-06-17 / LLMGAUGE MODEL TEST

Test Context

Model family JetBrains Mellum2
Variants Instruct / Thinking
Quant Q4_K_M
Suite agent-backend-v1
Hardware RTX 5070 12GB
Result 64k fit confirmed

64k Summary

Mellum2 Instruct Q4_K_M

  • Run: 5/5 complete, 0 failed
  • Generation: 251.0-257.2 tok/s
  • Prompt eval: 1603.1-2187.6 tok/s
  • Peak VRAM: 9203 MiB
  • Headroom: 3024 MiB

Mellum2 Thinking Q4_K_M

  • Run: 5/5 complete, 0 failed
  • Generation: 254.9-259.2 tok/s
  • Prompt eval: 1681.2-2274.9 tok/s
  • Peak VRAM: 9203 MiB
  • Headroom: 3024 MiB

Context Ladder

Both variants completed the 8k / 16k / 32k agent-backend context ladder.

Instruct ladder 8192 / 16384 / 32768
Thinking ladder 8192 / 16384 / 32768
Additional fit checks

The Instruct model completed 64k fake-tool and synthetic-agent-preload checks.

The Thinking model completed the same fit/performance path.

These checks confirm that both variants fit and run through the selected LLMGauge paths on WumboJetsII. They do not automatically prove that either model is safe or preferable for daily use.

Initial Read

Both Mellum2 variants look like strong fit/performance candidates for 64k local agent-backend testing on WumboJetsII.

The result is still a lab note, not a final quality verdict. Fit, speed, and artifact validity are only part of the decision. Shell safety, tool honesty, long-context retention, and practical output quality still need manual review before treating either model as a preferred daily driver.

Early qualitative read: Mellum2 Instruct is the safer candidate to prefer unless manual scoring shows otherwise. Mellum2 Thinking is also fast and stable, but it needs stricter review around shell safety, tool-use restraint, and whether its extra reasoning behavior actually improves outputs.