WUMBOLABS / EVALUATIONS CANONICAL MODEL EVALUATION

/evaluations/mellum2-12b-a2.5b

Mellum2 12B-A2.5B

Mellum2 12B-A2.5B — the current WumboLabs evidence state on one page: tested profiles, validated context, and the full chronological testing history. Each value is attributed to the profile and event that measured it.

LIMITED ROLE ONLY / latest evidence 2026-09-16

WumboLabs tests Mellum2 12B-A2.5B on real consumer hardware. This is the canonical model page: current state first, then every tested profile and every evidence event. Values are attributed to the profile and event that measured them; historical findings remain the evidence of their tested stack and are never silently replaced.

Current state

Tested profiles

llama.cpp Q4_K_M (Instruct) — CURRENT

Profile identity: mellum2-12b-a2.5b-llamacpp-q4km-instruct.

FieldValue
Runtimellama.cpp b9672 (74ade5274), CUDA SM120
ArtifactMellum2-12B-A2.5B-Instruct-Q4_K_M.gguf; reacquired from official repo after local+archive absence; SHA-256 verified against official LFS
PrecisionQ4_K_M

Status: current canonical/recommended tested surface.

Canonical Evidence Profile Metadata

Events on this profile:

llama.cpp Q4_K_M (Thinking) — CURRENT-ALTERNATE

Profile identity: mellum2-12b-a2.5b-llamacpp-q4km-thinking.

FieldValue
Runtimellama.cpp b9672 (74ade5274), CUDA SM120
ArtifactMellum2-12B-A2.5B-Thinking-Q4_K_M.gguf; locally reacquired and SHA-256 verified against the pinned official LFS object
PrecisionQ4_K_M

Status: validated alternate tested surface.

Canonical Evidence Profile Metadata

Events on this profile:

Testing history

Newest first. Each event is one immutable testing/publication event; the exact scientific report lives in the canonical evidence repository linked at the top of each event.

2026-09-16 — Current-WELP recharacterization

WELP Recharacterization — llama.cpp Q4_K_M (Thinking) · profile: llama.cpp Q4_K_M (Thinking) · maturity: CURRENT_WELP · status: NOT_READY

Canonical Evidence / Full Report View on GitHub

Identity
FieldValue
ModelMellum2 12B-A2.5B (Thinking)
ProducerJetBrains
Official modelJetBrains/Mellum2-12B-A2.5B-Thinking @ 71a489e7b95efacf89feaaa6fe3b2995f3542409 (tested GGUF repo JetBrains/Mellum2-12B-A2.5B-Thinking-GGUF-Q4_K_M)
Tested artifactMellum2-12B-A2.5B-Thinking-Q4_K_M.gguf; locally reacquired and SHA-256 verified against the pinned official LFS object
PrecisionQ4_K_M
Artifact SHA-256489cf0d7ca86ef4683e34e2efe8a46a9b52573bfe994a6ad9c86bf57c7173ccb
Campaignmellum2-12b-a2.5b-thinking-rtx5070-welp-recharacterization-2026-09-16
Record date2026-09-16
Runtime and hardware
FieldValue
Enginellama.cpp
Runtime versionb9672 (74ade5274), CUDA SM120
HardwareWumboJetsII (NVIDIA GeForce RTX 5070 12GB)
Hardware notesfull GPU residency; f16 KV; FA; single slot; one heavy CUDA workload at a time
WELP outcome
  • Outcome: PASS — RECHARACTERIZED (Thinking profile)
  • Classification: NOT_READY
  • Artifact classification: current

Publication state: published — canonical evidence: https://github.com/WumboLabs/evaluations/blob/a0ed8d8218d000c40fb2be157f4377d73c67af45/models/mellum2-12b-a2.5b/events/mellum2-12b-a25b-thinking-rtx5070-welp-recharacterization-2026-09-16/REPORT.md

Context profile
FieldValue
Practical defaultnot recorded tokens
Guarded contextnot recorded tokens
Native model-card maximum131072 tokens
Model-card envelope completeYES
Native maximum dispositionVALIDATED capacity: exact native 131,072 at 99.660–99.758% near-full occupancy, two reps, cache-free; useful-context gates fail under frozen 512-token reserve.
Headline performance
SurfaceTTFTPrefillDecode
Short0.043 s—267.9 tok/s
Moderate (4438-token input)0.574 s7830.6 tok/s245.4 tok/s

cache-free (--no-cache-prompt --cache-ram 0, cached_tokens 0 verified); short decode 267.8-268.3 tok/s across 5 reps, moderate 245.4-246.9; llama-bench canonical pp512 7,710.53 ± 18.61 / tg128 270.75 ± 0.66 tok/s.

Quality and capabilities
  • Constrained result: 11/12 frozen mechanical quality screen; Q12 empty at fixed cap
  • reasoning: explicit-CoT syllogism PASS
  • coding: moving_sum executable oracle PASS
  • native tools: exact call/arguments and grounded continuation PASS
Guardrails and limitations
  • Reliability 3/20 for both required seeds; every request reaches the fixed completion limit in reasoning.
  • No useful-context maximum validates under the frozen 512-token completion reserve; do not infer usefulness from 131,072 capacity.
  • NOT_READY for a current-WELP agent role (human-accepted 2026-09-16); narrow capability successes are not role qualification.

Reliability: Frozen 20-task corpus, official sampler, seeds 42 and 314159: 3/20 each; strict interfaces 3/3 only; all 40 requests length-truncated.

LocalMaxxing
FieldValue
StatusMEASURED_NOT_SUBMITTED
Canonical contextnot recorded tokens
tok/s out270.75
TTFTnot recorded
Submission referencenot recorded
verifiedRunnull (not claimed)

Local canonical llama-bench benchmark (pp512 7,710.53 ± 18.61 / tg128 270.75 ± 0.66 tok/s, evidence runs/llama-bench-canonical.txt); no LocalMaxxing service contact or submission occurred.

Canonical evidence

Canonical public evidence: https://github.com/WumboLabs/evaluations/blob/a0ed8d8218d000c40fb2be157f4377d73c67af45/models/mellum2-12b-a2.5b/events/mellum2-12b-a25b-thinking-rtx5070-welp-recharacterization-2026-09-16/REPORT.md

This event section is a human-readable derivative of the accepted local WELP campaign evidence named above; the campaign's REPORT.md is the authoritative scientific source. Results are bounded by the tested artifact, runtime, hardware, configuration, and protocol snapshot, and are not universal model rankings.

2026-09-15 — Current-WELP recharacterization

WELP Recharacterization — llama.cpp Q4_K_M (Instruct) · profile: llama.cpp Q4_K_M (Instruct) · maturity: CURRENT_WELP · status: LIMITED_ROLE_ONLY

Canonical Evidence / Full Report View on GitHub

Identity
FieldValue
ModelMellum2 12B-A2.5B
ProducerJetBrains
Official modelJetBrains/Mellum2-12B-A2.5B-Instruct @ 1236b4166ed6ab1d57e4be9bcc19f4899c190cbf (tested GGUF repo)
Tested artifactMellum2-12B-A2.5B-Instruct-Q4_K_M.gguf; reacquired from official repo after local+archive absence; SHA-256 verified against official LFS
PrecisionQ4_K_M
Artifact SHA-256b04281c27de5d968d577f310d982273b1b13bdbd8117b3ecffffeebfe222f0a7
Campaignmellum2-12b-a2.5b-rtx5070-welp-recharacterization-2026-09-15
Record date2026-09-15
Runtime and hardware
FieldValue
Enginellama.cpp
Runtime versionb9672 (74ade5274), CUDA SM120
HardwareWumboJetsII (NVIDIA GeForce RTX 5070 12GB)
Hardware notesfull GPU residency -ngl 99; f16 K/V with FA; single slot; one heavy CUDA workload at a time; --no-cache-prompt --cache-ram 0 for uncached measurement
WELP outcome
  • Outcome: PASS
  • Classification: LIMITED_ROLE_ONLY

Publication state: published — canonical evidence: https://github.com/WumboLabs/evaluations/blob/adad9217547bf22cea9f2bd27e79a929374a20a9/models/mellum2-12b-a2.5b/events/mellum2-12b-a25b-rtx5070-welp-recharacterization-2026-09-15/REPORT.md

Context profile
FieldValue
Practical default16384 tokens
Guarded context8192 tokens
Native model-card maximum131072 tokens
Model-card envelope completeYES
Native maximum dispositionFAILED on useful context with capacity+performance validated: exact 131,072 admits at f16 KV (9,850 MiB load, 2,377 MiB free) and passes near-full performance (99.66-99.76% occupancy, 4,872-4,883 tok/s prefill); the frozen two-seed useful-context gate failed 0/2 with a replicated recency-bias signature. Not a FIT_LIMIT and not a validation - a measured negative. Useful context is VALIDATED at 8,192 and 16,384 only.
Quality and capabilities
  • Constrained result: 11/12 frozen mechanical quality screen; sole miss was the three-word lowercase instruction-retention task.
  • coding: PASS (moving_sum executed correctly on all oracle cases)
  • tool formatting: PASS (native hermes tool_call, exact arguments, grounded continuation)
  • reasoning (direct mode, REASONING_OFF baseline): PASS (multi-hop syllogism correct, no think-tag leakage on the no-CoT checkpoint)
Guardrails and limitations
  • reliability 7/20 at seed 42 and 8/20 at seed 314159; hallucination category 0/4 and 1/4 - REPLICATED fabrication of documentation for nonexistent APIs (fake CUDA signature, fake requests function, fake git command, nonexistent commit summary)
  • sycophancy 0/3 and uncertainty 0/3 at both seeds; 7/20 outputs truncated at the frozen tight token caps
  • useful-context ceiling 16,384 on this stack despite the 131,072 advertised window
  • keep context <=16,384 (guarded 8,192); use the separate official Thinking checkpoint for explicit-CoT roles; that profile's current-WELP characterization is still open

Reliability: Strict interfaces 3/3 both seeds; evidence discipline 2/3 both; Git safety 1/1 both. The replicated fabrication and zero-pass sycophancy/uncertainty categories are classification-determining (LIMITED_ROLE_ONLY).

LocalMaxxing
FieldValue
StatusMEASURED_NOT_SUBMITTED
Canonical contextnot recorded tokens
tok/s outnot recorded
TTFTnot recorded
Submission referencenot recorded
verifiedRunnull (not claimed)

local canonical-profile benchmark pp512 7,744.16 ± 15.01 tok/s / tg128 271.66 ± 0.57 tok/s (5 reps); no external submission or fabricated verification fields; service mutation requires separate authorization.

Canonical evidence

Canonical public evidence: https://github.com/WumboLabs/evaluations/blob/adad9217547bf22cea9f2bd27e79a929374a20a9/models/mellum2-12b-a2.5b/events/mellum2-12b-a25b-rtx5070-welp-recharacterization-2026-09-15/REPORT.md

This event section is a human-readable derivative of the accepted local WELP campaign evidence named above; the campaign's REPORT.md is the authoritative scientific source. Results are bounded by the tested artifact, runtime, hardware, configuration, and protocol snapshot, and are not universal model rankings.

2026-07-04 — LocalMaxxing LMX speed runs (Instruct + Thinking)

Benchmark Only — LMX local speed · profile: llama.cpp Q4_K_M (Instruct) · maturity: BENCHMARK_ONLY · status: BENCHMARK_ONLY (LMX local speed)

Canonical Evidence / Full Report View on GitHub

Identity and scope

  • Profile: mellum2-12b-a2.5b-llamacpp-q4km-instruct — llama.cpp Q4_K_M (Instruct)
  • Evidence maturity: BENCHMARK_ONLY
  • Evidence scope: performance
  • Hardware: WumboJetsII (RTX 5070 12GB)

LMX local speed evidence: Instruct 261.41 tok/s out; Thinking 265.47 tok/s out (llama.cpp, Q4_K_M). Measured locally; submission disposition covered by the 2026-09-10 backfill scope for canonical profiles only.

This event is a bounded public-safe summary derived from retained WumboLabs evidence. It is not a WELP characterization, a universal model ranking, or a production-readiness proof. Values are attributed to the tested artifact, runtime, hardware, suite, and settings.

2026-06-17 — Fake-tool honesty runs (64k)

Specialized Test — fake-tool honesty · profile: llama.cpp Q4_K_M (Instruct) · maturity: SPECIALIZED_TEST · status: SPECIALIZED_TEST / UNSCORED_PROBES

Canonical Evidence / Full Report View on GitHub

Identity and scope

  • Profile: mellum2-12b-a2.5b-llamacpp-q4km-instruct — llama.cpp Q4_K_M (Instruct)
  • Evidence maturity: SPECIALIZED_TEST
  • Evidence scope: specialized, reliability
  • Hardware: WumboJetsII (RTX 5070 12GB)

Fake-tool honesty probes at 64k for both Mellum2 variants; feeding the agent-backend manual scores.

This event is a bounded public-safe summary derived from retained WumboLabs evidence. It is not a WELP characterization, a universal model ranking, or a production-readiness proof. Values are attributed to the tested artifact, runtime, hardware, suite, and settings.

2026-06-17 — Agent backend fit test (64k, Instruct + Thinking)

Specialized Test — LLMGauge agent-backend-v1 · profile: llama.cpp Q4_K_M (Instruct) · maturity: SPECIALIZED_TEST · status: SPECIALIZED_TEST / AGENT_BACKEND_FIT_TEST (not a general quality verdict) · related profiles: mellum2-12b-a2.5b-llamacpp-q4km-thinking

Canonical Evidence / Full Report View on GitHub Profile Metadata: mellum2-12b-a2.5b-llamacpp-q4km-thinking

Long-form report: Mellum2 Agent Backend Test

Identity and scope

  • Profile: mellum2-12b-a2.5b-llamacpp-q4km-instruct — llama.cpp Q4_K_M (Instruct)
  • Evidence maturity: SPECIALIZED_TEST
  • Evidence scope: specialized, agent-backend
  • Hardware: WumboJetsII (RTX 5070 12GB)

Instruct + Thinking Q4_K_M through LLMGauge agent-backend-v1 at 64k on WumboJetsII. 64k fit confirmed (Instruct 5/5 complete, 251.0-257.2 tok/s out, 9203 MiB peak VRAM). Manual scores: Instruct preferred (overall trust 3.7/5) over Thinking; neither safe for unsupervised shell/systemd operations. Not a general model-quality verdict.

This event is a bounded public-safe summary derived from retained WumboLabs evidence. It is not a WELP characterization, a universal model ranking, or a production-readiness proof. Values are attributed to the tested artifact, runtime, hardware, suite, and settings.

Related tested profile: mellum2-12b-a2.5b-llamacpp-q4km-thinking — both variants were tested in the same event.

Shared comparison events

This model appears in shared multi-model comparisons. Each report is stored once in WumboLabs/evaluations/shared-events/; the model's measured entries remain attributed to the same historical event:

Canonical evidence

All canonical public evidence lives in WumboLabs/evaluations. Each event links an immutable full-commit/path citation; each profile remains a distinct scientific identity, not a separate repository.

Legacy provenance