WUMBOLABS / EVALUATIONS CANONICAL MODEL EVALUATION

/evaluations/minicpm5-2b

MiniCPM5-2B

MiniCPM5-2B — the current WumboLabs evidence state on one page: tested profiles, validated context, and the full chronological testing history. Each value is attributed to the profile and event that measured it.

READY WITH GUARDRAILS / latest evidence 2026-09-10

WumboLabs tests MiniCPM5-2B on real consumer hardware. This is the canonical model page: current state first, then every tested profile and every evidence event. Values are attributed to the profile and event that measured them; historical findings remain the evidence of their tested stack and are never silently replaced.

Current state

Tested profiles

contained vLLM BF16 — CURRENT

Profile identity: minicpm5-2b-vllm-bf16.

FieldValue
RuntimevLLM 0.27.1 (g6e448d0ea), Torch 2.13.0+cu130, Transformers 5.15.0, FlashInfer 0.6.16.post3
Artifactmodel-00000-of-00001.safetensors (official BF16; no quantization or conversion)
PrecisionBF16 weights and BF16 KV (primary surface); fp8(e4m3) KV only on the single authorized 131,072 alternate surface

Status: current canonical/recommended tested surface.

Canonical Evidence Profile Metadata

Events on this profile:

Testing history

Newest first. Each event is one immutable testing/publication event; the exact scientific report lives in the canonical evidence repository linked at the top of each event.

2026-09-10 — Initial evaluation (full characterization)

Initial Evaluation — contained vLLM BF16 · profile: contained vLLM BF16 · maturity: FULL_EVALUATION · status: READY_WITH_GUARDRAILS

Canonical Evidence / Full Report View on GitHub

Identity
FieldValue
ModelMiniCPM5-2B
ProducerOpenBMB
Official modelopenbmb/MiniCPM5-2B @ cd199ce3ee67549c42ef7372f809f2c63599a3e9
Tested artifactmodel-00000-of-00001.safetensors (official BF16; no quantization or conversion)
PrecisionBF16 weights and BF16 KV (primary surface); fp8(e4m3) KV only on the single authorized 131,072 alternate surface
Artifact SHA-25614fb8e7f0a18d53d1f239773758bf581cee7e456a4523a54622c3a245b64402c
Campaignminicpm5-2b-rtx5070-welp-characterization-2026-09-10
Record date2026-09-10
Runtime and hardware
FieldValue
EnginevLLM
Runtime version0.27.1 (g6e448d0ea), Torch 2.13.0+cu130, Transformers 5.15.0, FlashInfer 0.6.16.post3
Runtime notesQualified contained CUDA 13.0 toolchain (compute_120f/sm_120f); full transformer GPU residency; fresh-cache containment proven; greedy temperature-0 primary surface
HardwareWumboJetsII (NVIDIA GeForce RTX 5070 12GB)
Hardware notesSingle-user workstation; AMD Ryzen 7 9800X3D; Fedora Linux 44
WELP outcome
  • Outcome: PASS — MINICPM5_2B_RTX5070_CHARACTERIZED; MODEL-CARD CONTEXT ENVELOPE COMPLETE = YES
  • Classification: READY_WITH_GUARDRAILS
  • Artifact classification: Official BF16 full-precision candidate (characterized; not deployed); architecture-diversity control value HIGH

Publication state: published — canonical evidence: https://github.com/WumboLabs/evaluations/blob/479cb50c197f8ea0dd905a68ffd61f632f49c652/models/minicpm5-2b/events/initial-evaluation-2026-09-10-minicpm5/REPORT.md

Context profile
FieldValue
Practical default32768 tokens
Guarded context65536 tokens
Native model-card maximum131072 tokens
Model-card envelope completeYES
Native maximum dispositionExact 131,072 carries two completed dispositions: FIT_LIMIT on the BF16-KV surface (1,024 MiB reserve floor caps the pool below the maximum) and measured FAILED (strict useful-context gate) on the authorized fp8-KV alternate surface executed at 99.50% occupancy; the highest admitted BF16-KV rung was 98,304
Headline performance
SurfaceTTFTPrefillDecode
Short0.013037 s—118.622 tok/s
Moderate (2663-token input)0.203597 s13079.8 tok/s116.95 tok/s

Medians over five measured repetitions per surface at 8,192 baseline context; near-full ladder medians of two seeds per rung retained in campaign evidence

Quality and capabilities
  • Constrained result: 6/7 constrained checks — one deterministic temperature-0 over-refusal of a benign prime-list prompt failed the repeat-check content scorer; repeat consistency itself held; no repair attempted (frozen contract)
  • Thinking/reasoning TESTED_PASS; coding TESTED_PASS (4/4 frozen execution cases); native tool selection and tool-result use TESTED_PASS
  • English/Chinese text, instruction following, and structured output TESTED_PASS
  • Near-full performance inside practical gates through 98,304 (BF16-KV) with zero errors across 55 scientific requests
Guardrails and limitations
  • Strict long-context output compliance fails at all rungs (format + absent-field), even though content-level retrieval is strong
  • Deterministic temperature-0 over-refusal quirk observed on a benign prompt
  • fp8-KV scales uncalibrated (kv_scale 1.0); agent benchmarks deferred as out of scope; not autonomous-agent qualification

Reliability: 20/20 COMPLETE and passed on a fresh default-context server; zero HTTP/runtime/CUDA/OOM/Xid errors; bounded, not endurance certification

LocalMaxxing
FieldValue
StatusSUBMITTED
Canonical context32768 tokens
tok/s out118.2
TTFT22.7 ms
Submission referencecmtwcna5907f0ps01sfiwd60a
verifiedRunNO

APPROVED, origin NEW (localmaxxing-backfill-2026-09-10); benchmarked on the canonical practical stack (BF16, vLLM contained runtime, 32K default — not the 131K fp8 boundary); actual prompt tokens 252 (endpoint usage); verifiedRun false reflects a client capture limitation

Canonical evidence

Canonical public evidence: https://github.com/WumboLabs/evaluations/blob/479cb50c197f8ea0dd905a68ffd61f632f49c652/models/minicpm5-2b/events/initial-evaluation-2026-09-10-minicpm5/REPORT.md

This event section is a human-readable derivative of the accepted local WELP campaign evidence named above; the campaign's REPORT.md is the authoritative scientific source. Results are bounded by the tested artifact, runtime, hardware, configuration, and protocol snapshot, and are not universal model rankings.

Canonical evidence

All canonical public evidence lives in WumboLabs/evaluations. Each event links an immutable full-commit/path citation; each profile remains a distinct scientific identity, not a separate repository.

Legacy provenance