WUMBOLABS / EVALUATIONS CANONICAL MODEL EVALUATION

/evaluations/apodex-1.1-mini

Apodex 1.1 mini

Apodex 1.1 mini — the current WumboLabs evidence state on one page: tested profiles, validated context, and the full chronological testing history. Each value is attributed to the profile and event that measured it.

BOUNDED EARLY STOP / PHASE 3 DO NOT ADVANCE / latest evidence 2026-08-25

WumboLabs tests Apodex 1.1 mini on real consumer hardware. This is the canonical model page: current state first, then every tested profile and every evidence event. Values are attributed to the profile and event that measured them; historical findings remain the evidence of their tested stack and are never silently replaced.

Current state

  • Classification: BOUNDED_EARLY_STOP / PHASE_3_DO_NOT_ADVANCE — Initial evaluation (bounded early stop) (2026-08-25), profile llama.cpp IQ1_M (community conversion)
  • Recommended profile: llama.cpp IQ1_M (community conversion) (apodex-1.1-mini-llamacpp-iq1m, current) — canonical evidence
  • Latest evidence: 2026-08-25 — Initial evaluation (bounded early stop)

Tested profiles

llama.cpp IQ1_M (community conversion) — CURRENT

Profile identity: apodex-1.1-mini-llamacpp-iq1m.

Runtime and artifact identity are described inside the event sections below (hand-authored records; no machine-readable export).

Status: current canonical/recommended tested surface.

Canonical Evidence Profile Metadata

Events on this profile:

Testing history

Newest first. Each event is one immutable testing/publication event; the exact scientific report lives in the canonical evidence repository linked at the top of each event.

2026-08-25 — Initial evaluation (bounded early stop)

Initial Evaluation — Early Stop · profile: llama.cpp IQ1_M (community conversion) · maturity: EARLY_STOP · status: BOUNDED_EARLY_STOP / PHASE_3_DO_NOT_ADVANCE

Canonical Evidence / Full Report View on GitHub

Identity

FieldValue
ModelApodex 1.1 mini
ProducerApodex; base: Qwen/Qwen3.5-35B-A3B; ~36B total / A3B active MoE
Evaluated artifactCommunity conversion abenzerps/Apodex-1.1-mini-GGUF @ 59afa57852525f79a9634d5f80dd639cceee572c (not official)
FileApodex-1.1-mini-IQ1_M.gguf
SHA-2561e84d8adf7837e96fb18712882a8a114becc7e53554372e2f612c5e0c6276cd4
Evaluation date2026-08-25
ProtocolWELP end-to-end (frozen snapshot welp-next-snapshot-2026-08-25-end-to-end)
Highest phase reachedPhase 3 (gate decision)
Campaign classificationBOUNDED_EARLY_STOP / PHASE_3_DO_NOT_ADVANCE

Hardware

FieldValue
MachineWumboJetsII
GPUNVIDIA GeForce RTX 5070 12GB (SM120)
VRAM12,227 MiB physical
CPUAMD Ryzen 7 9800X3D
OSFedora Linux 44
Driver / CUDANVIDIA 610.57.04
Runtimellama.cpp pinned at f280b26983ad0fdb705a0d9ebf0503e76f2899b0 (2026-08-24)

Headline Verdict

DO_NOT_ADVANCE at Phase 3 (early stop, frozen gate).

The model failed the applicable advancement gate. Per WELP protocol, later phases were not run. This is a bounded result, not a complete capability review.

The early-stop behavior is a feature of WELP transparency: the model failed the gate, so the evaluation stopped. This is not a protocol failure — it is the protocol working as designed.

Performance

MetricResult
Decode throughput169.36 tok/s (median 169.46, sigma 1.70)
Prefill (~5.6K tok)1,798.76 tok/s
TTFT (short prompt)~125 ms
VRAM idle9.7 GB

Phase 3 Practical Viability — DO_NOT_ADVANCE

Contract WELP Practical Viability 0.1.2-draft, scorer score_pv.py (self-test 35/35), 30 tasks x seeds {42,43,44}, thinking OFF baseline.

GateThresholds42s43s44Verdict
G1 aggregate>=0.750.6670.7000.733FAIL (all)
G2 false-premise>=0.600.4000.2000.600FAIL (42,43)
G3 factual-uncertainty>=0.500.5000.7500.500PASS (boundary)
G4 structured-output>=0.700.7140.7140.714PASS

Decision identical across seeds => stable FAIL. Bounded failure review: ~19 genuine failures (fabricated package description, explained nonexistent git flag, asserted 25 prime, invented phone number, echoed un-reversed word); ~8 scorer lexicon artifacts. Verdict robust: correcting every artifact still leaves seed-42 aggregate <0.75.

What Was NOT Tested

  • Native Tools module
  • Coding module
  • Reasoning module (dedicated)
  • Context / Variance / Optimization / Stability
  • Soak
  • OMP Stage 2 (agent)

Do not fabricate missing results. This evaluation establishes that the strongest representation satisfying the frozen WumboJetsII full-GPU baseline (IQ1_M, 1.75-bpw) does not clear the frozen viability gate. It does NOT establish that BF16/FP8/GPTQ or Agent Team deployments would fail.

Role Classification

Supported roles: none beyond "runs coherently on 12GB at extreme quantization."

NOT_REACHED: Phases 4–10, OMP Stage 2, all capability modules, context/variance/optimization/soak/classification engine.

Important Limitations

  • Extreme quantization forced by hardware: single 12GB consumer GPU required IQ1_M (1.75-bpw). Producer positioning rests on BF16-class deployments plus an agent harness that was not authorized to run.
  • Phase-3 failure at IQ1_M does NOT establish that BF16/FP8/GPTQ or Agent Team deployments would fail. It establishes that the strongest representation satisfying the frozen full-GPU baseline does not clear the frozen viability gate.
  • Scorer lexicon defect understates false-premise performance modestly; verdict unaffected.
  • No reliability, coding, tools, context, or soak results exist.
  • Canonical evaluation repo: https://github.com/WumboLabs/evaluations/tree/553a68d915b9ea1c9c9b3be6fa65c16f13527c24/models/apodex-1.1-mini/profiles/apodex-1.1-mini-llamacpp-iq1m
  • WELP protocol: https://github.com/WumboLabs/welp
  • Labs catalog: https://github.com/WumboLabs/evaluations/tree/553a68d915b9ea1c9c9b3be6fa65c16f13527c24/models
  • LocalMaxxing: Speed result SUBMITTED/APPROVED (ID cmt9ijytg00xali017f46xk25, 182.31 tok/s p512/n128). Benchmark suites NOT_SUBMITTED due to early WELP stop.

Reproduction

Direct link to canonical reproduction material: Central evidence

This Lab Record is a summary; the canonical repo is the source of truth.

Canonical evidence

All canonical public evidence lives in WumboLabs/evaluations. Each event links an immutable full-commit/path citation; each profile remains a distinct scientific identity, not a separate repository.

Legacy provenance