WUMBOLABS / LABS READ-ONLY / EVIDENCE INDEX

WUMBOLABS / LABS

Labs

WumboLabs model evaluation lab records. Human-readable summaries of canonical GitHub evidence from real-hardware testing.

WumboLabs Evaluation Lab Records

WumboCore Lab Records are human-readable summaries of canonical evidence published in the WumboLabs/eval-* GitHub repositories. They are not independent evaluation artifacts.

Canonical evidence: GitHub eval repositories Human-readable summary: this page Standardized benchmarks: LocalMaxxing (where recorded)

Each record stays bounded by the tested artifact, runtime, hardware, configuration, and protocol snapshot. Results are not universal model rankings.


WELP — WumboLabs Evaluation Lifecycle Protocol

WELP (WumboLabs Evaluation Lifecycle Protocol) is the reproducible, phase-gated evaluation lifecycle used for WumboLabs model testing. It is published at https://github.com/WumboLabs/welp.

What it is: a fixed testing protocol that runs a model through ordered phases — provenance, admission, performance, practical viability, reliability, capability modules, context, variance, optimization, and stability. Each phase has a deterministic gate.

Why phase gates exist: the protocol is frozen before a model is evaluated. A failed gate is a valid result. This prevents post-hoc threshold tuning and makes early-stop behavior transparent.

Why early-stop results are still valuable: a model that stops at Phase 3 still produces bounded evidence about admission, performance, and practical viability. The absence of later-phase data is itself a finding, not a gap to hide.

Why campaign depth differs: different models reach different WELP depths. Qwen3.8-27B completed a deep end-to-end campaign. Nemotron 3 Nano 4B stopped at a protocol-defined gate. Apodex 1.1 mini failed a viability gate and was not advanced. These differences are features of the protocol, not inconsistencies in effort.

Where the canonical specification lives: https://github.com/WumboLabs/welp (DRAFT — not yet frozen as v1.0).


Current Lab Records

The table below summarizes each evaluation. Read individual records for evidence boundaries, hardware, and detailed findings.

Model Status Hardware Headline Repo
Qwen3.8-27B COMPLETED_DEEP_EVALUATION WumboJetsII (RTX 5070 12GB) Strong reviewed local coding/technical assistant. Not recommended as unguarded daily driver or unattended autonomous agent. GitHub
Nemotron 3 Nano 4B PARTIAL_PROTOCOL_DEVELOPMENT WumboJetsII (RTX 5070 12GB) Protocol development campaign. Phases 0–2 passed; stopped at Phase 3 gate decision. Not a complete WELP capability review. GitHub
Apodex 1.1 mini BOUNDED_EARLY_STOP / PHASE_3_DO_NOT_ADVANCE WumboJetsII (RTX 5070 12GB) Early stop at Phase 3 gate. Model failed the applicable advancement gate; later phases were not run. Bounded result only. GitHub
LFM2.5 2.6B COMPLETE_SMALL_MODEL_EVALUATION WumboJetsII (RTX 5070 12GB) QAD Q4_0 is the rational ultra-fast deployment quant when 1.59 GB is the hard model-size target. No material practical-quality win over PTQ Q4_0 was reproduced. GitHub
LFM2.5 1.2B COMPLETE_SMALL_MODEL_EVALUATION WumboJetsII (RTX 5070 12GB) Best always-resident micro-sidecar candidate in the family. Hallucination resistance absent at 1.2B scale — every quant including BF16 fabricates confidently on invented subjects. GitHub
Ornith 1.5 9B COMPLETE / FINAL_WELP_READINESS_NOT_READY WumboJetsII (RTX 5070 12GB) Full WELP campaign completed. Phase 3 practical viability ADVANCE (0.922); Phase 4 reliability gate FAIL (DO_NOT_ADVANCE) closes all deployment roles: FINAL WELP READINESS: NOT_READY. Later-phase evidence is characterization only. GitHub

Lab Records

COMPLETED_DEEP_EVALUATION / 2026-08-21

Qwen3.8-27B Lab Record

Complete WumboLabs evaluation of Qwen3.8-27B on WumboJetsII (RTX 5070 12GB). Strong reviewed local coding/technical assistant. Not recommended as unguarded daily driver.

PARTIAL_PROTOCOL_DEVELOPMENT / 2026-08-25

Nemotron 3 Nano 4B Lab Record

WumboLabs evaluation of NVIDIA Nemotron 3 Nano 4B on WumboJetsII (RTX 5070 12GB). Partial protocol development campaign — not a complete WELP capability review.

BOUNDED_EARLY_STOP / PHASE_3_DO_NOT_ADVANCE / 2026-08-25

Apodex 1.1 mini Lab Record

WumboLabs evaluation of Apodex 1.1 mini on WumboJetsII (RTX 5070 12GB). Bounded early stop at Phase 3 — DO NOT ADVANCE.

COMPLETE_SMALL_MODEL_EVALUATION / 2026-08-26

LFM2.5 2.6B QAD Lab Record

WumboLabs evaluation of Liquid AI LFM2.5 2.6B QAD Q4_0 on WumboJetsII (RTX 5070 12GB). QAD preserves Q4_0 speed/memory but no material practical-quality win over PTQ was reproduced.

COMPLETE_SMALL_MODEL_EVALUATION / 2026-08-26

LFM2.5 1.2B QAD Lab Record

WumboLabs evaluation of Liquid AI LFM2.5 1.2B QAD Q4_0 on WumboJetsII (RTX 5070 12GB). Best always-resident micro-sidecar candidate in the family, but hallucination resistance is absent at this scale.

COMPLETE / FINAL_WELP_READINESS_NOT_READY / 2026-08-27

Ornith 1.5 9B Lab Record

WumboLabs full-lifecycle WELP evaluation of Ornith 1.5 9B (Q4_K_M) on WumboJetsII (RTX 5070 12GB). Complete campaign; mechanical outcome NOT_READY — Phase 4 reliability gate failure closes all deployment roles.