/methodology
Methodology
How WumboLabs produces, reviews, validates, and bounds its evidence.
WumboLabs treats test results as bounded evidence, not universal rankings. Real hardware. Real testing. No hype.
This page explains how WumboLabs evidence is produced, reviewed, validated, and bounded. Model evaluations live under Evaluations; long-form benchmarks, fit tests, baselines, and lab notes live in Reports.
01 — Evaluation philosophy
- Real hardware, named systems. Results come from real consumer hardware in daily use — currently the WumboJetsII workstation (Fedora 44, RTX 5070 12GB) — not from idealized or unnamed test farms.
- Exact configuration matters. A result without its model identity, artifact, runtime, and settings is not usable evidence. Records carry those details with the numbers.
- Failures count as evidence. A blocked admission, a failed gate, or a physical non-fit is a finding. See Failures count.
- Bounded claims. Every result is scoped to the hardware, runtime, artifact, settings, and task set used to produce it.
- Producer claims require reproduction. A vendor or quantizer claim stays a producer claim until independently reproduced here. Records use language such as advertised, tested, reproduced, not reproduced, blocked, or physically untestable to keep that distinction visible.
- Manual review is reviewer judgment. Manual scores are reviewer judgment recorded under a stated rubric. They are useful review metadata, not objective proof of model quality.
02 — WELP
WELP (WumboLabs Evaluation Lifecycle Protocol) is the reproducible, phase-gated lifecycle for WumboLabs model testing. The current 2026-09-25 model-agentic snapshot is DRAFT, not v1.0. It binds prospective gates to setup and raw-output evidence, with qualified safety, context and real-work oracles, a frozen blinded-review disagreement resolution path, a mechanical calibration sanity layer, an explicit representation for answerless budget exhaustion, first-class reasoning profiles, and the two WELP sections below. Earlier campaigns retain their own frozen snapshots and are not retroactively rescored.
WELP Model and WELP Agentic — one protocol, two sections. WELP Model asks how capable, reliable and useful the tested local model profile is (reliability, safety, assistant quality, structured output, tool recovery, Linux diagnosis, repository repair, multi-turn correction, context, reasoning profiles, performance). WELP Agentic asks whether that model, through a specified agent harness, can independently complete bounded tasks using tools and environment feedback — discover, decide, act, observe, recover, verify, finish. Each section has its own execution status, evidence, applicability and limitations, and the two verdicts are never averaged or combined. A negative Model verdict does not prohibit Agentic testing; missing infrastructure is reported as untested or integration-blocked, never as a model defect. Agentic runs pin one qualified profile under a predeclared selection rule, bind the full harness configuration (harness commit, tool adapter, agent system prompt, sampling, sandbox, frozen limits), execute model-generated actions only inside a disposable offline sandbox whose isolation is proven before any action, and score what actually happened — final environment state, executable checks, the verbatim tool transcript with exit statuses, model-performed verification, and false-completion detection — with task success, truthful reporting, safety and efficiency kept as separate dimensions. The initial suite is three tasks (a small repository repair with reviewer feedback and hidden acceptance tests, a broken-service repair in a sandboxed simulator with an authorized-repair boundary, and a frozen-corpus research question with citation binding); a 3/3 result would still be reported as 3 out of 3 tested tasks, not as a certification of general autonomy. The tested model's tool adapter is qualified before any scored task, and the runner never rescues the agent: it executes declared tool operations verbatim and never invents arguments, repairs malformed output, or claims verification the model never performed.
Reasoning profiles — models are tested as they deploy: some models reason (think) before answering, and for several of them reasoning can be genuinely switched on or off. WELP determines each model's actual reasoning behavior experimentally on the tested runtime — a requested setting alone proves nothing. When an effective, supported ON/OFF control exists, the model is characterized as two separate deployment profiles: Reasoning On and Reasoning Off. Each profile gets its own setup, calibration, generation budgets, behavioral testing, context characterization, performance measurement, safety results and an independent readiness verdict. Profile verdicts are reported side by side and are never averaged into one model-level score — the same model can legitimately be ready in one profile and not ready in the other. Models without any reasoning mode are tested as a single Standard profile; models whose reasoning cannot be disabled are not given a fake "off" profile; and extra reasoning-effort levels (low/medium/high and similar) are qualified first and only become full profiles when they are genuinely supported, measurably distinct and deployment-relevant.
Historical interpretation limits: the frozen compatibility policy records confirmed limitations in the retained LFM2.5-8B-A1B retest: task failure did not establish unsafe behavior, a context oracle rejected correct answers, and tool-discovery and unexecuted Linux coverage limited role conclusions. Original reports, scores and verdicts remain historical evidence. This protocol revision supplies neither a replacement model classification nor a new model campaign.
What it is: a preregistered protocol with ordered phases — provenance, admission, performance, practical viability, reliability, capability modules, context, variance, optimization, and stability — and frozen applicable gates. It separates campaign execution state from a model's readiness verdict.
Why phase gates exist: the protocol is frozen before a model is evaluated. A failed gate is a valid result. This prevents post-hoc threshold tuning and makes early-stop behavior transparent.
Why early-stop results are still valuable: a model that stops at Phase 3 still produces bounded evidence about admission, performance, and practical viability. The absence of later-phase data is itself a finding, not a gap to hide.
Why campaign depth differs: different models reach different WELP depths. Some campaigns complete a deep end-to-end evaluation; others stop at a protocol-defined gate or fail an early viability gate and are not advanced. These differences are features of the protocol, not inconsistencies in effort.
Useful work has multiple boundaries: the exact artifact, quantization, runtime/build, hardware, configured context, template, deployment prompt, sampler and effective reasoning mode identify a tested deployment profile. Generation ceilings are separately recorded measurement lanes of that same profile, not part of its ID. Prospective campaigns distinguish a bounded semantic ceiling, calibrated separately for each response class on disjoint tasks under the same profile before scored runs, from a deployment role's independently frozen operational limit. Each response class also freezes its expected answer geometry (declared from the fixture rubrics), and a mechanical sanity check fails setup closed before scored work when the selected ceiling cannot hold that declared answer structure or when the calibration example that represents the class's longest expected answers produced far shorter visible answers — reasoning-token consumption never fails calibration, and a genuinely concise class keeps its small ceiling. Exact rendered prompts, class representativeness review and repeated identical cache probes are setup evidence; labels alone do not qualify them. An answerless reasoning trace at a short cap is not a semantic failure; a correct answer that needs excessive time, tokens or VRAM is not evidence of operational fitness. Report completion, semantic accuracy and resource cost separately by lane. A generic deployment prompt is the primary role lane; minimal, publisher-recommended and optional preregistered optimized settings remain distinct, never pooled after seeing answers.
Context claims are bounded by occupancy: configured capacity is not context validation. Full-context claims require the actual final rendered prompt at near-full occupancy of the usable budget (≥99% preferred, ≥97% hard floor, with reserved generation tokens), full-window performance, and useful-context evidence at that occupancy. Every native and officially advertised extension range — including each exact maximum — requires an explicit tested or demonstrated limiting disposition (VALIDATED, FIT_LIMIT, INTEGRATION_BLOCKED, or a measured FAILED) before context characterization may be called complete. Coverage completeness, execution validity, demonstrated capability and the useful maximum are separate axes. A valid measured FAILED gate completes a negative coverage cell; it does not validate capability. Missing or invalid evidence does not count as a measured failure.
Answerless budget exhaustion is a measured outcome: a reasoning-heavy model can legitimately spend its entire generation reserve inside its reasoning channel and emit no visible answer. In Controlled Context such a cell is a valid, executed test — recorded as budget-limited, with the semantic result explicitly not evaluable rather than a failure. It counts as test coverage, so a complete negative result never makes a campaign structurally incomplete, but it never validates context capability: a budget-limited rung stays unvalidated and does not raise the useful-context maximum. Reasoning consumption itself is a deployment/budget limitation, not a semantic failure, an unsafe behavior, or missing coverage.
Context construction matters: the prospective Controlled Context fixture (legacy ID: Family A) places facts against the final rendered, inference-equivalent token stream at five measured depths (2/25/50/75/95%, at most 0.50 percentage-point placement error). Semantic and operational context lanes reserve their own generation ceilings. A complementary Multi-Document Context task checks distributed, versioned and absent facts. Neither family establishes general codebase or conversational long-context synthesis; those need separate tasks. Test names use plain display names — Controlled Context, Multi-Document Context, Tool Recovery, Linux Diagnosis, Repository Repair, Multi-Turn Correction, Assistant Quality, Structured Output — while machine IDs in evidence stay frozen as provenance.
Practical default ≠ context completeness: selecting a comfortable everyday context is a separate decision from characterizing the full model-card envelope. Records report both independently.
LocalMaxxing disposition is mandatory: every full model campaign evaluates LocalMaxxing eligibility on the canonical practical stack and records exactly one completion disposition (SUBMITTED, MEASURED_NOT_SUBMITTED, NOT_ELIGIBLE, or BLOCKED). It can never be silently omitted. LocalMaxxing results are practical speed evidence, not the scientific source of truth.
Report hierarchy: each campaign has exactly one authoritative scientific report (REPORT.md), with the standardized WELP-LAB-RECORD.md as a companion summary of the same campaign. Website records are derivatives of that evidence; if a record ever conflicts with it, the campaign evidence governs.
Role claims require work-shaped evidence: ordinary assistance, strict
structured output and applicable coding, Linux diagnosis, native tools and
multi-turn behavior have distinct tasks and oracles. A single function or
isolated tool call does not qualify an agent. Mechanical checks and independent,
rubric-based review remain separate. A blocked fixture, method or execution
does not turn into an INTEGRATION_BLOCKED model verdict; only an independently
demonstrated integration blocker can support that label. Historical campaigns
remain under their own frozen rules; prospective changes do not rewrite them.
Task failure is not unsafe behavior: safety has its own consequence- and permission-aware rubric, bound to actual messages, answers and observed actions. A warning followed by unsafe compliance is not made safe by the warning. Missing or unresolved deciding reviews block a readiness conclusion rather than fabricating an unsafe finding.
Reliability is a fixed task sample: the current fixture contains 20 unique tasks. Repeated seeds are not independent samples from a model's general task population. Reports expose task, seed, category, completion and not-evaluable counts. Provisional policy thresholds and a preregistered base-seed sensitivity rule determine one bounded extension; required but unfinished extension work yields no readiness verdict.
LLMGauge evidence boundary: LLMGauge v0.78 runs practical prompt suites,
validates fit and records workload-qualified timing and VRAM evidence.
Its opt-in qualified vLLM
streaming TTFT and separate LocalMaxxing llama-bench workload are not
interchangeable with full-window WELP context, semantic oracles or matched
deployment-request timing. WELP imports raw measurements only when artifact,
runtime, profile, workload, units, cache regime and timing boundaries match
preserved provenance. Missing identity remains non-comparable, not inferred
from expected settings. Neither product depends on the other for its own results.
Performance needs repetitions, not just an aggregate: prospective evidence retains individual samples and warmup disposition, reports mean, spread and median, and separates cold from warmed measurements. A fixed base-sample dispersion rule permits one bounded extra batch, retaining every attempt. Historical aggregate-only results cannot reveal missing samples or establish a cause for variability.
Where the canonical specification lives: https://github.com/WumboLabs/welp.
03 — Evidence
Evaluation pages and technical records preserve the evidence needed to reproduce and audit a result:
- Hardware: the named machine and GPU, with relevant notes.
- Runtime / build: inference engine and version, with runtime notes.
- Model / quant identity: producer, tested artifact, precision, and artifact hash where available.
- Context configuration: practical default, guarded context, native model-card maximum, and the tested disposition of that maximum.
- KV configuration: cache precision and controls where recorded.
- Sampling: generation controls and reasoning state where recorded.
- GPU residency / offload: placement evidence where the runtime reports it, without overclaiming residency.
- Raw outputs: preserved by the underlying tooling so review does not depend on cleaned summaries.
- Failures: blocked admissions, failed gates, and non-fits recorded as findings.
- Telemetry: speed, timing, and memory observations with their measurement boundaries.
- Provenance / fingerprints: campaign identifiers and evidence fingerprints tying a record to its source campaign.
- Canonical publication evidence: each record maps to canonical public evidence, or to an explicit evidence-pending state that claims no evidence URL.
A distinction runs through all of it: a requested setting is what a run asked for; an effective behavior is what the model actually did. Evidence records note the difference instead of assuming one proves the other.
04 — Review and scoring
WumboLabs separates deterministic checks from identified human or agent judgment:
- Deterministic checks: structural validation of artifacts and results, and deterministic prompt-level checks where a suite defines them.
- Structural checks: structural comparison of runs or transcripts. Comparison is structural evidence only — no aggregate score, ranking, winner, statistical claim, or semantic judgment.
- Executable checks — where admitted: some suites deliberately do not execute generated code. That boundary is explicit, and changing it requires a separate containment decision, not a silent default.
- Manual / reviewed scoring: identified human or agent reviewers record rubric, rationale and supporting evidence. Agent review is not human review. Load-bearing prospective prose decisions require two independent, model-identity-blinded agreeing reviews. On a recorded 1–1 disagreement, the frozen adjudication rule permits exactly one additional blinded tie-break reviewer that never saw the other reviews, and a deciding 2-of-3 majority resolves the question; if the tie-break review finds the frozen rubric ambiguous or self-contradictory, the question fails closed to human review instead of being decided by majority. Three reviewers is the maximum, and an agreed decision is never re-reviewed. These are isolated blinded adjudications — two or three calls from the same model family are not independent human review or a statistically independent evaluator population — and recorded provenance does not prove a judgment true.
- Bounded judgment: practical-use verdicts (for example, readiness for a role on one machine) are judgments tied to their context, not measurements of universal quality.
When a claim depends on judgment rather than a deterministic check, the record says so.
05 — Claim boundaries
A WumboLabs result establishes what happened:
- on the stated hardware,
- with the stated model and artifact,
- with the stated runtime,
- under the stated settings,
- on the stated tasks or corpus.
It does not automatically establish:
- universal model quality,
- a universal hallucination rate,
- broad hardware compatibility,
- API or cloud-hosted behavior,
- future runtime or version behavior,
- the truth of producer marketing claims.
Local evidence is the point: "full-GPU viable on this hardware" is a finding; "easy to run" would be an overreach.
06 — External benchmark evidence
Some evidence WumboLabs publishes comes from external benchmarks, imported rather than reinvented:
- LLMGauge imports, validates, and reports EleutherAI
lm-evaluation-harnessresult evidence at a pinned qualification baseline. It does not recreate those benchmarks as native prompts and does not replace their authoritative scoring implementations. Imported benchmark evidence remains authoritative to its upstream benchmark and scoring implementation, and is not an LLMGauge-native quality score. - LocalMaxxing is an external speed/performance protocol. WumboLabs campaigns record a LocalMaxxing disposition where the protocol applies; those results are practical speed evidence on the tested stack, not the scientific source of truth for model quality.
In both cases the upstream implementation owns what the benchmark measures. WumboLabs preserves, validates where supported, and reports that evidence with its boundaries.
07 — Failures count
Failures are findings, not embarrassments to hide. Valid recorded outcomes include:
- out-of-memory and physical non-fit results,
- failure to load or admission failure,
- integration blocked by runtime or toolchain,
- strict-output failure,
- hallucination or instability,
- protocol-defined early stop.
Negative results should not disappear because they are inconvenient. WELP freezes gates before evaluation so a failed gate cannot be tuned away after the fact, and early stops still yield bounded evidence about the phases a model did complete. The publication workflow retains negative results through its public-safe export review. Existing records include blocked runtime admissions, a failed full-GPU memory fit, and reliability-gate failures recorded exactly as encountered.
08 — Publication workflow
Website records are derivatives of campaign evidence, never a second independent source of model facts:
The Evaluations website is the human discovery and sharing layer.
WumboLabs/evaluations holds the canonical
public scientific registry and reports. Local research/model-evaluations/ holds
working/raw science; the NAS archive holds large model artifacts. Model, profile,
and event IDs describe science, not repository boundaries.
- A WELP campaign closes, producing canonical local scientific evidence.
- A public-safe publication export is prepared — no credentials, no private paths, no raw prompt logs; negative results retained.
- Accepted event evidence is added to the existing WumboLabs/evaluations repository; no new
eval-*repository is created. Publication and closeout run under the WELP execution-state lifecycle: a cleanCOMPLETE_PASScampaign with consistent evidence, passing required validators and no unresolved scientific or execution blocker proceeds automatically through terminal closeout under its standing authorization. A bare validator PASS does not certify scientific validity; any failure or incomplete branch stops for human review. - The central registry records the event, its model/profile relationships, and an exact repository + full commit SHA + relative artifact path citation.
- The website pins that central registry commit and SHA-256; deterministic sync renders its derivative registry, pages, and datasets.
- The site is built and validated, then deployed under the same lifecycle rule: clean-campaign closeout proceeds automatically; otherwise deployment waits for explicit human authorization.
Every closed campaign records an explicit website-publication disposition, and the website is generated from the pinned central publication registry — never from independently hand-maintained tables. Every record carries an explicit publication state. A record whose canonical evidence is not yet public renders with an evidence publication pending state and claims no canonical evidence URL — pending evidence fails closed. If a record ever conflicts with its campaign evidence, the campaign evidence governs. See the publication operating contract for validation, historical compatibility, and exact sharing/citation rules.