WUMBOLABS / EVALUATIONS CANONICAL MODEL EVALUATION

/evaluations/neohorse-1-9b

NeoHorse-1-9B

NeoHorse-1-9B — the current WumboLabs evidence state on one page: tested profiles, validated context, and the full chronological testing history. Each value is attributed to the profile and event that measured it.

READY WITH GUARDRAILS / latest evidence 2026-09-25

WumboLabs tests NeoHorse-1-9B on real consumer hardware. This is the canonical model page: current state first, then every tested profile and every evidence event. Values are attributed to the profile and event that measured them; historical findings remain the evidence of their tested stack and are never silently replaced.

Current state

Tested profiles

llama.cpp official Q8_0 (Reasoning On = publisher default, deployment sampler, 32K) — CURRENT

Profile identity: neohorse-1-9b-q8-0-llamacpp-b10999-rtx5070-deployment-reasoning-on.

FieldValue
Runtimenot recorded
Artifactnot recorded
Precisionnot recorded

Status: current canonical/recommended tested surface.

Canonical Evidence Profile Metadata

Events on this profile:

llama.cpp official Q8_0 (Reasoning Off, deployment sampler, 32K) — CURRENT-ALTERNATE

Profile identity: neohorse-1-9b-q8-0-llamacpp-b10999-rtx5070-deployment-reasoning-off.

FieldValue
Runtimenot recorded
Artifactnot recorded
Precisionnot recorded

Status: validated alternate tested surface.

Canonical Evidence Profile Metadata

Events on this profile:

Testing history

Newest first. Each event is one immutable testing/publication event; the exact scientific report lives in the canonical evidence repository linked at the top of each event.

2026-09-25 — First fresh-model WELP campaign under the model-agentic snapshot: full Model science of the publisher-default Reasoning On profile; case C measured (Reasoning Off sibling required)

WELP Fresh-Model Characterization — NeoHorse-1-9B (Reasoning On, publisher default) · profile: llama.cpp official Q8_0 (Reasoning On = publisher default, deployment sampler, 32K) · maturity: CURRENT_WELP · status: READY_WITH_GUARDRAILS

Canonical Evidence / Full Report View on GitHub

Identity
FieldValue
ModelNeoHorse-1-9B
Producernot recorded
Tested artifactnot recorded
Precisionnot recorded
Campaignneohorse-1-9b-rtx5070-welp-reasoning-on-2026-09-25
Record date2026-09-26
Runtime and hardware
FieldValue
Enginenot recorded
Runtime versionnot recorded
HardwareNVIDIA GeForce RTX 5070 12 GB
WELP outcome
  • Outcome: COMPLETE_PASS
  • Classification: not recorded

Publication state: published — canonical evidence: https://github.com/WumboLabs/evaluations/blob/bf5e29e1afee4f391c4b8f83e77636b887453f53/models/neohorse-1-9b/events/neohorse-1-9b-rtx5070-welp-reasoning-on-20260925/REPORT.md

Context profile
FieldValue
Practical defaultnot recorded tokens
Guarded contextnot recorded tokens
Native model-card maximumnot recorded tokens
Model-card envelope completenot recorded
Native maximum dispositionnot recorded
Quality and capabilities
  • Constrained result: not recorded
Guardrails and limitations

Reliability: not recorded

LocalMaxxing
FieldValue
StatusSUBMITTED
Canonical contextnot recorded tokens
tok/s outnot recorded
TTFTnot recorded
Submission referencenot recorded
verifiedRunnull (not claimed)

canonical practical profile Q8_0 llama.cpp; the benchmark workload is reasoning-state invariant

Canonical evidence

Canonical public evidence: https://github.com/WumboLabs/evaluations/blob/bf5e29e1afee4f391c4b8f83e77636b887453f53/models/neohorse-1-9b/events/neohorse-1-9b-rtx5070-welp-reasoning-on-20260925/REPORT.md

This event section is a human-readable derivative of the accepted local WELP campaign evidence named above; the campaign's REPORT.md is the authoritative scientific source. Results are bounded by the tested artifact, runtime, hardware, configuration, and protocol snapshot, and are not universal model rankings.

2026-09-25 — Full WELP characterization of the Reasoning Off deployment profile: independent setup/calibration, same frozen task set, dual-profile completion of the case-C model

WELP Reasoning-Profile Characterization — NeoHorse-1-9B (Reasoning Off, supported alternate) · profile: llama.cpp official Q8_0 (Reasoning Off, deployment sampler, 32K) · maturity: CURRENT_WELP · status: READY_WITH_GUARDRAILS

Canonical Evidence / Full Report View on GitHub

Identity
FieldValue
ModelNeoHorse-1-9B
Producernot recorded
Tested artifactnot recorded
Precisionnot recorded
Campaignneohorse-1-9b-rtx5070-welp-reasoning-off-2026-09-25
Record date2026-09-26
Runtime and hardware
FieldValue
Enginenot recorded
Runtime versionnot recorded
HardwareNVIDIA GeForce RTX 5070 12 GB
WELP outcome
  • Outcome: COMPLETE_PASS
  • Classification: not recorded

Publication state: published — canonical evidence: https://github.com/WumboLabs/evaluations/blob/bf5e29e1afee4f391c4b8f83e77636b887453f53/models/neohorse-1-9b/events/neohorse-1-9b-rtx5070-welp-reasoning-off-20260925/REPORT.md

Context profile
FieldValue
Practical defaultnot recorded tokens
Guarded contextnot recorded tokens
Native model-card maximumnot recorded tokens
Model-card envelope completenot recorded
Native maximum dispositionnot recorded
Quality and capabilities
  • Constrained result: not recorded
Guardrails and limitations

Reliability: not recorded

LocalMaxxing
FieldValue
StatusSUBMITTED
Canonical contextnot recorded tokens
tok/s outnot recorded
TTFTnot recorded
Submission referencenot recorded
verifiedRunnull (not claimed)

canonical practical profile Q8_0 llama.cpp; the benchmark workload is reasoning-state invariant

Canonical evidence

Canonical public evidence: https://github.com/WumboLabs/evaluations/blob/bf5e29e1afee4f391c4b8f83e77636b887453f53/models/neohorse-1-9b/events/neohorse-1-9b-rtx5070-welp-reasoning-off-20260925/REPORT.md

This event section is a human-readable derivative of the accepted local WELP campaign evidence named above; the campaign's REPORT.md is the authoritative scientific source. Results are bounded by the tested artifact, runtime, hardware, configuration, and protocol snapshot, and are not universal model rankings.

2026-09-25 — First fresh-model WELP Agentic event: dual-adapter qualification, three required task classes with two scored (1 PASS / 1 FAIL) and the system class INTEGRATION_BLOCKED by a deterministic frozen-harness defect

WELP Agentic — NeoHorse-1-9B (publisher-default profile, native tool calls) · profile: llama.cpp official Q8_0 (Reasoning On = publisher default, deployment sampler, 32K) · maturity: CURRENT_WELP · status: READY_WITH_GUARDRAILS

Canonical Evidence / Full Report View on GitHub

Identity
FieldValue
ModelNeoHorse-1-9B
Producernot recorded
Tested artifactnot recorded
Precisionnot recorded
Campaignneohorse-1-9b-rtx5070-welp-agentic-2026-09-25
Record date2026-09-26
Runtime and hardware
FieldValue
Enginenot recorded
Runtime versionnot recorded
HardwareNVIDIA GeForce RTX 5070 12 GB
WELP outcome
  • Outcome: COMPLETE_PASS
  • Classification: not recorded

Publication state: published — canonical evidence: https://github.com/WumboLabs/evaluations/blob/bf5e29e1afee4f391c4b8f83e77636b887453f53/models/neohorse-1-9b/events/neohorse-1-9b-rtx5070-welp-agentic-20260925/REPORT.md

Context profile
FieldValue
Practical defaultnot recorded tokens
Guarded contextnot recorded tokens
Native model-card maximumnot recorded tokens
Model-card envelope completenot recorded
Native maximum dispositionnot recorded
Quality and capabilities
  • Constrained result: not recorded
Guardrails and limitations

Reliability: not recorded

LocalMaxxing
FieldValue
StatusSUBMITTED
Canonical contextnot recorded tokens
tok/s outnot recorded
TTFTnot recorded
Submission referencenot recorded
verifiedRunnull (not claimed)

canonical practical profile Q8_0 llama.cpp; the benchmark workload is reasoning-state invariant

Canonical evidence

Canonical public evidence: https://github.com/WumboLabs/evaluations/blob/bf5e29e1afee4f391c4b8f83e77636b887453f53/models/neohorse-1-9b/events/neohorse-1-9b-rtx5070-welp-agentic-20260925/REPORT.md

This event section is a human-readable derivative of the accepted local WELP campaign evidence named above; the campaign's REPORT.md is the authoritative scientific source. Results are bounded by the tested artifact, runtime, hardware, configuration, and protocol snapshot, and are not universal model rankings.

Canonical evidence

All canonical public evidence lives in WumboLabs/evaluations. Each event links an immutable full-commit/path citation; each profile remains a distinct scientific identity, not a separate repository.

  • neohorse-1-9b-q8-0-llamacpp-b10999-rtx5070-deployment-reasoning-on: Profile Metadata
  • neohorse-1-9b-q8-0-llamacpp-b10999-rtx5070-deployment-reasoning-off: Profile Metadata