Skip to content

Fused comprehensive reference runs

These completed local reference receipts were captured from August 18 through August 20, 2026. They are not upstream leaderboard scores. They measure the pinned Mere comprehensive subset, scorers, sampling profiles, and runtime listed on this page. Runner, fixture, sandbox, model, or profile changes require a new dated receipt; they do not replace these results.

Completed profile scope

Laguna XS 2.1 and Nemotron Lightning each completed their single native profile. The Qwen3.8 plan declares low, medium, and xhigh reasoning profiles; this receipt intentionally stops after all 550 low-profile rows. Medium and xhigh remain at zero and are not represented by the score.

Model and profileCompleted scopeScoredUnscoredStrict passesStrict pass rateMean score
Qwen3.8 27B, low reasoning550/550 low; 550/1,650 plan5500404/55073.5%0.8170
Laguna XS 2.1, native550/55050050247/50049.4%0.6473
Nemotron 3.5 Lightning, native550/55050050172/50034.4%0.5794

passed is the all-or-nothing case contract. score retains partial credit where a scorer supports it, so strict pass rate and mean score answer different questions. Unscored capability rows are reported separately and are not converted into failures.

Qwen3.8 advertised image input and therefore scored all 50 vision rows. Laguna and Nemotron did not, so their vision rows were unscored. The fair like-for-like comparison removes Qwen's vision rows and uses the same 500 text, code, and tool rows for every model:

Model and profileNon-vision strict passesStrict pass rateMean score
Qwen3.8 27B, low reasoning359/50071.8%0.8031
Laguna XS 2.1, native247/50049.4%0.6473
Nemotron 3.5 Lightning, native172/50034.4%0.5794

Results by source family

Each cell shows strict passes / scored rows (rate) · mean score. Counts are case-trials, not unique prompts.

Source familyRowsQwen3.8 lowLaguna XS 2.1Nemotron Lightning
Mere authored chat260199/260 (76.5%) · 0.9223155/260 (59.6%) · 0.8478105/260 (40.4%) · 0.7448
Mere authored tools5045/50 (90.0%) · 0.980037/50 (74.0%) · 0.941010/50 (20.0%) · 0.7517
BFCL v35020/50 (40.0%) · 0.400020/50 (40.0%) · 0.40007/50 (14.0%) · 0.1400
OpenAI HumanEval1515/15 (100%) · 1.000010/15 (66.7%) · 0.666712/15 (80.0%) · 0.8000
EvalPlus HumanEval+3023/30 (76.7%) · 0.76675/30 (16.7%) · 0.16676/30 (20.0%) · 0.2000
EvalPlus MBPP+3025/30 (83.3%) · 0.833311/30 (36.7%) · 0.366723/30 (76.7%) · 0.7667
LiveCodeBench3022/30 (73.3%) · 0.73338/30 (26.7%) · 0.26678/30 (26.7%) · 0.2667
LongBench v13510/35 (28.6%) · 0.22091/35 (2.9%) · 0.06181/35 (2.9%) · 0.0706
Mere authored vision5045/50 (90.0%) · 0.95600 scored; 50 unscored0 scored; 50 unscored

Within these exact receipts, Qwen3.8 low had the highest strict-pass result in every scored source family except BFCL, where it tied Laguna. This comparison does not predict the unfinished Qwen medium or xhigh profiles and is not a claim about the complete upstream benchmark collections.

Shared run contract

FieldRecorded value
HostApple M4 Max, 16 logical processors, 128 GiB unified memory, arm64
Operating systemmacOS 26.5.2, build 25F84
mere.run0.40.1
Runner executable204,648,000 bytes; SHA-256 71cb98eeda7fb71f6e92399de7753faadcfc84c5b6b6f19da107a2d7e7633b5f
Suite manifestmere-fused-v1 1.2.0; SHA-256 4db423313708db492caf9e5eb70ab0fb2e18021a1eb0dee74c83b3b2b5d5e954
Per-profile shape110 cases × 5 trials = 550 rows
Quality laneSampled final target, exact policy, logprob summaries
Context32,768 tokens
Code execution/usr/bin/python3, 5-second timeout, macos-sandbox-exec
Response captureDisabled; the receipts do not contain visible model responses
Duplicate rowsZero in each schema-v2 checkpoint

Qwen3.8 27B low-reasoning receipt

  • Model ID: vision-chat-q38-27b.
  • Artifact: Qwen/Qwen3.8-27B.
  • Artifact revision: 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0.
  • Runtime-manifest SHA-256:c5813b1044dff6f7aec2464c31d829f96aba19671839b70d83d617a6a1c42ec0
  • Profile: qwen3.8-native-low; reasoning effort 0.2, thinking enabled, temperature 1.0, top-p 0.95, top-k 20, min-p 0
  • Checkpoint window: 2026-08-20T01:15:56Z to 2026-08-20T09:21:19Z.
  • Completion scope: Low 550/550; medium 0/550; xhigh 0/550.
  • Plan SHA-256:64a9b0fa53478c177947101cb385ef73ac6025827fbb8be2d9031a719c13d5b2
  • Frozen receipt filename: qwen38-low-comprehensive-b944c0e7.json.
  • Receipt SHA-256:c6b70396f183ec7183563d0b2789a32893841285e843cafb8d3ff796e18698f4

Laguna XS 2.1 receipt

  • Model ID: text-chat-laguna-xs-2-1.
  • Artifact: poolside/Laguna-XS-2.1-NVFP4-mlx.
  • Artifact revision: 841778bda563a36104dd521e37d99218e46f4f25.
  • Runtime-manifest SHA-256:213b9c6ee642824a2294715aa37c84ee5772e45ca2d671f3cc2aafa2e5c368f7
  • Profile: Temperature 1.0, top-p 1.0, top-k 20, min-p 0.02.
  • Checkpoint window: 2026-08-19T12:47:30Z to 2026-08-19T14:47:50Z.
  • Plan SHA-256:ba0c0785c17f857a58428c024eae9448130b749e39c4e525eb60910c16720af1
  • Receipt filename: laguna-xs-2-1-comprehensive-b944c0e7.json.
  • Receipt SHA-256:6b8ddccd034cd00515becbb6f7eb31d12a2f8aae2ccb93d00cc7212237dd522d

Nemotron Lightning receipt

  • Model ID: text-chat-nemotron-35-lightning.
  • Managed artifact:Sawfwair/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4-MLX
  • Managed artifact revision:6699e5fd3f0c5b392bb3f8bac2443276bb41958a
  • Pinned base-model revision:nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4@e0b753dc24903ad4d62f5696077da22020eca89a
  • Runtime-manifest SHA-256:369da6748557f0a3a05256cbeea933e789b7a9b15c5e08403a24946ec6b453e7
  • Profile: Temperature 1.0, top-p 0.95, top-k disabled, min-p 0.
  • Checkpoint window: 2026-08-18T18:20:13Z to 2026-08-19T12:24:41Z.
  • Plan SHA-256:bd04d29ac3a67b23d9c23ffc9477b7284274adf22de14af65cd7ac14de5aa8a7
  • Receipt filename: nemotron-lightning-comprehensive-b944c0e7.json.
  • Receipt SHA-256:b1631e7504c74fcf41a34b44c1969dc2d6984b75e43a3bdac29fe9894b3ecd68

Interpretation limits

  • These results apply to the exact Mere subset and hashes in the shared run contract, not the full HumanEval+, MBPP+, LiveCodeBench, BFCL, or LongBench leaderboards.
  • Qwen3.8 medium and xhigh were intentionally deferred. The Qwen checkpoint is complete for low reasoning and partial for its three-profile plan.
  • The two text-only catalog profiles skipped vision by capability declaration. Compare their scored text/code/tool rows; do not call the vision lane failed.
  • Nemotron and Qwen were paused and resumed. Their checkpoint timestamps are preservation windows, not throughput measurements. Logprob capture also makes these quality runs unsuitable as performance benchmarks.
  • Failing code-row diagnostics on this host often included an xcrun_db cache warning before the candidate's Python syntax, name, assertion, or test error. Passing code rows were still recorded. Treat any rerun after sandbox or toolchain changes as a new receipt rather than rewriting a historical one.
  • The raw checkpoints and normalized external fixtures remain local, ignored artifacts. Their hashes are recorded here so a later summary cannot silently substitute different result bytes.

See Mere fused model evaluation for the source mix, question and answer examples, scoring semantics, and reproduction procedure.

Released under the MIT License.