Skip to content

Instantly share code, notes, and snippets.

@plainOldCode
Created September 27, 2026 23:55
Show Gist options
  • Select an option

  • Save plainOldCode/313ef3aab705475beba6c1afeee26100 to your computer and use it in GitHub Desktop.

Select an option

Save plainOldCode/313ef3aab705475beba6c1afeee26100 to your computer and use it in GitHub Desktop.
mia-ai-lab-vs-myllmbox-qwen-3.8-flash-next-vlllm-0.30-comparison
# Qwen3.8 Mia AI ↔ hibrid48 benchmark
## Executive summary
- The cluster deployment was faster in all 16 matched throughput cells, although the speed sweep is a **deployment comparison**, not an isolated checkpoint/quantization comparison. The old full Mia throughput sweep used MTP3/server cap 8; hibrid48 used MTP5/server cap 64.
- On the synthetic quality suite, hibrid48 passed 36/60 short coding attempts vs Mia/NVIDIA 40/60. The paired task-bootstrap interval for the −6.7 percentage-point difference is −18.3 to +5.0 pp, so this small suite does not establish a reliable coding-quality regression or equivalence.
- Long-context quality was 22/24 for hibrid48 vs 20/24 for Mia/NVIDIA. The paired interval is −8.3 to +20.8 pp (only 8 unique tasks); do not treat the point difference as a general capability claim.
- Both models truncated many coding attempts at the default 8,192-token budget. The README-recommended paired 16,384-token coding supplement is underway; hibrid48 is graded at 47/60, while the preserved Mia deployment is restarting for its matched run.
This is the package's original synthetic Core60 diagnostic suite, **not** an official benchmark, contamination-free certification, broad writing evaluation, or proof of model equivalence. The benchmark's own protocol warns against a single global winner.
## Conditions
| | Mia AI / NVIDIA NVFP4 | Cluster / hibrid48 |
|---|---|---|
| Checkpoint revision | `fc694b54fb0174e0913e6adf86691ef85a4ead47` | `e316057ddc794cbf76c3ef28bfd1c7b62b20128c` |
| Serving runtime | vLLM 0.30.0, `vllm/vllm-openai` image | vLLM 0.30.0, cluster recipe `4d28f44ac221086b20f4b95e959a81db66f1c25e`, image v6 |
| Quality-baseline settings | MTP2, max-num-seqs 4, KV auto→BF16, `qwen3_coder` / `qwen3` parsers | MTP5, max-num-seqs 64, KV auto with model dtype BF16, `qwen3_xml` / `qwen3` parsers |
| GPU SM clock cap | 2,000 MHz | 2,000 MHz |
Quality is a deployment comparison: the required runtimes, MTP settings, parsers, and server caps differ. No server setting, MTP count, clock cap, or host tuning was changed to improve the result. The Mia containers were preserved while cluster was active.
## Throughput and GPU energy
Thinking ON; same frozen English/Korean/Chinese/code prompt fixtures; 256-token warmup per domain/concurrency; 3 measured batches at each C1/C2/C4/C8; each request forced to 2,048 output tokens at temperature 0. The primary output rate is summed API completion tokens divided by batch wall time (includes prefill/reasoning, not pure decode). Code measures reasoning-token throughput, not code-quality accuracy. Both nodes' GPU power samples were integrated at about 1-second intervals; system/CPU power is excluded and idle power is not subtracted.
Each speed cell shows **Mia baseline → hibrid48**. These Mia rates are from the existing full C1–C8 sweep, whose server cap was 8 and MTP was 3; hibrid48 used cap 64 and MTP5. The current final Mia recipe is MTP2/cap4, so this is not a one-variable comparison.
### Aggregate output tok/s
| Domain | C1 | C2 | C4 | C8 |
|---|---:|---:|---:|---:|
| English prose | 40.04 → 51.59 (1.29×) | 73.72 → 81.84 (1.11×) | 108.64 → 119.44 (1.10×) | 177.51 → 180.55 (1.02×) |
| Korean prose | 24.63 → 46.98 (1.91×) | 39.90 → 73.64 (1.85×) | 66.59 → 109.58 (1.65×) | 100.55 → 163.26 (1.62×) |
| Chinese prose | 26.84 → 47.80 (1.78×) | 41.89 → 73.43 (1.75×) | 69.31 → 111.55 (1.61×) | 109.67 → 162.04 (1.48×) |
| Coding prompt (reasoning only) | 45.79 → 74.43 (1.63×) | 77.71 → 109.71 (1.41×) | 129.66 → 159.83 (1.23×) | 200.21 → 239.92 (1.20×) |
### GPU output tok/J
| Domain | C1 | C2 | C4 | C8 |
|---|---:|---:|---:|---:|
| English prose | 1.355 → 1.625 | 2.316 → 2.366 | 3.212 → 3.294 | 5.088 → 4.526 |
| Korean prose | 0.789 → 1.426 | 1.233 → 2.120 | 1.955 → 3.014 | 2.877 → 4.109 |
| Chinese prose | 0.872 → 1.448 | 1.293 → 2.105 | 2.023 → 3.067 | 3.130 → 4.019 |
| Coding prompt (reasoning only) | 1.456 → 2.251 | 2.378 → 3.133 | 3.772 → 4.403 | 5.700 → 5.999 |
hibrid48's measured English C8 throughput was only about 2% higher, and its GPU tok/J was lower in that cell. The largest relative throughput differences were at lower concurrency and in Korean/Chinese; those languages were included as comparison fixtures, not as primary quality claims. Cluster run: 48/48 measured batches completed; zero API errors and zero flagged external-load/lag batches. The API-reported cached-input share was 0.0% for the fixture sweep; do not infer that every internal prefix-cache mechanism was inactive.
## Quality results (8,192-token default budget)
Sampling was matched: temperature 0.6, top-p 0.95, top-k 20, Thinking ON; C1, 3 attempts per task. Coding/reasoning requests were capped at 8,192 output tokens, other single-turn categories at 4,096; long tool episodes use the protocol's bounded episode budget. Truncations remain scored failures. Code grading uses isolated Docker sandboxes. No judge-infrastructure errors occurred.
| Category | Unique tasks | Mia/NVIDIA | hibrid48 | Paired B−A pass-rate difference (95% task-bootstrap CI) | Median latency ratio B/A on both-pass attempts |
|---|---:|---:|---:|---:|---:|
| Coding | 20 | 40/60 (66.7%) | 36/60 (60.0%) | −6.7 pp (−18.3, +5.0) | 0.648× |
| Instruction | 8 | 23/24 (95.8%) | 24/24 (100%) | +4.2 pp (0, +12.5) | 0.732× |
| Korean | 4 | 12/12 (100%) | 12/12 (100%) | 0 pp (0, 0) | 0.617× |
| Reasoning | 10 | 30/30 (100%) | 30/30 (100%) | 0 pp (0, 0) | 0.674× |
| Tool | 10 | 30/30 (100%) | 30/30 (100%) | 0 pp (0, 0) | 0.785× |
| Long context | 8 | 20/24 (83.3%) | 22/24 (91.7%) | +8.3 pp (−8.3, +20.8) | 0.795× |
Primary short coding statuses were 41 `ok` + 19 `generation_truncated` for Mia/NVIDIA and 36 `ok` + 24 `generation_truncated` for hibrid48. Long-context status was API `ok` for 24/24 attempts on each model; all 8 prepared token-ID counts/hashes matched the running tokenizer on both deployments (8,164 through 200,012 input tokens). Thus, the two long-context quality failures on hibrid48 were judge/answer failures, not request or tokenizer infrastructure errors.
The cluster 8-case parser smoke passed 8/8 quality checks, including C14; the Mia smoke passed 7/8, with C14 truncated at 8,192. Smoke cases are diagnostic only and are not mixed into the paired Core60 table.
## 16K paired coding supplement
The protocol says to run a new paired larger-budget experiment when many cases truncate and to preserve the original failures. The new configs use the same requests/sampling/repeats/concurrency on each model and set the coding output cap to 16,384; the 8K runs above are untouched.
- hibrid48: 60/60 captured; 47/60 passed, 12 `generation_truncated`, one non-truncated quality failure, zero judge infrastructure failures. Relative to its 8K run this is +11 passing attempts and half as many truncations (24 → 12); this remains a per-attempt result, not proof that all failures are fixed.
- Mia/NVIDIA: not run. The preserved `vllm-fn` containers were restarted, but the 8888 endpoint never became healthy after about 11 minutes; vLLM repeatedly logged the 60-second shared-memory-broadcast warning during FlashInfer autotuning while a worker remained CPU-active. The containers were stopped and preserved, and hibrid48 was restored. Therefore the cluster-only 16K score is descriptive, not a paired comparison.
## Interpretation and limitations
- The throughput gain belongs to the complete deployed stack (checkpoint, MTP, scheduler, kernels, parser/runtime, and concurrency limits). It cannot be attributed to FP4 weight format alone.
- The quality suite is a small synthetic regression set. Exact ties on simple tasks do not prove equivalence; one task moves a category by 5 percentage points. The coding CI crosses zero and the long-context CI is wide.
- Truncation-heavy 8K coding results measure performance under a fixed budget, not unconstrained coding ability. The 16K paired supplement is needed before making a stronger truncation/quality claim.
- Throughput prompts are repetitive, relatively short contexts; code throughput is reasoning-only. Tok/s across languages is not a comparison of equal semantic/text volume.
- GPU tok/J is a two-GPU measurement only, excludes CPU/system power, and does not subtract idle draw. Quality response latency is end-to-end and benefits from repeated prompt caching; it is not a pure decode benchmark.
## Artifacts
- Task/checking instructions: `README.md`, `PROTOCOL.md`, `VALIDATION.md`, `TODO.md`.
- Cluster throughput raw results and summary: `../bench/qwen-update-20260927/cluster/results/hibrid48-20260927T152612Z/`.
- Quality runs: `runs/nvidia-core/`, `runs/cluster-core/`, `runs/nvidia-long/`, `runs/cluster-long/`.
- Primary paired comparisons: `runs/comparison-core/COMPARISON.md`, `runs/comparison-long/COMPARISON.md`.
- Higher-budget configs/runs: `configs/*-highbudget.json`, `runs/cluster-code16k/`; paired Mia run pending.
- Frozen long-context prompts: `prepared/core60-full.jsonl` (shared unchanged between deployments).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment