Experiment 004 · GLM-5.3 · 8× MI355X · 1 gate breached of 5

The gate that missed by 0.0202 pp差 0.0202 pp 越线的那道门槛

A controlled A/B between the official FP8 checkpoints of GLM-5.3 and their MXFP4 conversions, on eight MI355X, against five thresholds written down before the run. Four gates pass. One misses by a fifth of a percent of itself. This page is about why that is reported as a breach, why a checkpoint 42% smaller on disk streams more bytes per generated token, and why half the experiment is not allowed to conclude anything.

GLM-5.3 官方 FP8 checkpoint 与其 MXFP4 转换版在 8 张 MI355X 上的受控 A/B, 对照的是开跑前就写下的五条门槛。 四条通过, 一条以自身的千分之二越线。 这篇要讲的是: 为什么它被记成越线而不是被四舍五入抹掉; 为什么一个盘上小 42% 的 checkpoint 每生成一个 token 反而要读更多字节; 以及为什么这场实验的另一半不被允许得出任何结论。

Continuation now live后续评测直播中

The full-model Terminal-Bench 4.0 continuation now runs as two monitored pools. Follow scored attempts, queue pressure, ETA, and the issue ledger on the live dashboard →

Full 模型的 Terminal-Bench 4.0 后续评测正在两个受监控的 pool 上运行。 可在实时看板 →查看计分尝试、队列压力、预计完成时间和问题台账。

Status at stop停机时状态

Paused at a completed case boundary. Server exited through its cleanup trap, port 31103 closed, all eight GPU locks released. The Flash A/B is complete and carries one boundary warning. The Full A/B has no candidate side at all — the partial FP8 performance records are marked not publishable, and this page honours that mark.

在一个已完成的 case 边界暂停。 服务经 cleanup trap 正常退出, 端口 31103 关闭, 八张 GPU 的锁全部释放。 Flash 的 A/B 已经完整, 带一条边界告警。 Full 的 A/B 连候选侧都没有启动—— 那批 FP8 的 partial performance 记录被标记为不可发布, 这一页遵守这个标记。

Silicon硬件
8× MI355X · gfx950 · mia1-p02-g23
Runtime运行时
host ROCm 7.2.0 · torch 2.9.1 · SGLang + AITER source overlay
Window窗口
21:02:18 · 2026-08-31 07:33:49 → 09-01 04:36:07 UTC
Data数据
data/glm53-mxfp4-mi355x
Status状态

Where it stopped, and what that leaves standing停在哪里, 以及还剩下什么

Four things were supposed to be measured: two model sizes, each in two numeric formats. Two of the four are finished. One is finished on the baseline side only. The fourth was never started.

本来要测四样东西: 两个模型规模, 每个两种数值格式。 四样里完成了两样, 第三样只完成了基线一侧, 第四样根本没有启动。

The point of saying that first is that it decides what the rest of the page is allowed to claim. A quantization A/B is a paired comparison; its unit of evidence is a pair. One complete side is not half an answer, it is zero answers about the format — it is only a fact about the baseline. So the campaign produced exactly one comparison, on the smaller model, and one unpaired baseline on the larger one.

先说这个, 是因为它决定了这一页后面可以主张什么。 量化 A/B 是配对比较, 它的证据单位是「一对」。 只完成一侧不是半个答案, 而是关于这个数值格式的零个答案—— 它只是关于基线的一个事实。 所以这轮跑出来的是: 小模型上的一个完整比较, 加上大模型上的一个没有对照的基线。

Variant变体 CheckpointCheckpoint AccuracyAccuracy PerformancePerformance Standing判定
GLM-5.3-Flash FP8 zai-org/GLM-5.3-Flash 1198 rows complete1198 行完成 3 rounds × 3 concurrencies + 2 canaries3 轮 × 3 并发 + 双 canary usable可用
GLM-5.3-Flash MXFP4 OneNexus/GLM-5.3-Flash-MXFP4 1198 rows complete1198 行完成 3 rounds × 3 concurrencies + 2 canaries3 轮 × 3 并发 + 双 canary boundary warning边界告警
GLM-5.3 Full FP8 zai-org/GLM-5.3 2017 rows complete2017 行完成 stopped mid-sequence: canary-before, c1-r1, c8-r1中途停止: canary-before、c1-r1、c8-r1 accuracy keeps · perf not publishableaccuracy 保留 · 性能不可发布
GLM-5.3 Full MXFP4 OneNexus/GLM-5.3-MXFP4 not started未启动 not started未启动 no Full A/B exists不存在 Full A/B

The evidence matrix as it stands. Every row of this table is reconstructible from the archive: the accuracy summaries in accuracy/, the perf summaries in perf/, and the stop record in INTERMEDIATE_STOP.json.

当前的证据矩阵。 表里每一行都能从归档里重建: accuracy 汇总在 accuracy/, 性能汇总在 perf/, 停机记录在 INTERMEDIATE_STOP.json。

The stop itself was deliberate rather than a crash, and the difference matters for what can be reused. The performance runner was allowed to finish the case it was inside — c8-r1 at ISL 8192 / OSL 1024 — and its record was written out completely before the runner exited; c32-r1 was never launched. The server then went down through its own cleanup path, which is what released the eight per-GPU locks. A crash mid-case would have left a truncated JSONL that looks like a valid measurement, and a leaked lock that would have poisoned the next run's idle check.

这次停机是主动的, 不是崩溃, 而这个区别决定了哪些东西还能复用。 性能 runner 被允许跑完它当时所在的那个 case—— ISL 8192 / OSL 1024 下的 c8-r1—— 记录完整写出后 runner 才退出; c32-r1 从未启动。 随后服务经自己的 cleanup 路径下线, 八个 per-GPU lock 因此被释放。 如果是 case 中途崩溃, 留下的会是一个看起来像有效测量的截断 JSONL, 外加一把泄漏的锁, 让下一次运行的空闲检查失去意义。

The rule the stop record carries停机记录里写下的规则

The stop marker does not just say when it stopped; it says what may be done with what survived. Accuracy for Full FP8 may be reused only if the frozen checkpoint, SGLang, AITER, kernel runtime, datasets and generation protocol are all unchanged. The partial performance records may not be appended to a final table under any circumstances — the sequence must be re-run into a fresh output directory so all three repeats and both canaries belong to one uninterrupted run.

停机标记不只记录了何时停止, 还写明了幸存下来的东西可以怎么用。 Full FP8 的 accuracy 仅当冻结的 checkpoint、SGLang、AITER、kernel 运行时、数据集与生成协议全部不变时才可复用。 那批 partial performance 记录在任何情况下都不得追加到最终表里—— 整段序列必须重跑进一个全新的输出目录, 使三轮重复和两次 canary 属于同一次不中断的运行。

Protocol协议

Five lines drawn before the data existed在数据出现之前画好的五条线

A quantized checkpoint is a trade, and a trade needs a price agreed in advance. These five thresholds were written into the protocol before the first formal run started, which is the only property that makes them worth anything.

量化 checkpoint 是一笔交易, 而交易需要事先谈好价格。 这五条门槛是在第一次正式运行之前写进协议的—— 这是它们唯一有价值的属性。

Why five and not one. A single scalar — "accuracy recovery" — can be held constant by a model that has quietly changed behaviour underneath it. Each of these gates watches a different way the trade can go bad, and each one is the cheapest available detector for its own failure. Per-benchmark accuracy catches a regression concentrated in one domain that an aggregate would average away. Combined correct-count recovery catches a broad shallow loss that no single benchmark reaches significance on. Length-finish rate catches the failure where the model does not get wrong, it gets long — a slightly less confident model reasons further, hits the token cap, and is scored as wrong for running out of room rather than for being mistaken. Throughput ratio is the reason for doing any of this. And canary drift catches the case where the machine, not the model, was the variable.

为什么是五条而不是一条。 单一标量—— 比如「accuracy recovery」—— 完全可能在模型底层行为已经改变的情况下保持不变。 这五条门槛各自盯住交易变坏的一种方式, 而且每一条都是它所对应故障的最便宜的探测器。 逐项 accuracy 抓的是集中在某一个领域的回退, 这种回退会被汇总值平均掉。 合计正确数 recovery 抓的是分散而浅的整体损失, 单个 benchmark 上达不到显著性。 length finish rate 抓的是模型没有变错、而是变长的那种故障—— 稍微不那么确定的模型会多推一段, 撞上 token 上限, 于是因为「写不下」而不是因为「答错了」被判错。 吞吐比是做这一切的理由本身。 而 canary 漂移抓的是变量其实是机器而不是模型的那种情况。

Plate I · The gate ledger4 pass · 1 breach
GATE MEASURED AGAINST THE LINE VERDICT accuracy delta WORST BENCHMARK +0.20 −0.60 −1.52 −3.0 pp FLOOR −2.0 pp +1.0 pp PASS correct-count RECOVERY, COMBINED 99.5301% 96% FLOOR 98% 101% PASS length-finish rate INCREASE, WORST +0.40 ×2 +2.0202 0 pp CEILING 2.0 pp +3.0 pp BREACH throughput ratio WORST CONCURRENCY 91.78 · 92.21 · 92.45% 85% FLOOR 90% 100% PASS canary drift BEFORE VS AFTER +0.46 +2.30 0% CEILING 5% 6% PASS DETAIL · GATE 3 · THE READINGS 198 ROWS CAN ACTUALLY PRODUCE GPQA Diamond 198 ROWS CEILING 2.0000 pp 0.0202 pp — the entire breach 1.5152 pp · 3 MORE ROWS 2.0202 pp · 4 MORE ROWS 2.5253 pp · 5 MORE ROWS On a 198-row sample the reading moves in steps of 0.50505 pp. The 2.0 pp ceiling falls in a gap between two attainable values — it names a number this instrument cannot report.

Each row is drawn on its own scale, with the failing side washed in carmine and the preset line dashed. Four measurements land clear of their line. The third does not, and the detail strip below shows by how much: the ceiling was set at a value that a 198-row benchmark can never return, because truncation count is an integer and each row is worth 0.50505 pp.

每一行按各自的刻度绘制, 不通过的一侧用绛红色打底, 预设线用虚线标出。 四项测量落在线内。 第三项没有, 下方的细节条标出差了多少: 这条上限被设在了一个 198 行 benchmark 永远返回不了的数值上, 因为截断计数是整数, 每一行值 0.50505 pp。

Gate门槛 What it detects它探测什么 Line线 Measured实测 Mark判定
Per-benchmark accuracy delta逐项 accuracy 差值 a regression concentrated in one domain集中在单一领域的回退 ≥ −2.0 pp −1.52 pp pass
Combined correct-count recovery合计正确数 recovery a broad shallow loss no single set resolves单个数据集分辨不出的广泛浅层损失 ≥ 98% 99.5301% pass
Length-finish-rate increaselength finish rate 增幅 answers scored wrong for running out of room因为写不下而被判错的答案 ≤ 2.0 pp 2.0202 pp breach
Median total throughput ratio中位总吞吐比 the trade failing at its own purpose这笔交易在自己的目的上失败 ≥ 90% 91.78% pass
Before/after canary drift前后 canary 漂移 the machine, not the model, as the variable变量其实是机器而不是模型 ≤ 5% 2.30% pass

Each "measured" cell is the worst reading across the benchmarks or concurrencies the gate covers, since a gate is only satisfied when every case it covers satisfies it.

「实测」一列取的是该门槛所覆盖的 benchmark 或并发中最差的一个读数, 因为一条门槛只有在它覆盖的每个 case 都满足时才算满足。

Why the canary gate is not ceremony为什么 canary 这条门槛不是仪式

The FP8 side drifted +2.30% between its opening and closing canary — a short fixed 512/128 probe run before and after the whole sequence. That is not a small number in context: it is larger than the 1.1% round-to-round spread of the c1 measurement it brackets, and it is a third of the 7.6–8.2% effect the experiment is trying to resolve. Without the bracket there is no way to tell a real 8% difference from a machine that warmed up. With it, the two sides can be compared, because both were measured inside a window whose endpoints are known.

FP8 一侧在开场与收场 canary 之间漂移了 +2.30%—— canary 是整段序列前后各跑一次的固定 512/128 短探针。 放在上下文里这不是小数: 它大于它所夹住的 c1 测量本身 1.1% 的轮间离散度, 也相当于这个实验试图分辨的 7.6%–8.2% 效应的三分之一。 没有这层括号, 就无法区分「真实的 8% 差异」和「机器热起来了」。 有了它, 两侧才可比, 因为它们都是在一个端点已知的窗口里测出来的。

The breach越线

A threshold that bends when it is barely exceeded was never a threshold一碰就弯的门槛, 从来就不是门槛

GPQA Diamond truncated 55 of 198 answers on FP8 and 59 on MXFP4. Four more rows out of 198 is an increase of 2.0202 percentage points against a ceiling of 2.0. The overshoot is one part in a hundred of the ceiling itself.

GPQA Diamond 在 FP8 上有 55/198 个回答被截断, 在 MXFP4 上是 59 个。 198 里多了 4 行, 增幅 2.0202 个百分点, 上限是 2.0。 越线的幅度是上限自身的百分之一。

Every instinct says round it. The measurement has no error bar attached, the two numbers are indistinguishable in any practical sense, and no downstream decision changes if the gate is marked green. All of that is true, and none of it is the point.

所有直觉都在说: 四舍五入掉吧。 这个测量本身没有附误差棒, 两个数在任何实际意义上都不可区分, 把这条门槛标绿也不会改变下游的任何决策。 这些都成立, 但都不是重点。

A threshold does exactly one job: it converts a number into a decision without the person holding the number getting to choose. Its entire value comes from being fixed before the result is known. The moment it is allowed to move after the result is known — even by a hundredth of itself, even in a case where moving it is obviously harmless — it stops being a decision rule and becomes a description of what happened. And the damage is not local. A reader who sees one gate relaxed has to assume every gate on the page was evaluated against a line that could have moved, which means none of the four passes carry information any more. The 0.0202 pp is not what is being protected. The other four gates are.

门槛只做一件事: 在持有数字的人无权选择的前提下, 把一个数字变成一个决定。 它的全部价值来自「在结果已知之前就被固定」。 一旦它被允许在结果已知之后移动—— 哪怕只移动自身的百分之一, 哪怕在这个具体 case 里移动它显然无害—— 它就不再是决策规则, 而变成了对已发生之事的描述。 而且损害不是局部的。 一个读者只要看到有一条门槛被放宽, 就必须假设这一页上每条门槛都是对着一条可以移动的线评估的, 于是那四个「通过」也不再携带任何信息。 被保护的不是那 0.0202 pp, 是另外四条门槛。

The threshold was mis-specified, which is a different problem这条门槛本身写错了, 那是另一个问题

There is a real defect here, and it is worth separating from the reading. Truncation count on GPQA Diamond is an integer over 198 rows, so the measurable grid is 0.50505 pp wide. Between the FP8 baseline and any candidate, the attainable increases are 1.5152 pp (three more truncations) and 2.0202 pp (four). A ceiling of 2.0 pp sits in the gap between them. It cannot ever be met with equality, and the difference between "just inside" and "just outside" is decided by a single question out of 198 flipping from a complete answer to a truncated one.

这里确实有一个真实缺陷, 值得和读数本身分开来说。 GPQA Diamond 上的截断计数是 198 行上的整数, 所以可测的格点宽度是 0.50505 pp。 在 FP8 基线与任何候选之间, 可达的增幅是 1.5152 pp(多三次截断)与 2.0202 pp(多四次)。 2.0 pp 的上限落在这两者之间的空隙里。 它永远无法被恰好取到, 而「刚好在内」与「刚好在外」之间的差别, 取决于 198 道题里的某一道从完整作答翻成截断。

The fix is not to loosen the number. It is to state the gate in the units the instrument actually reports — at most three additional truncations on a 198-row set — or to choose the sample size that makes 2.0 pp resolvable. Both changes have to be made before the next run, not after this one, or they are the same act of moving the line that the previous paragraph rejects.

修法不是把数字放松, 而是用仪器实际报告的单位来陈述这条门槛—— 在 198 行的集合上最多多出三次截断—— 或者选一个能让 2.0 pp 可分辨的样本量。 这两种改动都必须在下一次运行之前做, 而不是在这一次之后做, 否则它就是上一段所拒绝的那个「移动线」的动作。

What the gate is actually watching for这条门槛真正在盯什么

Length-finish rate is the only gate here that watches a failure mode capable of hiding inside a passing accuracy score. A quantized model that has lost a little confidence does not answer wrong — it reasons for longer, hits the 16K output cap, and returns nothing scoreable. On GPQA under reasoning_effort=max the FP8 baseline already truncates 27.78% of the time, so the metric sits on a steep part of its own curve: small drifts in reasoning length translate directly into scored-wrong answers. That is why the ceiling on this gate is tight and why an overshoot, however small, is worth writing down rather than rounding away.

length finish rate 是这里唯一一条盯着「能藏在及格 accuracy 分数背后」的故障模式的门槛。 一个稍微失去一点确定性的量化模型不会答错—— 它会多推一段, 撞上 16K 输出上限, 然后返回一个无法评分的东西。 在 reasoning_effort=max 下的 GPQA 上, FP8 基线本身就有 27.78% 的截断率, 所以这个指标正处在自己曲线的陡峭段: 推理长度上的微小漂移会直接转化成被判错的答案。 这就是为什么这条门槛的上限收得很紧, 也是为什么越线不管多小都值得写下来而不是抹平。

The resulting sentence about this checkpoint is longer than anyone would like, and it is the only accurate one available: Flash MXFP4 passes the accuracy-recovery and throughput gates and carries a boundary warning on truncation rate. It is not "equivalent to FP8". It is not "failed". Both of those are shorter, and both would be reporting a decision the data did not make.

于是关于这个 checkpoint 的表述比谁都希望的要长, 而它是唯一准确的一句: Flash MXFP4 通过了 accuracy recovery 与吞吐门槛, 在截断率上带一条边界告警。 它不是「与 FP8 等价」, 也不是「未通过」。 那两句都更短, 而且都是在报告一个数据没有做出的决定。

Result结果

The one comparison this campaign actually produced这一轮真正跑出来的那一个比较

Same checkpoint family, same runtime, same datasets, same generation settings, same three-round performance protocol on both sides. 1198 paired accuracy rows and nine performance runs per side, bracketed by canaries.

同一 checkpoint 家族、同一运行时、同一数据集、同一生成参数、两侧同一套三轮性能协议。 每侧 1198 行配对 accuracy 和九次性能运行, 前后由 canary 夹住。

Accuracy, and what the delta hidesAccuracy, 以及差值掩盖了什么

Dataset数据集 FP8 MXFP4 Delta差值 McNemar pMcNemar p truncated截断 Gate 3门槛 3
GSM8K · 500 489 · 97.80% 490 · 98.00% +0.20 pp 1.000 1 → 3 +0.40 pp
MMLU · 500 437 · 87.40% 434 · 86.80% −0.60 pp 0.690 37 → 39 +0.40 pp
GPQA Diamond · 198 138 · 69.70% 135 · 68.18% −1.52 pp 0.7283 55 → 59 +2.0202 pp
Combined · 1198合计 · 1198 1064 1059 −0.42 pp — 93 → 101 99.5301%

From flash-accuracy-comparison.json. The combined recovery is 1059/1064 = 0.9953007518796992. No benchmark shows a significant difference under an exact McNemar test.

来自 flash-accuracy-comparison.json。 合计 recovery 为 1059/1064 = 0.9953007518796992。 在精确 McNemar 检验下没有任何一个 benchmark 显示出显著差异。

The line worth pausing on is not the −5 answers. It is that the two models disagreed on 67 of the 1198 rows and the net came out at −5. Thirty-six questions FP8 got right and MXFP4 got wrong; thirty-one the other way. The aggregate delta is a difference of two large, nearly equal counts, and it is small for the same reason a coin lands near even — not because the underlying behaviour is unchanged.

值得停下来看的不是「少了 5 个正确答案」, 而是两个模型在 1198 行里有 67 行给出了不同的对错结果, 而净值落在 −5。 36 道 FP8 答对而 MXFP4 答错, 31 道反过来。 这个汇总差值是两个很大且几乎相等的计数之差, 它小的原因和硬币接近对半开是同一个原因—— 不是因为底层行为没有变。

Plate II · Paired disagreement, not net difference1198 rows · exact McNemar
GSM8K · 500 FP8 CORRECT FP8 WRONG MXFP4 CORRECT MXFP4 WRONG 485 5 4 6 DISCORDANT 9 · p = 1.000 MMLU · 500 FP8 CORRECT FP8 WRONG MXFP4 CORRECT MXFP4 WRONG 423 11 14 52 DISCORDANT 25 · p = 0.690 GPQA DIAMOND · 198 FP8 CORRECT FP8 WRONG MXFP4 CORRECT MXFP4 WRONG 120 15 18 45 DISCORDANT 33 · p = 0.7283 ALL 1198 PAIRED ROWS 1028 BOTH CORRECT 103 BOTH WRONG 36 FP8 only 31 MXFP4 only

McNemar looks only at the two off-diagonal cells, because the rows both models agree on carry no information about a difference between them. Sixty-seven rows changed side; the reported delta of −5 is what is left after 36 and 31 cancel.

McNemar 只看两个非对角格, 因为两个模型答案一致的行不携带任何关于「它们之间有差异」的信息。 有 67 行换了边; 报告出来的 −5 是 36 与 31 相抵之后剩下的量。

Why churn matters more than the delta为什么 churn 比差值更重要

If quantization were a small uniform blur, the two models would disagree only on questions sitting near a decision boundary, and the disagreements would be roughly symmetric — which is what 36 against 31 looks like. That is the benign reading, and it is consistent with three p-values above 0.69. But the same net delta could also be produced by a model that lost a whole capability and gained an unrelated one, and no aggregate score can tell the two apart. The 2×2 tables can. Reporting only the delta throws away the evidence that decides which story is true.

如果量化只是一层均匀的轻微模糊, 那两个模型只会在靠近决策边界的题目上分歧, 而且分歧大致对称—— 36 对 31 正是这个样子。 这是良性解读, 也与三个都在 0.69 以上的 p 值一致。 但同样的净差值也可以由「丢掉了一整块能力、又获得了另一块不相干能力」的模型产生, 而任何汇总分数都分辨不出这两者。 2×2 表可以。 只报告差值, 就把决定哪个故事为真的证据扔掉了。

Serving throughput服务吞吐

Three rounds at each concurrency, ISL 8192 and OSL 1024, median across rounds, with a fixed 512/128 canary before and after the whole sequence.

每个并发三轮, ISL 8192、OSL 1024, 取轮间中位数, 整段序列前后各跑一次固定的 512/128 canary。

Concurrency并发 FP8 totalFP8 总吞吐 MXFP4 totalMXFP4 总吞吐 Ratio比值 TPOT FP8 → MXFP4 Round spread轮间离散
1124.93 tok/s114.66 tok/s91.78%71.81 → 77.97 ms1.12% · 0.80%
8943.52 tok/s870.06 tok/s92.21%75.27 → 81.63 ms0.82% · 0.44%
323624.11 tok/s3350.44 tok/s92.45%75.62 → 82.08 ms0.06% · 0.18%

From flash-fp8-perf-summary.json and flash-mxfp4-perf-summary.json. Round spread is the p10–p90 band as a fraction of the median, FP8 first.

来自 flash-fp8-perf-summary.json 与 flash-mxfp4-perf-summary.json。 轮间离散是 p10–p90 带宽占中位数的比例, FP8 在前。

Plate III · Throughput against the 90% floor, and the time per output tokenISL 8192 · OSL 1024 · median of 3
MXFP4 TOTAL THROUGHPUT AS A FRACTION OF FP8 · BAR HEIGHT NORMALISED PER GROUP GATE FLOOR · 90% OF FP8 100% 0 124.93 114.66 concurrency 1 91.78% 943.52 870.06 concurrency 8 92.21% 3624.11 3350.44 concurrency 32 92.45% tok/s RATIO MEDIAN TIME PER OUTPUT TOKEN · MILLISECONDS · LOWER IS BETTER 70 75 80 85 71.81 75.27 · 75.62 77.97 81.63 · 82.08 TPOT

The bars are normalised inside each group, so what they show is the ratio, not the absolute rate — the absolute figures are printed above them and span a factor of thirty. The dot plot underneath is the same result seen per token: all three FP8 points sit below 76 ms and all three MXFP4 points above 77 ms, with no overlap.

柱子在每组内部各自归一化, 所以它们展示的是比值而不是绝对速率—— 绝对数值印在柱子上方, 三组之间相差三十倍。 下方的点图是同一个结果按每 token 来看: 三个 FP8 点全部低于 76 ms, 三个 MXFP4 点全部高于 77 ms, 没有重叠。

The loss is 7.55% to 8.22% and it barely moves with concurrency. That flatness is itself a result: an overhead that scaled with batch size would point at the matrix work, and one that vanished with batch size would point at launch cost. A constant fraction across a thirty-fold range in load points instead at something paid once per weight element per step — which is what unpacking four-bit values and applying their block scales is.

损失在 7.55% 到 8.22% 之间, 而且几乎不随并发变化。 这种平坦本身就是一个结果: 随批大小放大的开销会指向矩阵计算, 随批大小消失的开销会指向 launch 成本。 在三十倍负载范围内保持恒定比例, 指向的是「每步每个权重元素付一次」的东西—— 而解包 4 bit 值并施加它们的块 scale 正是这样的东西。

Mechanism机制

Why a checkpoint 42% smaller on disk can stream more bytes per token为什么盘上小 42% 的 checkpoint 每 token 反而读更多字节

The natural inference is that halving the numbers halves the reading, and halving the reading halves the time. Both halves of that sentence are wrong here, and they are wrong for different reasons.

很自然的推论是: 数字减半则读取减半, 读取减半则时间减半。 这句话的两半在这里都不成立, 而且不成立的原因还不一样。

A file is a census; a decode step is a sample文件是普查, 解码步是抽样

In a sparse mixture-of-experts model these are two very different quantities. The file has to contain every expert, because any of them might be routed to. A single decode step touches the shared trunk plus the handful of experts the router picked for that one token. In GLM-5.3-Flash that is 8 of 288 experts, and the expert weights are 92.7% of the file. So compressing the experts — which is where essentially all of the file lives — moves the per-token read by 8/288 of what it moves the file by. The remaining 280/288 of the saving is real, and it buys HBM capacity and load time. It does not buy decode bandwidth, because those bytes were never going to be read on this token.

在稀疏 MoE 模型里, 这是两个非常不同的量。 文件必须装下每一个 expert, 因为任何一个都可能被路由到。 而一次解码步只碰共享主干, 加上 router 为这一个 token 挑出的少数几个 expert。 在 GLM-5.3-Flash 上是 288 选 8, 而 expert 权重占文件的 92.7%。 所以压缩 expert—— 文件的绝大部分都在那里—— 对每 token 读取量的影响, 只有它对文件大小影响的 8/288。 剩下 280/288 的节省是真实的, 它买到的是 HBM 容量和加载时间, 而不是解码带宽, 因为那些字节本来就不会在这个 token 上被读。

Written out, the active set for one token is

写出来, 一个 token 的活跃权重集是

active_bytes = always_active_bytes + (top_k / num_experts) * routed_expert_bytes

which is the quantity computed for all four checkpoints in weights/, directly from the safetensors headers rather than from any config claim.

这就是 weights/ 里为四个 checkpoint 分别算出的量, 直接读 safetensors header, 而不是依据 config 里的任何声明。

The baseline is already one byte, and the exclusions are two基线已经是一字节, 而例外是两字节

This is the part the file size actively hides. The comparison is not MXFP4 against BF16, it is MXFP4 against FP8 — a baseline that already stores everything at one byte per parameter. The MXFP4 conversion applies four-bit packing plus a shared block scale to the expert tensors and leaves everything it excludes in BF16, at two bytes. Relative to the FP8 baseline the quantized parts roughly halve and the excluded parts double.

这正是文件大小主动掩盖的部分。 这里比较的不是 MXFP4 对 BF16, 而是 MXFP4 对 FP8—— 一个已经把所有东西都按每参数一字节存储的基线。 MXFP4 转换把 4 bit 打包加共享块 scale 应用到 expert 张量上, 而所有被排除的部分保留 BF16, 每参数两字节。 相对于 FP8 基线, 被量化的部分大致减半, 被排除的部分翻倍。

And the excluded set is not a random subset. Embeddings, attention projections, norms, the multi-token-prediction head — the things a converter is most reluctant to quantize — are precisely the tensors that every token reads. So the format change makes the sparse part cheaper and the dense part more expensive, and the decode step is dominated by the dense part.

而被排除的那一组并不是随机子集。 embedding、attention 投影、norm、多 token 预测头—— 转换工具最不愿意量化的那些—— 恰好就是每个 token 都要读的那些张量。 于是格式转换让稀疏部分变便宜、让稠密部分变昂贵, 而解码步恰恰由稠密部分主导。

Plate IV · What the file measures against what the step readsfrom safetensors headers
A · BYTES ON DISK INDEX TOTAL, ALL EXPERTS B · BYTES READ PER GENERATED TOKEN TRUNK + 8 OF N EXPERTS Flash FP8 Flash MXFP4 Full FP8 Full MXFP4 328.3 GB 227.5 GB · −30.71% 755.6 GB 438.0 GB · −42.04% 0 800 GB 22.41 GB 21.95 GB · −2.06% 41.38 GB 43.17 GB · +4.33% 0 48 GB the file shrank 42.04% and the per-token read grew 4.33% ALWAYS ACTIVE — TRUNK, ATTENTION, NORMS, MTP HEAD ROUTED SHARE — (top_k / num_experts) × ALL EXPERT BYTES

Panel A is what a model card advertises. Panel B is what the hardware does on every token. For the Flash pair they agree in sign and disagree by a factor of fifteen in magnitude; for the Full pair they disagree in sign outright.

A 面板是 model card 会宣传的东西。 B 面板是硬件在每个 token 上真正做的事。 Flash 这一对的符号一致但量级差了十五倍; Full 那一对连符号都相反。

CheckpointCheckpoint File文件 Always active恒定活跃 Routed share路由份额 Active per token每 token 活跃 Bytes / active param字节 / 活跃参数
Flash FP8328.327 GB13.957 GB8.458 GB22.415 GB1.245
Flash MXFP4227.486 GB16.574 GB5.379 GB21.953 GB1.220
Full FP8755.617 GB18.729 GB22.655 GB41.383 GB1.061
Full MXFP4438.002 GB31.141 GB12.032 GB43.174 GB1.107

Flash has 288 experts with top-8 routing over 45 layers; Full has 256 experts with top-8 over 78. The always-active column is where the story is: it grows 18.7% on Flash and 66.3% on Full when moving from FP8 to MXFP4, while the routed share falls 36.4% and 46.9%.

Flash 是 45 层、288 experts、top-8; Full 是 78 层、256 experts、top-8。 故事在「恒定活跃」这一列: 从 FP8 换到 MXFP4 时它在 Flash 上涨 18.7%、在 Full 上涨 66.3%, 而路由份额分别降 36.4% 与 46.9%。

Flash survives because its expert bank is enormous relative to its trunk, so the routed saving is just large enough to cover the trunk's expansion — net −2.06%. Full does not: its trunk is 78 layers deep and its expert bank, while bigger in absolute terms, is spread over 256 experts of which 8 are read. The trunk grows by 12.4 GB per token and the routed share only falls by 10.6 GB. Net +4.33%.

Flash 之所以能扛住, 是因为它的 expert 库相对于主干极大, 路由侧的节省刚好够覆盖主干的膨胀—— 净值 −2.06%。 Full 扛不住: 它的主干有 78 层深, 而 expert 库虽然绝对值更大, 却摊在 256 个 expert 上、每次只读 8 个。 主干每 token 涨了 12.4 GB, 路由份额只降了 10.6 GB。 净值 +4.33%。

And even the 2% Flash saves does not appear as speed而且 Flash 省下的那 2% 也没有变成速度

Flash MXFP4 reads 2.06% fewer bytes per token and runs 8.22% slower at concurrency 1. So the active-byte model, having correctly demolished the file-size argument, does not then get to become the new prediction. Check it against the clock: at TP8 each rank holds an eighth of the active set, 2.80 GB, and the measured TPOT is 71.81 ms. That is 39 GB/s per rank against HBM3E measured in thousands. This operating point is nowhere near the bandwidth roof — it is bound by launch and dependency latency down a 45-layer chain, and a 2% change in bytes cannot be what sets the time.

Flash MXFP4 每 token 少读 2.06% 的字节, 在并发 1 下慢 8.22%。 所以活跃字节这个模型在正确地拆掉文件大小论证之后, 并不因此就变成新的预测器。 用时钟核对一下: TP8 下每个 rank 持有活跃集的八分之一, 2.80 GB, 实测 TPOT 是 71.81 ms。 那是每 rank 39 GB/s, 而 HBM3E 的带宽以千计。 这个工作点离带宽上限还远得很—— 它受限于 45 层链路上的 launch 与依赖延迟, 2% 的字节变化不可能是决定时间的那个量。

Three models, none of which predicts the clock三个模型, 没有一个能预测时钟

File size predicts −30.71%. Active bytes predict −2.06%. The measurement says +8.22% in time. Each model is correct about the quantity it names and wrong about the one that matters, because the cost that actually dominates is not in any of them: unpacking two four-bit values from each byte and applying a per-block scale is arithmetic that FP8 does not perform, paid once per weight element per step. That is also why the penalty holds flat from concurrency 1 to 32 — it scales with weights read, and the weights are read once per step regardless of how many tokens ride along.

文件大小预测 −30.71%。 活跃字节预测 −2.06%。 实测是时间上 +8.22%。 每个模型对自己命名的那个量都是对的, 对真正重要的那个量都是错的, 因为真正主导的成本不在它们任何一个里面: 从每个字节里解出两个 4 bit 值并施加逐块 scale, 是 FP8 根本不做的算术, 而且每步每个权重元素都要付一次。 这也解释了为什么这个代价从并发 1 到 32 保持平坦—— 它随读入的权重量缩放, 而不管随行的 token 有多少, 权重每步都只读一次。

The transferable form of this, with the proper nouns deleted: compressing a sparse structure changes storage in proportion to what is stored and changes runtime in proportion to what is visited, and the two are related by the sparsity factor, not by one. Any claim that moves from the first to the second without dividing by that factor is unsupported, and if the compression is not uniform — if some tensors are excluded and stored wider than the baseline — the second can move the other way entirely.

把专有名词删掉之后, 这件事的可迁移形式是: 压缩一个稀疏结构, 对存储的影响正比于「存了什么」, 对运行时的影响正比于「访问了什么」, 而两者之间差着稀疏因子而不是 1。 任何从前者直接推到后者、中间不除以这个因子的论断都是没有支撑的; 而如果压缩不是均匀的—— 如果有些张量被排除、并且存得比基线还宽—— 后者甚至可能整个反向。

Incomplete未完成

The half that is not allowed to conclude anything不被允许下结论的那一半

The 755 GB model produced a clean, complete accuracy baseline and three performance records. The accuracy baseline is a result. The three performance records are not, and the difference is not one of degree.

那个 755 GB 的模型跑出了一份干净完整的 accuracy 基线, 和三条性能记录。 accuracy 基线是一个结果。 那三条性能记录不是, 而且这个区别不是程度上的。

What the Full FP8 baseline does sayFull FP8 基线说了什么

Dataset数据集 Score分数 Correct正确 Wall clock墙钟 Completion tokens生成 token Truncated截断
GSM8K · 131997.1948%12820:48:36257,3516 / 1319
MMLU · 50085.20%4261:19:001,051,08047 / 500
GPQA Diamond · 19843.9394%872:11:122,033,151109 / 198

From full-fp8-accuracy-summary.json. Request error rate 0 on all three. Generation was temperature=0, top_p=0.95, reasoning_effort=max, 16K max tokens on GPQA.

来自 full-fp8-accuracy-summary.json。 三项的请求错误率均为 0。 生成参数为 temperature=0、top_p=0.95、reasoning_effort=max, GPQA 上限 16K token。

The 43.94% on GPQA needs its footnote immediately, because read naively it looks like a broken model. It is not a quantization result — this is the official FP8 checkpoint, and there is no candidate to compare it to. It is a protocol result: 109 of 198 answers, 55.05%, hit the 16K output ceiling and were scored on a truncated reasoning trace. A benchmark where more than half the responses are cut off mid-thought is measuring the token budget at least as much as the model. Before the Full A/B resumes, either the ceiling moves or the benchmark is reported as a truncation rate rather than an accuracy.

GPQA 上的 43.94% 需要立刻加脚注, 因为直观读起来像一个坏掉的模型。 它不是量化结果—— 这是官方 FP8 checkpoint, 而且根本没有候选者与之对照。 它是协议结果: 198 个回答里有 109 个、也就是 55.05%, 撞上了 16K 输出上限, 是在被截断的推理轨迹上被评分的。 一个超过一半回答被拦腰截断的 benchmark, 测量 token 预算的成分至少不亚于测量模型。 在 Full 的 A/B 恢复之前, 要么把上限抬高, 要么把这一项作为截断率而不是准确率来报告。

Why the three performance records stay in the archive为什么那三条性能记录只留在归档里

A performance number in this campaign is defined by its protocol: the median of three independent rounds at a given concurrency, taken inside a window bracketed by an opening and a closing canary, with a post-run quiescence check on the node. What exists for Full FP8 is one round at two of three concurrencies, an opening canary, and no closing canary. That is not a noisier version of the defined quantity. It is a different quantity that happens to be printed in the same unit.

在这轮实验里, 一个性能数值是由它的协议定义的: 给定并发下三次独立运行的中位数, 取自一个由开场与收场 canary 夹住的窗口, 并在运行后对节点做静默检查。 Full FP8 现有的是: 三个并发里两个的一轮, 一次开场 canary, 没有收场 canary。 这不是同一个被定义的量的噪声版本, 而是恰好用同一个单位打印出来的另一个量。

And the campaign already measured why that matters. The FP8 canary on the Flash side drifted +2.30% across its own sequence. A single round carries that drift with no way to detect it, because detection is exactly what the second canary and the two extra rounds are for. So a lone c1 number could be off by a third of the effect the experiment exists to resolve, and nothing in the record would say so.

而且这轮实验已经把「为什么这很重要」测出来了。 Flash 一侧的 FP8 canary 在自己的序列里漂移了 +2.30%。 单独一轮会把这个漂移原样带进去而无从察觉, 因为「察觉」正是第二次 canary 和另外两轮存在的目的。 于是一个孤立的 c1 数值可能偏离这个实验存在的意义所要分辨的效应的三分之一, 而记录里不会有任何东西提示这一点。

Why publishing them is worse than publishing nothing为什么发布它们比什么都不发布更糟

A missing number is self-describing: anyone who needs it knows it has to be measured. A number published with a caveat loses the caveat on its first copy into a slide, and from then on it is indistinguishable from a protocol-conformant measurement — same units, same model, same hardware, same plausible magnitude. Absence is a state the next run can repair. A wrong figure circulating under this page's name is not, because there is no mechanism to recall it. So the three records live in perf/full-fp8-partial/ next to the stop marker that disqualifies them, where anyone planning the resume can read them for scheduling and no one can mistake them for a result.

缺失的数字是自我描述的: 需要它的人知道它必须被测出来。 而带着限定语发布的数字, 在第一次被复制进幻灯片时就会丢掉限定语, 从那之后它与一个符合协议的测量再也无法区分—— 同样的单位、同样的模型、同样的硬件、同样看似合理的量级。 缺失是下一次运行可以修复的状态; 一个挂着这一页名义流传出去的错误数值不是, 因为没有任何机制能把它召回。 所以那三条记录躺在 perf/full-fp8-partial/ 里, 紧挨着那个取消它们资格的停机标记, 谁要规划恢复都可以读它们来排期, 而谁也不会把它们误当成结果。

Plate V · What a performance result is made of, and which pieces existFlash · Full FP8
FLASH · EACH SIDE COMPLETE FULL FP8 STOPPED AT A CASE BOUNDARY c1 c8 c32 c1 c8 c32 CANARY BEFORE ROUND 1 ROUND 2 ROUND 3 CANARY AFTER CANARY BEFORE ROUND 1 ROUND 2 ROUND 3 CANARY AFTER 512/128 512/128 11 of 11 · median of 3 is defined 512/128 NEVER RUN 3 of 11 · median of 3 is undefined

Eleven runs per side make one publishable performance result: two canaries and three rounds at each of three concurrencies. Full FP8 stopped after three of them. The empty cells are not error bars — they are the definition of the statistic that was supposed to be reported.

每一侧十一次运行才构成一个可发布的性能结果: 两次 canary, 三个并发各三轮。 Full FP8 跑到第三次就停了。 那些空格不是误差棒—— 它们是本应报告的那个统计量的定义本身。

The same reasoning governs the delivery state. The generator that would have written the public benchmark page is built to fail closed when any of the four variants lacks complete accuracy and performance evidence, and it was not bypassed. The working branch feat/glm53-mxfp4-benchmarks has no new commit, no push and no pull request. That is not an omission; it is the mechanism doing its job.

同样的逻辑也决定了交付状态。 那个本来会写出公开 benchmark 页面的 generator, 设计上会在四个变体中任何一个缺少完整 accuracy 与性能证据时 fail closed, 而这次没有绕过它。 工作分支 feat/glm53-mxfp4-benchmarks 没有新 commit、没有 push、没有 PR。 这不是疏漏, 而是机制在正常工作。

Time时间

Where 21 hours went, and which part of it was avoidable21 小时去了哪里, 其中哪一段本可以不花

The auditable window runs from the earliest checkout artifact in the working directory to the moment the server released its last GPU lock: 21:02:18. Roughly 37% of it produced no number that survives.

可审计窗口从工作目录里最早的 checkout 工件开始, 到服务释放最后一把 GPU 锁为止: 21:02:18。 其中大约 37% 没有产出任何幸存下来的数字。

The reason to account for this precisely is that the intuitive explanation — a very large model is slow — is measurably wrong, and acting on it would fix nothing. The three formal accuracy runs generated 7,319,672 completion tokens in 7:53:58 of model time. That is real work at a real rate; the machine was not stalled. The hours that did not survive went somewhere else.

要把这笔账算清楚, 是因为直觉上的解释—— 模型太大所以慢—— 可以被测量证伪, 而按它去行动什么也修不好。 三次正式 accuracy 运行在 7:53:58 的模型时间里生成了 7,319,672 个 completion token。 那是以真实速率进行的真实工作; 机器没有卡住。 没能幸存下来的那些小时花在了别处。

Plate VI · The 21-hour window, three ways2026-08-31 07:33:49 → 09-01 04:36:07 UTC
A · CHRONOLOGY · 21:02:18 · MARKERS ARE THE TWELVE TRANSITIONS 01 02 03 04 05 06 07 08 09 10 11 12 07:33 UTC 12:00 18:00 09-01 00:00 04:36 B · THE SAME WINDOW ROLLED UP BY CATEGORY FREEZE 0:37 EXPLORATION, FAILED, INVALIDATED · 7:44 FLASH FP8 · 2:40 FLASH MXFP4 · 2:56 FORMAL FULL FP8 · 7:02 C · INSIDE THE FORMAL FULL FP8 WINDOW · 7:01:56 LOAD 0:25:37 ACCURACY 4:18:50 · 61.3% · 3,341,582 TOKENS IDLE 1:57:03 · 27.7% PARTIAL PERF 0:17:50 · STOP 0:02:24

Band A is chronological and band B is the same 21:02:18 sorted into five buckets; the two agree almost exactly, which is the check that the buckets are complete. Band C expands the largest bucket and finds that the single biggest block inside it is nothing happening at all.

A 带按时间顺序排列, B 带是同一个 21:02:18 归入五个桶后的样子; 两者几乎完全对齐, 这正是「桶是完整的」这一点的校验。 C 带展开了最大的那个桶, 发现里面最大的一块是什么都没发生。

Marker标记 UTC Transition转折
0108-31 07:33Environment and checkpoint freeze begins — SGLang, AITER, a checkout-matched kernel package, and four weight sets verified.环境与 checkpoint 冻结开始—— 核对 SGLang、AITER、与 checkout 匹配的 kernel 包, 以及四组权重。
0208:11First Flash FP8 round starts. It will be invalidated mid-run when SGLang PR #36507 force-updates from cfb1c833 to b9c0e90d.第一轮 Flash FP8 开始。 运行中 SGLang PR #36507 从 cfb1c833 force-update 到 b9c0e90d, 这一轮随之作废。
0309:25Runtime and KDA diagnosis. An accidental import of the installed kernel package 0.4.5 triggers an HSA memory aperture violation; the split-grid program-id bug is found underneath it.运行时与 KDA 诊断。 误导入已安装的 0.4.5 kernel 包会触发 HSA memory aperture violation; 在它下面找到了 split-grid program id 覆盖问题。
0410:21Second Flash FP8 round. Accuracy completes, dataset paths and the local performance seed are corrected, and it is still invalidated — the PR and AITER combination underneath it kept moving.第二轮 Flash FP8。 accuracy 跑完, 数据集路径与本地性能 seed 被修正, 但仍然作废—— 底下的 PR 与 AITER 组合还在变。
0512:30A complete pre-refresh Flash FP8 run: accuracy plus three performance rounds. Discarded when the freeze point is refreshed to the newest upstream.一次完整的 pre-refresh Flash FP8 运行: accuracy 加三轮性能。 在冻结点被刷新到最新上游时整套丢弃。
0615:15Flash MXFP4 loader fix. First load fails on a 128-versus-256 shape mismatch between packed FP4 and a BF16 exclusion inside the same fused MoE tensor.Flash MXFP4 loader 修复。 首次加载在同一个 fused MoE 张量内部的 packed FP4 与 BF16 exclusion 之间撞上 128 对 256 的 shape mismatch。
0715:55Formal Flash FP8: smoke, 1198-row accuracy, three rounds at c1/c8/c32, both canaries.正式 Flash FP8: smoke、1198 行 accuracy、c1 / c8 / c32 各三轮、双 canary。
0818:37Formal Flash MXFP4 under the identical protocol. This pair is the campaign's only complete A/B.同协议下的正式 Flash MXFP4。 这一对是整轮唯一完整的 A/B。
0921:34Formal Full FP8 launches. 25 minutes to load and warm up, then 4:18:50 of accuracy.正式 Full FP8 启动。 25 分钟加载与预热, 随后 4:18:50 的 accuracy。
1009-01 02:18Accuracy finishes and nothing starts. The pipeline has no transition from the accuracy stage to the performance stage; the server holds eight GPUs idle for 1:57:03.accuracy 结束, 而没有任何东西启动。 pipeline 没有从 accuracy 阶段到 performance 阶段的衔接; 服务把八张 GPU 空持了 1:57:03。
1104:15A manual status check notices and starts performance. Before-canary, c1-r1 and c8-r1 complete.一次人工状态检查发现了这件事并启动性能测试。 before-canary、c1-r1、c8-r1 完成。
1204:33Intermediate stop requested. c32-r1 is prevented from starting, the server exits through its cleanup trap, all eight locks release at 04:36:07.请求中间停点。 c32-r1 被阻止启动, 服务经 cleanup trap 退出, 八把锁在 04:36:07 全部释放。

Two failures, and only one of them is about scheduling两个故障, 只有一个和调度有关

The 1:57:03 of idle is the easy one. Any pipeline whose next stage is started by a person is bounded below by that person's polling interval, not by the machine's readiness. The machine finished at 02:18 and the poll arrived at 04:15. Nothing here needs a better scheduler — it needs each stage to write a completion marker and the successor to be triggered by that marker, plus a watchdog that tears the server down if no marker appears within a stage's expected time. The cost of not having that was 27.7% of the campaign's largest window, on eight MI355X.

1:57:03 的空转是容易的那个。 任何「下一阶段由人启动」的 pipeline, 其下界都是那个人的轮询间隔, 而不是机器的就绪时刻。 机器在 02:18 完成, 轮询在 04:15 到达。 这里不需要更好的调度器—— 需要的是每个阶段写一个完成标记、后继阶段由这个标记触发, 再加一个看门狗: 若在某阶段的预期时间内没有出现标记就把服务拆掉。 没有这套东西的代价, 是整轮最大窗口的 27.7%, 乘以八张 MI355X。

The 7:44 of invalidated work is the harder one, and it has a single root. The baseline was specified as "the newest upstream SGLang, AITER and HF revision" rather than as a set of commit hashes. A definition that refers to a moving target cannot be satisfied by a run that takes three hours, because the target moves during the run — and it did, when a pull request under test was force-updated to a different commit mid-measurement. Three Flash FP8 runs were completed and thrown away for this reason before the fourth was performed against pinned hashes. A freeze is a decision to stop tracking, taken once and written down; "latest" is not a freeze, it is a subscription.

7:44 的作废工作是难的那个, 而它只有一个根因。 基线被规定成「最新的上游 SGLang、AITER 与 HF revision」, 而不是一组 commit hash。 一个指向移动目标的定义, 无法被一次耗时三小时的运行满足, 因为目标会在运行期间移动—— 事实也确实如此: 一个正在被测的 pull request 在测量中途被 force-update 到了另一个 commit。 因为这个原因, 有三次 Flash FP8 完整跑完又被丢掉, 第四次才是对着固定 hash 做的。 冻结是「决定停止跟踪」, 做一次并写下来; 「最新」不是冻结, 是订阅。

Where the model time actually goes模型时间实际花在哪里

Inside the 7:53:58 of real generation, one benchmark dominates: Full FP8 on GPQA Diamond took 2:11:12 to answer 198 questions, generating 2,033,151 tokens — an average of 10,268 per question, with 109 of them running into the 16K ceiling. The cost driver is the generation protocol, not the model. reasoning_effort=max against a hard token cap turns a 198-question set into the single most expensive item in a 21-hour campaign, and it does so while producing an accuracy figure that is half a truncation-rate measurement.

在 7:53:58 的真实生成里, 有一项 benchmark 占了压倒性比重: Full FP8 在 GPQA Diamond 上用 2:11:12 回答 198 道题, 生成 2,033,151 个 token—— 平均每题 10,268 个, 其中 109 道撞上 16K 上限。 成本驱动因素是生成协议, 不是模型。 reasoning_effort=max 配一个硬 token 上限, 把一个 198 题的集合变成 21 小时实验里最贵的单项, 而且同时产出的那个 accuracy 数字有一半是截断率的测量。

Findings发现

What the runtime cost, separately from what the model scored运行时的账单, 与模型分数无关的那部分

Four of the things this campaign found have nothing to do with either checkpoint's quality. They are properties of the software underneath, and three of them would have silently corrupted the measurement if they had not surfaced.

这轮实验发现的四件事和两个 checkpoint 的质量都无关。 它们是底层软件的性质, 而其中三件如果没有浮出水面, 会无声地污染整个测量。

A split grid that overwrote its own program id一个覆盖了自己 program id 的 split grid

The KDA kernel splits one logical launch across an extra grid dimension, and each program works out which slice of the problem it owns from its own program id. If that id is recomputed after the split rather than derived from both dimensions, two programs can resolve to the same slice — and the second one to finish overwrites the first. Nothing crashes. The output is the right shape, contains plausible numbers, and is wrong in a way that only shows up if you compare it against the unsplit path on the same input. That comparison is what caught it, on GPU. The fix is SGLang PR #37250, now merged.

KDA kernel 把一次逻辑 launch 拆到一个额外的 grid 维度上, 每个 program 从自己的 program id 推出它负责哪一片。 如果这个 id 是在拆分之后重新计算的、而不是由两个维度共同推导的, 那么两个 program 可能解析到同一片—— 后完成的那个会覆盖先完成的那个。 什么都不会崩。 输出形状正确、数值看起来合理, 而它的错误只有在同一输入上与未拆分路径逐个比对时才会显现。 抓到它的正是这个比对, 在 GPU 上跑的。 修复是 SGLang PR #37250, 已合并。

The class of bug this belongs to这个 bug 属于哪一类

Every parallel decomposition carries an implicit invariant: the map from program identity to work is injective. Nothing in the type system, the launch API or the output shape checks it, and a violation produces well-formed garbage rather than a fault. That is why a correctness regression that reruns the same computation through a different decomposition is worth more than any amount of shape assertion — it is the only cheap test that can see this class of failure at all.

每一种并行分解都带着一条隐式不变量: 从 program 身份到工作的映射是单射。 类型系统、launch API 和输出形状都不检查它, 而违反它产生的是格式正确的垃圾而不是一个错误。 这就是为什么「用另一种分解重跑同一计算」的正确性回归测试, 比任何数量的形状断言都值钱—— 它是唯一能看见这一类故障的廉价测试。

A fused tensor with two dtypes in it一个内部有两种 dtype 的 fused 张量

The first Flash MXFP4 load failed on a shape mismatch, 128 against 256 on the same axis. The cause is structural rather than incidental. A fused mixture-of-experts buffer packs every expert of a layer into one tensor, and a loader slices each expert out using one dtype and one stride for the whole buffer. Four-bit packing puts two values in a byte, so a packed expert's last dimension is half the logical width; an excluded expert stored in BF16 keeps the full width. One stride cannot be right for both, and the mismatch is exactly the factor of two. The loader has to carry the exclusion set at expert granularity rather than tensor granularity — and on the Full model that same requirement reappears independently, because its fused-MoE exclusions are also per-expert. Two local commits, cc981fdebd74 and f53571897, carry the verification; neither has been pushed.

第一次加载 Flash MXFP4 时在同一个轴上撞上 128 对 256 的 shape mismatch。 原因是结构性的而不是偶然的。 一个 fused MoE 缓冲把一层的所有 expert 打包进一个张量, 而 loader 用整个缓冲统一的一种 dtype 和一个 stride 去切出每个 expert。 4 bit 打包把两个值放进一个字节, 所以被打包的 expert 最后一维只有逻辑宽度的一半; 被排除、以 BF16 存储的 expert 保持完整宽度。 一个 stride 不可能同时对两者成立, 而错位恰好就是那个 2 倍。 loader 必须把 exclusion 集合按 expert 粒度而不是张量粒度来处理—— 而在 Full 模型上同样的要求会独立地再出现一次, 因为它的 fused-MoE exclusion 同样是逐 expert 的。 两个本地 commit cc981fdebd74 与 f53571897 带着验证; 两者都没有推送。

Plate VII · One stride cannot address two dtypesfused MoE buffer
A · WHAT A SINGLE-DTYPE LOADER ASSUMES · offset = expert_index × stride E0 · 128 E1 · 128 E2 · 128 E3 · 128 E4 · 128 E5 · 128 STRIDE B · WHAT THE SHARD ACTUALLY HOLDS E0 · U8 · 128 E1 · U8 · 128 E2 · BF16 EXCLUSION · 256 E3 · U8 E4 · U8 E5 · U8 BYTES the loader indexes E3 here — the middle of E2 every expert past the exclusion is off by one packed width

The reported mismatch is 128 against 256 on one axis, which is the signature of exactly this: four-bit packing halves the stored last dimension while a BF16 exclusion keeps it whole, so the first uniform stride that crosses an exclusion is wrong and stays wrong for the rest of the buffer.

报出来的错位是同一个轴上 128 对 256, 而这正是这件事的特征: 4 bit 打包把存储的最后一维减半, 而 BF16 exclusion 保持完整, 于是第一个跨过 exclusion 的统一 stride 就错了, 并且在这个缓冲的剩余部分里一直错下去。

The model card and the metadata disagreemodel card 和 metadata 对不上

The Full MXFP4 card states that the dense and shared MLP projections remain in BF16. Reading the safetensors headers instead: decoder layers 3 through 77 store their shared-expert gate_proj, up_proj and down_proj as packed U8 with an FP4 scale tensor alongside; layers 0 through 2 have no shared-expert projections at all; the only BF16 exclusion is the MTP head at layer 78. The full per-layer readout is in full-mxfp4-shared-experts.json.

Full MXFP4 的 card 声明 dense 与 shared MLP 投影保持 BF16。 改读 safetensors header: decoder 第 3 到 77 层的 shared-expert gate_proj、up_proj、down_proj 都是 packed U8 加一个并列的 FP4 scale 张量; 第 0 到 2 层根本没有 shared-expert 投影; 唯一的 BF16 exclusion 是第 78 层的 MTP 头。 完整的逐层读数在 full-mxfp4-shared-experts.json。

This matters twice. Operationally, a loader written to the card would size the shared MLP wrong and either fault or read garbage. Analytically, it changes nothing about the conclusion in the previous section, and that is worth stating: if the card were right and those projections really were BF16, the always-active set would be larger still, and the per-token read would grow by more than 4.33%. The direction of the result does not depend on which document is correct — only its magnitude does. Config, index and safetensors headers are the source of record; the card is a claim about them.

这件事有两重意义。 在工程上, 按 card 写的 loader 会把 shared MLP 的尺寸算错, 结果不是崩就是读到垃圾。 在分析上, 它对上一节的结论没有任何改变, 而这一点值得说出来: 如果 card 是对的、那些投影真的是 BF16, 那么恒定活跃集会更大, 每 token 读取量的增幅会超过 4.33%。 结论的方向不取决于哪份文档正确, 只有幅度取决于它。 config、index 与 safetensors header 是记录源; card 只是关于它们的一个声明。

What this run is not这次运行不是什么

Boundary of the claim主张的边界

The node has no container runtime — no Docker, Podman, Apptainer or Skopeo — so no ROCm 7.2.4 MI35x image was ever pulled or run. Everything here is host ROCm 7.2.0 with a source overlay of the frozen SGLang and AITER checkouts. That is a legitimate configuration and it is what the numbers describe, but it is not the same thing as validating a released container image, and describing it that way would be a claim the evidence cannot carry.

这个节点没有容器运行时—— 没有 Docker、Podman、Apptainer 或 Skopeo—— 所以从未拉取或运行任何 ROCm 7.2.4 MI35x 镜像。 这里的一切都是宿主机的 ROCm 7.2.0, 加上冻结的 SGLang 与 AITER checkout 的源码 overlay。 这是一个合法的配置, 也正是这些数字所描述的东西, 但它和「验证一个已发布的容器镜像」不是一回事, 那样描述会是一个证据支撑不了的主张。

Frozen冻结

The point everything was measured against一切测量所对照的那个点

The Flash A/B is only a comparison because both sides ran against the same software. These are the hashes that make the claim checkable — and the reason the earlier runs were discarded is that they were taken against a definition rather than against this list.

Flash 的 A/B 之所以是一个比较, 唯一的原因是两侧跑在同一套软件上。 下面这些 hash 让这个主张可以被核对—— 而早前那几次运行被丢弃, 正是因为它们对照的是一个定义而不是这张表。

Component组件 Frozen value冻结值
Node / GPU节点 / GPUmia1-p02-g23 · 8× AMD Instinct MI355X · gfx950:sramecc+:xnack-
SGLang · Fullcc981fdebd74582b26a6e0e3c1274c910df72bc7
SGLang · Flashe122ff45e901e714f4610b1c3a535cecdd1706b4
AITERea6868a02cf54e29730a38b9a3980e4342823482
sgl-evaldb1547d6098c791ecb3576353f8a5e9d06344e7c
ROCm · torch · Triton7.2.0 · 2.9.1+rocm7.2.0.git7e1940d4 · 3.6.0
FlyDSL · sglang-kernel0.3.2.dev853 · 0.4.6.post1
Server服务TP 8 · --attention-backend dsa · --moe-runner-backend aiter · KV cache BF16 · context 65536 · --disable-radix-cache

Recorded at run time in env/environment-preflight-20260901.json, which additionally captures the resolved import paths of sglang, aiter and sgl_kernel — the check that caught an accidental import of the wrong kernel package earlier in the campaign.

运行时记录在 env/environment-preflight-20260901.json 里, 它另外还记下了 sglang、aiter、sgl_kernel 解析后的导入路径—— 正是这项检查在实验前段抓到了误导入错误 kernel 包的问题。

CheckpointCheckpoint Index total索引体积 Shards分片 Tensors张量数 RevisionRevision
zai-org/GLM-5.3-Flash328.327 GB6276,10803eb5366286a
OneNexus/GLM-5.3-Flash-MXFP4227.486 GB12072,46621e1124f735f
zai-org/GLM-5.3755.617 GB141118,629187fb9fff631
OneNexus/GLM-5.3-MXFP4438.002 GB282117,410104690ed94d4

Full revisions and the SHA-256 of each weight index are in checkpoints.json. The index hash is what makes "same checkpoint" a testable statement rather than a name.

完整 revision 与每份权重索引的 SHA-256 在 checkpoints.json 里。 索引哈希让「同一个 checkpoint」成为一句可检验的陈述, 而不是一个名字。

Reproduce复现

Reproducing this, and resuming it复现, 以及恢复

Every number on this page comes from the 39 data files in data/glm53-mxfp4-mi355x/ — the accuracy summaries and the paired comparison, the per-run performance JSONL and the derived medians, the weight-header analyses, the environment preflight and the checkpoint manifest. README.md in that directory gives the layout, the exact server and benchmark commands, and which columns are measured against which are derived.

这一页上的每个数字都来自 data/glm53-mxfp4-mi355x/ 里的 39 个数据文件—— accuracy 汇总与配对比较、逐次运行的性能 JSONL 与推导出的中位数、权重 header 分析、环境预检与 checkpoint 清单。 该目录下的 README.md 给出目录结构、精确的服务与压测命令, 以及哪些列是实测、哪些是推导。

The shape of one side, for orientation — server, accuracy, then eleven benchmark runs:

一侧的形状, 用来建立方位感—— 起服务、跑 accuracy, 然后十一次压测:

python3 -m sglang.launch_server --model-path /data/GLM-5.3-Flash-MXFP4 \
    --tp-size 8 --attention-backend dsa --moe-runner-backend aiter \
    --kv-cache-dtype bfloat16 --context-length 65536 \
    --mem-fraction-static 0.8 --chunked-prefill-size 4096 \
    --max-running-requests 32 --disable-radix-cache --trust-remote-code \
    --served-model-name glm-5.3-flash-mxfp4 --port 31102

sgl-eval run gsm8k --base-url http://127.0.0.1:31102/v1 \
    --num-threads 32 --temperature 0 --top-p 0.95 \
    --reasoning-effort max --max-tokens 8192 --seed 0

python3 -m sglang.bench_serving --backend sglang --dataset-name random \
    --random-input-len 8192 --random-output-len 1024 --random-range-ratio 1.0 \
    --max-concurrency 32 --num-prompts 128 --port 31102 \
    --output-file perf-c32-r1-i8192-o1024.jsonl

GPQA runs at --max-tokens 16384; the canaries are the same benchmark command at 512/128 and concurrency 1, once before the first round and once after the last. Reading the archive needs no GPU. Re-measuring needs eight idle ones and, at the observed rates, close to seven hours per variant.

GPQA 用 --max-tokens 16384; canary 是同一条压测命令在 512/128、并发 1 下各跑一次, 一次在第一轮之前, 一次在最后一轮之后。 读归档不需要 GPU。 重新测量需要八张空闲卡, 按实测速率算, 每个变体接近七小时。

Resuming: pilot before full恢复: 先 pilot 再全量

Finishing the original plan needs roughly seven to eight more GPU wall-clock hours: a clean re-run of Full FP8 performance into a fresh directory, then Full MXFP4 from launch through accuracy and three performance rounds. Spending that before knowing whether the Full MXFP4 loader path even works would repeat this campaign's own mistake at larger scale.

按原计划跑完还需要大约七到八个 GPU 墙钟小时: 把 Full FP8 性能干净地重跑进一个新目录, 然后 Full MXFP4 从启动、accuracy 到三轮性能。 在还不知道 Full MXFP4 的 loader 路径能不能走通之前就花掉这些时间, 等于把这轮实验自己的错误在更大的尺度上重演一遍。

  1. A 60–90 minute pilot first. Route smoke on Full MXFP4, a fixed small sample of each accuracy set, one round at each of c1, c8 and c32. This is enough to see a loader failure, an obviously broken truncation rate, or a throughput collapse — the three ways the remaining six hours could be wasted.
  2. 先跑 60–90 分钟的 pilot。 Full MXFP4 的路由 smoke、每个 accuracy 集合固定的小样本、c1 / c8 / c32 各一轮。 这足以看出 loader 故障、明显异常的截断率, 或者吞吐崩塌—— 剩下六小时被浪费掉的三种方式。
  3. Only then the full sequence, on the same frozen commits and datasets, with stages chained automatically, a completion marker written at every stage boundary, and a watchdog that stops the server if the next marker does not appear in time. That single mechanism removes the largest avoidable block in the window above.
  4. 通过之后才跑全量序列, 用同一批冻结 commit 与数据集, 阶段自动串联, 每个阶段边界写一个完成标记, 再加一个看门狗: 下一个标记没有按时出现就停掉服务。 仅这一个机制就能消掉上面那个窗口里最大的一块可避免开销。
  5. If the generation protocol changes at all — reasoning effort, token cap, sample size — both sides have to be re-run. A candidate measured under a new protocol cannot be compared against a baseline measured under the old one, and the Full FP8 accuracy above stops being reusable the moment any of those move.
  6. 只要生成协议有任何改动—— reasoning effort、token 上限、样本量—— 两侧都必须重跑。 在新协议下测出的候选者不能与在旧协议下测出的基线比较, 上面那份 Full FP8 accuracy 在这些参数中的任何一个变动的瞬间就不再可复用。

And before the length-finish gate is evaluated again, restate it in the units the instrument reports. On a 198-row set the honest form is a maximum number of additional truncations, not a percentage that falls between two attainable readings.

另外, 在再次评估 length finish 这条门槛之前, 先用仪器实际报告的单位把它重新表述一遍。 在 198 行的集合上, 诚实的形式是「最多允许多出几次截断」, 而不是一个落在两个可达读数之间的百分比。

Close收尾

What an unfinished experiment is worth一个没跑完的实验值多少

Twenty-one hours on eight MI355X produced one complete paired comparison, one unpaired accuracy baseline, a merged kernel correctness fix, two unpushed loader fixes, and a documented disagreement between a model card and the tensors it describes. It did not produce the result it set out to produce, and the honest version of the write-up says so in the first paragraph rather than the last.

八张 MI355X 上的二十一个小时, 产出了一个完整的配对比较、一个没有对照的 accuracy 基线、一个已合并的 kernel 正确性修复、两个未推送的 loader 修复, 以及一份关于「model card 与它所描述的张量对不上」的记录。 它没有产出它出发时想要的那个结果, 而诚实的写法是在第一段而不是最后一段说明这一点。

Three things generalise past this checkpoint pair. The first is that a threshold's value lies entirely in its immovability; a gate relaxed by a hundredth of itself costs nothing locally and voids every other gate on the page. The second is that a sparse model's file size and its per-token read are separated by the sparsity factor, so a compression applied only to the sparse part can shrink the file by 42% while making the step read 4.33% more — and neither number predicted the 8% the clock actually showed. The third is that the largest single line item in this campaign's budget was one hour and fifty-seven minutes of a pipeline waiting for a person to notice it had finished.

有三件事可以推广到这一对 checkpoint 之外。 第一, 门槛的价值完全在于它不可移动; 放宽自身百分之一的一道门槛在局部零成本, 却让这一页上其他所有门槛失效。 第二, 稀疏模型的文件大小与它的每 token 读取量之间差着一个稀疏因子, 所以只作用于稀疏部分的压缩可以让文件小 42%、同时让单步多读 4.33%—— 而这两个数字都没有预测到时钟上真正显示的那 8%。 第三, 这轮实验预算里最大的单项, 是一个 pipeline 等了一小时五十七分钟, 等一个人注意到它已经跑完了。

The 0.0202 pp stays on the page. It is the smallest thing measured here and the only one that says anything about how the rest of the measurements should be read.

那 0.0202 pp 留在页面上。 它是这里测到的最小的东西, 也是唯一一个能说明「其余测量该怎么读」的东西。