Live evaluation · Terminal-Bench 4.0

GLM-5.3 FP8

The full zai-org/GLM-5.3 block-FP8 weights, 743B total and 39B active, served on one 8×MI355X node at TP=8. Same instrument and same 63-task denominator as the MXFP4 run, so the two boards are directly comparable. Scored model outcomes stay separated from infrastructure faults.

完整的 zai-org/GLM-5.3 block-FP8 权重, 总参数 743B、 激活 39B, 在单节点 8×MI355X 上以 TP=8 服务。 仪表与 63 任务分母都与 MXFP4 那一轮相同, 因此两块看板可直接比较。 模型得分结果与基础设施故障分开记录。

0%scored已计分

Attempt ledger尝试台账

The denominator is fixed at 63 tasks × 5 attempts: Terminal-Bench 4.0 ships 66, and three of them declare gpus = 1, which this node cannot honour while all eight GPUs are serving the model under test. An attempt advances this bar only when a verifier returns a reward; infrastructure failures remain visible but do not consume the target.

分母固定为 63 个任务 × 5 次尝试: Terminal-Bench 4.0 共 66 个任务, 其中 3 个声明了 gpus = 1, 而本节点的 8 张卡都在为被测模型提供服务, 无法满足。 只有 verifier 返回 reward 的尝试才推进这条进度; 基础设施故障保留记录, 但不占用目标次数。

0 / 315—
Scored attempts已计分尝试—
of 315 fixed target固定目标 315
Complete tasks完整任务—
five scored attempts each每项均有 5 次计分
Observed rate实测速率—
terminal trials per hour终态 trial / 小时
Main-run ETA主跑预计剩余—
at observed rate, regrade excluded按实测速率, 不含 regrade

Verified correctnessVerifier 正确率

Binary task rewards · live二元任务 reward · 实时
Current scored attempts当前已计分尝试 — —
Completed-task matched完整任务同项比较
—FP8 current
—Official API
—
Official same 63 CPU tasks官方同一组 63 个 CPU 任务 — —
Official full 66-task row官方完整 66 项结果 41.82% 138 / 330 · 95% CI ±3.23 pp
The comparison always uses a task-matched denominator. Before completion, the all-attempt rate is biased by task completion order and the matched card uses only finished tasks. At 63/63 complete it becomes the final CPU-subset comparison. The official row used an opaque hosted API; this run serves the published block-FP8 weights on our own SGLang stack, so the line is an external end-to-end reference, not a same-runtime A/B. 比较始终使用同 task 分母。 完成之前, 全部 attempt 比例会受到任务完成顺序的偏置, 同项卡片只使用已完成 task; 达到 63/63 后, 它才成为最终 CPU 子集比较。 官方结果来自不透明的 hosted API, 当前运行在我们自己的 SGLang 栈上服务官方 block-FP8 权重, 因此这是一条端到端外部参考线, 不是同 runtime 的 A/B。
Task任务 Current FP8当前 FP8 Official GLM-5.3 API官方 GLM-5.3 API Δ when complete完成项差值 State状态
No tasks match this view.没有符合条件的任务。

Serving pools模型服务池

One 8-GPU replica · max-running-requests 64单个 8 卡副本 · max-running-requests 64
Replica A · TP=8—
Terminal终态—
Running运行中—
Pending等待—
Retries重试—
Active requests活跃请求—
Waiting / KV tokens等待 / KV token—

Issue ledger问题台账

Faults are not silently scored故障不会被静默计分

Completion path完成路径

One owner per attempt每次尝试只有一个归属
Evaluation completion flow 66 TASKS LESS 3 GPU TASKS ONE TP=8 REPLICA 8x MI355X · FP8 KV CURRENT RUN REPAIR / REGRADE ONLY MISSING REWARDS 315 SCORED

Why the count is not the Harbor terminal count. A model outcome may score zero and still count. A container or verifier fault without a reward does not count and is replayed from the narrowest durable boundary.

为什么这里不直接使用 Harbor 的终态数量。 模型结果即使得 0 分也可以计数; 没有 reward 的容器或 verifier 故障不计数, 并从最窄的持久化边界补跑。

What this public view excludes. Task prompts, agent trajectories, internal addresses, credentials, per-request content, and unpublished benchmark artifacts never enter the feed.

公开看板不会包含什么。 Task prompt、agent trajectory、内部地址、凭据、逐请求内容以及未发布的 benchmark artifact 都不会进入数据源。