Skip to content
Aevonix ResearchContact
Menu
All benchmark runsServing performance / Model endpoint

GLM-5.3 Flash

stock · Warm-engine unique-prefix pilot; about 4K input, client C1

completedPerformanceActual inference
glm53-flash-stock-serving-4k-c1-pilot-119f8a2564a2
Observed outcomes

6 / 8 observed passes

6 / 8 verified primary passes. Every declared case stays in the denominator; unverified attribution does not mean an observed task failed.

2fail
6pass

End-to-end completion throughput

21.183 tok/s

Server-reported completion tokens divided by the entire measured batch duration, including prefill, waiting and failed requests. Token accounting is stated in the serving cell.

8 samples

First generated output, median

2,563.189 ms

First nonempty reasoning or content delta; this may precede any visible answer. Includes every response with that timing; missing timings are not zero.

8 samples

First final content, median

3,322.551 ms

First nonempty final-content delta; arrival does not establish correctness or usefulness. Includes every response with that timing; missing timings are not zero.

8 samples

Request completion, median

4,155.302 ms

Full request duration, including reasoning; includes failures with recorded duration. Includes every response with that timing; missing timings are not zero.

8 samples

Request completion, p95

6,909.581 ms

Linear-interpolated p95 over all recorded request durations, including failures. Small samples give an unstable tail estimate.

8 samples

Stream chunk interval, median

87.761 ms

Observed SSE intervals reported by the serving client. Speculative decoding can emit several tokens per chunk; this is not per-token latency.

198 samples

Measured batch duration

37.153 s

Client measurement interval, excluding process setup. This is the denominator for aggregate completion throughput.

1 samples

Case results

Observed outcome and primary-model attribution are separate.

request-01Synthetic request 01pass
Primary outcome
pass
Failure stage
Not recorded
Duration
3,923.377 ms
request-02Synthetic request 02fail
Primary outcome
fail
Failure stage
answer_checks
Duration
6,931.26 ms
request-03Synthetic request 03pass
Primary outcome
pass
Failure stage
Not recorded
Duration
3,801.035 ms
request-04Synthetic request 04pass
Primary outcome
pass
Failure stage
Not recorded
Duration
4,393.57 ms
request-05Synthetic request 05fail
Primary outcome
fail
Failure stage
answer_checks
Duration
4,387.226 ms
request-06Synthetic request 06pass
Primary outcome
pass
Failure stage
Not recorded
Duration
3,203.412 ms
request-07Synthetic request 07pass
Primary outcome
pass
Failure stage
Not recorded
Duration
3,635.69 ms
request-08Synthetic request 08pass
Primary outcome
pass
Failure stage
Not recorded
Duration
6,869.321 ms

Deployment and test conditions

Configured capacity is distinct from successfully tested context or concurrency.

Model
GLM-5.3 Flash
Checkpoint revision
eb9eb208eb0d988989d07a6a12d0fdeb5f52574a
Variant
stock
Quantization
Native FP8 weights; FP8 E4M3 KV cache
Hardware
Four NVIDIA DGX Spark systems, GB10, 128 GB unified memory per system
Node count
4
Runtime
SGLang
Runtime version
0.0.0.dev1+gd6ab04bdf1
Configured context
262144
Configured request slots
8
Profile
Warm-engine unique-prefix pilot; about 4K input, client C1
Weight identity
Verified; see evidence
Suite version
serving-performance-v1
Started
Sep 20, 2026, 2:31 PM UTC
Finished
Sep 20, 2026, 2:32 PM UTC
Input-size target
4096
Observed input-token range
3704–3735
Client concurrency limit
1
Observed peak request overlap
Not reported
Cache condition
Warm engine, unique prefixes
Transport completions
8 / 8
Completion-token accounting
Reasoning and final content combined
Stream intervals
SSE chunks, not individual tokens
Timing observer SHA-256
8a29437cb3e4b806f4be78454483721e9c31b09a7063f55bd9ec50081ffc3bef
Dataset SHA-256
53ddfba9e794833f0ddb92f9785dc1c012c566ef6d438df0b4dd153ca074860f
Correctness scorer
frozen-evidence-fields-v1
Peak-overlap basis
Not recorded; serving-client peak estimates are excluded
Runtime image digest
sha256:f5de20f9b4a806a94012d405f5cb8734c129eb69a7b3875d5362210527bad575
Weight verification
All checkpoint files rehashed against the pinned artifact on all four replicas
Serving layout
TP4, four hosts; prefill chunk 2048; scheduler cap 8; static memory fraction 0.80
Observed KV token pool
1204416 tokens total; not eight full-length sessions
Parsers
glm45 reasoning; glm47 tools
Speculation
Adaptive NEXTN: launch 5 steps/6 draft tokens; observed 3/4 at readiness
Thinking request
High reasoning; thinking enabled
Request settings
Temperature 0; seed 20260920; high reasoning; completion cap 4096 tokens; streaming with server usage
Input preparation
Identical frozen UTF-8 documents across tokenizers; approximate 4K target;8 independent development tasks
Cache state
Unique document prefixes first used on an already warm engine; no warmup requests or cache flush
Measurement path
CPU-only client over local network HTTP; no SSH tunnel in the measured inference path
Available scorer identity
Named scorer version; original grading-script source hash was not recorded

Provenance

Suite SHA-256
7604be858e8b928dd7c2acb188dd25ec8b46bf26d544872220e18c20344db1d6
Evaluator SHA-256
1f2d6733c17c647cba8599f8b560fdb2529b2fc9e710f8dd6ad71c68d606bf40
Comparison conditions SHA-256
Comparability not established

No additional public evidence links were attached. The run JSON contains all published fields.

Public records exclude production conversations, private memories, network endpoints and hidden model reasoning.