Open models tested for reliable agent work. Compare the shared screening results, inspect hardware and serving recipes, and see where each model still fails.
Serving-performance gridMeasured context and request-load cells, separate from quality scoresMeasurements
Serving measurements
Context and concurrent requests
Speed and answer correctness are shown together. These are measured workloads, not maximum-capacity claims.
Context rows are approximate input-size targets. Columns are client concurrency limits, separate from server slots and observed overlap. Open a run for actual token counts, timings and limitations.
GLM-5.3 Flash stock
Warm-engine unique-prefix pilot; about 4K input, client C1
Four NVIDIA DGX Spark systems, GB10, 128 GB unified memory per system · Warm engine, unique prefixes
Input target
C1Client limit
C4Client limit
C8Client limit
4K
6/8 correct8/8 transport completed6/8 verified primary21.18 tok/sIncludes reasoning + final tokens
First generated
2.56 s
First final content
3.32 s
Median request timings3,704–3,735 actual input tokensPeak overlap not measured
Not measured
Not measured
32K
Not measured
Not measured
Not measured
128K
Not measured
Not measured
Not measured
First generated output may be hidden reasoning. First final-content arrival is not proof of a useful or correct answer. Throughput spans the full measured batch. Speculative streaming intervals describe chunks, not individual tokens.
Testing the agent systemWhat memory, shared state and learning add to the same modelMethods
System contribution
Protagine's contribution
The same model can behave differently when memory, shared tasks and recollection are available.
Matched comparisons
For the finalists, compare base Hermes with Hermes plus Protagine on 48 state-dependent scenarios. A further 12-scenario control withholds relevant memories or substitutes unrelated ones.
Learning beyond repetition
Provide the same experience and feedback, close the session, then test untouched task variants. Repeat a subset with a different processor. Measure useful transfer and regressions separately.
These are planned system-contribution experiments. Endpoint or extraction tests alone do not demonstrate persistent-agent behavior. Learning uses Protagine's ordinary pathways, with invented people and isolated state.
About ProtagineMethodology and test inventoryBoundaries, scoring, limitations and the 180-scenario planProtocol
Protocol v1.0
Methodology and test inventory
180 planned scenarios. A scenario may span several turns, tools or sessions. This is an agent-work benchmark, not an AGI score.
Three distinct boundaries
Model endpoint, native Hermes, and Hermes plus Protagine. Consumer tests remain labeled separately. A valid JSON response is not evidence that the stored fact is correct.
Development and held-out tests
A 24-case screen comes from 120 development scenarios. Freeze settings before opening 60 held-out scenarios. Three repetitions measure consistency, not three times as many independent tasks.
Scoring and uncertainty
Check tool effects, memory changes and code with executable oracles first. Free-form answers use blind semantic rubrics. Publish sample counts, disagreements and task-level uncertainty when measured.
Practical serving performance
Report queue delay, first useful content, completed tasks and latency under mixed load. Separate cold prefill from warm caches, single-request speed from aggregate throughput, and configured context from demonstrated usable context.
Fair comparisons
Pin fixtures, supporting models, runtime and budgets. Record exact checkpoints and quantizations. Modified weights are a separate deployment. Keep failed configurations and distinguish model errors from invalid tests.
Public evidence, private deployment
Publish named result fields and synthetic examples. Owner conversations, memories, credentials, internal network details and private reasoning are excluded. Downloadable artifacts identify what was actually tested.
View the 180-scenario test list
G01–G12
Grounding
Evidence, uncertainty, false premises and honest completion claims.
12
T01–T16
Tools and output contracts
Correct actions, arguments, structured output, streaming and recovery.
16
F01–F16
Memory formation
Useful facts, attribution and temporal scope, without admitting fiction or clutter.
16
R01–R20
Recollection and corrections
Automatic recall, updated facts, forgetting and a different model in a new session.
20
L01–L12
Long context
Evidence use, contradictions, synthesis and compression, beyond finding a hidden phrase.
12
P01–P12
Planning and execution
Dependencies, changed constraints, delegation and verified final task state.
12
U01–U12
Unified state
Concurrent channels, shared commitments, steering and durable task ownership.
12
A01–A16
Identity and authority
Synthetic contacts, legitimate access, impersonation and private-data boundaries.
16
I01–I08
Identity, opinions and style
Complete concise answers, justified disagreement and stable state across processors.
8
E01–E12
Experience and learning
Feedback applied to unseen variants, delayed recall, transfer and regressions.
12
C01–C12
Coding
Small and multi-file repairs, instructions, executable tests and accurate reports.
12
V01–V12
Vision and voice
Grounded interpretation, ambiguity, interruption and conversation continuity.
12
O01–O08
Operational behavior
Timeouts, bounded retries, fallback attribution and isolated service recovery.
8
X01–X12
Interactive multitasking
Foreground conversation and steering while substantial background work continues.
12
Recipes, case-set hashes, evaluator versions and limitations accompany each published run. Snapshots update after test batches; this page does not expose a live inference endpoint. Historical measurements elsewhere on the site are not scores in this suite.