Skip to content
Aevonix ResearchContact
Menu
Aevonix Research

Model benchmarks

Open models tested for reliable agent work. Compare the shared screening results, inspect hardware and serving recipes, and see where each model still fails.

Updated Sep 20, 2026, 4:26 PM UTCDevelopment results; full qualification ongoing

Loading model results…

Loading benchmark filters…

Serving-performance gridMeasured context and request-load cells, separate from quality scoresMeasurements
Serving measurements

Context and concurrent requests

Speed and answer correctness are shown together. These are measured workloads, not maximum-capacity claims.

Context rows are approximate input-size targets. Columns are client concurrency limits, separate from server slots and observed overlap. Open a run for actual token counts, timings and limitations.

GLM-5.3 Flash stock

Warm-engine unique-prefix pilot; about 4K input, client C1

Four NVIDIA DGX Spark systems, GB10, 128 GB unified memory per system · Warm engine, unique prefixes

Input targetC1Client limitC4Client limitC8Client limit
4K
6/8 correct8/8 transport completed6/8 verified primary21.18 tok/sIncludes reasoning + final tokens
First generated
2.56 s
First final content
3.32 s
Median request timings3,704–3,735 actual input tokensPeak overlap not measured
Not measuredNot measured
32KNot measuredNot measuredNot measured
128KNot measuredNot measuredNot measured

First generated output may be hidden reasoning. First final-content arrival is not proof of a useful or correct answer. Throughput spans the full measured batch. Speculative streaming intervals describe chunks, not individual tokens.

Testing the agent systemWhat memory, shared state and learning add to the same modelMethods
System contribution

Protagine's contribution

The same model can behave differently when memory, shared tasks and recollection are available.

Matched comparisons

For the finalists, compare base Hermes with Hermes plus Protagine on 48 state-dependent scenarios. A further 12-scenario control withholds relevant memories or substitutes unrelated ones.

Learning beyond repetition

Provide the same experience and feedback, close the session, then test untouched task variants. Repeat a subset with a different processor. Measure useful transfer and regressions separately.

These are planned system-contribution experiments. Endpoint or extraction tests alone do not demonstrate persistent-agent behavior. Learning uses Protagine's ordinary pathways, with invented people and isolated state.

About Protagine
Methodology and test inventoryBoundaries, scoring, limitations and the 180-scenario planProtocol
Protocol v1.0

Methodology and test inventory

180 planned scenarios. A scenario may span several turns, tools or sessions. This is an agent-work benchmark, not an AGI score.

Three distinct boundaries

Model endpoint, native Hermes, and Hermes plus Protagine. Consumer tests remain labeled separately. A valid JSON response is not evidence that the stored fact is correct.

Development and held-out tests

A 24-case screen comes from 120 development scenarios. Freeze settings before opening 60 held-out scenarios. Three repetitions measure consistency, not three times as many independent tasks.

Scoring and uncertainty

Check tool effects, memory changes and code with executable oracles first. Free-form answers use blind semantic rubrics. Publish sample counts, disagreements and task-level uncertainty when measured.

Practical serving performance

Report queue delay, first useful content, completed tasks and latency under mixed load. Separate cold prefill from warm caches, single-request speed from aggregate throughput, and configured context from demonstrated usable context.

Fair comparisons

Pin fixtures, supporting models, runtime and budgets. Record exact checkpoints and quantizations. Modified weights are a separate deployment. Keep failed configurations and distinguish model errors from invalid tests.

Public evidence, private deployment

Publish named result fields and synthetic examples. Owner conversations, memories, credentials, internal network details and private reasoning are excluded. Downloadable artifacts identify what was actually tested.

View the 180-scenario test list
G01–G12

Grounding

Evidence, uncertainty, false premises and honest completion claims.

12
T01–T16

Tools and output contracts

Correct actions, arguments, structured output, streaming and recovery.

16
F01–F16

Memory formation

Useful facts, attribution and temporal scope, without admitting fiction or clutter.

16
R01–R20

Recollection and corrections

Automatic recall, updated facts, forgetting and a different model in a new session.

20
L01–L12

Long context

Evidence use, contradictions, synthesis and compression, beyond finding a hidden phrase.

12
P01–P12

Planning and execution

Dependencies, changed constraints, delegation and verified final task state.

12
U01–U12

Unified state

Concurrent channels, shared commitments, steering and durable task ownership.

12
A01–A16

Identity and authority

Synthetic contacts, legitimate access, impersonation and private-data boundaries.

16
I01–I08

Identity, opinions and style

Complete concise answers, justified disagreement and stable state across processors.

8
E01–E12

Experience and learning

Feedback applied to unseen variants, delayed recall, transfer and regressions.

12
C01–C12

Coding

Small and multi-file repairs, instructions, executable tests and accurate reports.

12
V01–V12

Vision and voice

Grounded interpretation, ambiguity, interruption and conversation continuity.

12
O01–O08

Operational behavior

Timeouts, bounded retries, fallback attribution and isolated service recovery.

8
X01–X12

Interactive multitasking

Foreground conversation and steering while substantial background work continues.

12

Recipes, case-set hashes, evaluator versions and limitations accompany each published run. Snapshots update after test batches; this page does not expose a live inference endpoint. Historical measurements elsewhere on the site are not scores in this suite.