Skip to content
Aevonix ResearchContact
Menu
Aevonix Research

Open source · Apache-2.0

Model recipes GLM-5.3 on RTX PRO 6000.

Each recipe is a GitHub repository that serves GLM-5.3 with TensorFold.

Recipes

4x RTX PRO 6000

Full GLM-5.3

EXL3Recipe 1.1.0TensorFold v0.6.6

Prose, single stream
92 tok/s
Prose, 5 streams
140 tok/s
Structured output, 5 streams
195 tok/s
Prefill, up to
2,081 tok/s
  • Up to 5 concurrent requests
  • 1M context
View on GitHub
4x RTX PRO 6000

GLM-5.3-Flash

EXL3Recipe 1.2.0TensorFold v0.6.6

Prose, single stream
266 tok/s
Prose, 16 streams
1,087 tok/s
Structured output, 16 streams
2,054 tok/s
Prefill, up to
9,135 tok/s
  • Up to 40 concurrent requests
  • 1M context
View on GitHub
2x RTX PRO 6000

GLM-5.3-Flash

EXL3Recipe 1.1.0TensorFold v0.6.6

Prose, single stream
238 tok/s
Prose, 4 streams
449 tok/s
Structured output, 4 streams
651 tok/s
Prefill, up to
6,908 tok/s
  • Up to 4 concurrent requests
  • 512K context
View on GitHub

01

Run a recipe

From a recipe's repository:

./start.sh

One command builds, downloads and serves.

03

Credits

TensorFold by Ash Hart.

Built in collaboration with Mia's AI Lab (EXL3 checkpoints and base patches).

License
Apache-2.0