INFERENCE LAB / BENCHMARKS
Benchmarks
Performance claims need a workload, an environment and the raw evidence.
The benchmark notebook is open. Results are not in yet.
Benchmark data pending real-world testing. We publish measured results only, with the environment and raw evidence needed to reproduce them.
What every report will include
| Measurement | Reporting requirement |
|---|---|
| TTFT | Time to first token, including measurement boundaries |
| TPOT | Time per output token; report the distribution |
| Prefill / decode throughput | Separate input processing from output generation |
| VRAM usage | Peak and steady-state memory, with measurement method |
| Workload | Concurrency, context length, output lengths and arrival pattern |
| Environment | Model, GPU, framework, quantization, TP, versions and commands |
Compare like with like
Future reports can be filtered by model, GPU, framework, precision, tensor parallelism, context and concurrency. No synthetic scores or unsupported rankings are shown.
Read our engine evaluation method →