INFERENCE LAB / BENCHMARKS

Benchmarks

Performance claims need a workload, an environment and the raw evidence.

The benchmark notebook is open. Results are not in yet.

Benchmark data pending real-world testing. We publish measured results only, with the environment and raw evidence needed to reproduce them.

What every report will include

MeasurementReporting requirement
TTFTTime to first token, including measurement boundaries
TPOTTime per output token; report the distribution
Prefill / decode throughputSeparate input processing from output generation
VRAM usagePeak and steady-state memory, with measurement method
WorkloadConcurrency, context length, output lengths and arrival pattern
EnvironmentModel, GPU, framework, quantization, TP, versions and commands

Compare like with like

Future reports can be filtered by model, GPU, framework, precision, tensor parallelism, context and concurrency. No synthetic scores or unsupported rankings are shown.

Read our engine evaluation method →