Prompt processing
tokens/s
Choose comparisons
Loading…
Models and toolboxes
Models
Toolboxes
Choose models and toolboxes
Charts and URL match this selection.
Benchmark tuning
Calibrated settings
Chart view
Text generation
tokens/s
Generation throughput
Exact measurements
Selected results
Each point measures a 2,048-token prefill or 128-token generation after the listed starting depth. The KV-cache fill is not timed. Values after ± are standard deviations; percentages compare against the selected baseline toolbox.
Reproducibility