Fastest - Higher is better
Median tokens per second while the model is generating. Higher is better.
STAC benchmark
The recorded workload* is a long-context, long-trajectory agentic system in which a DeepSeek LLM iteratively develops, debugs, and evaluates trading models.
Median tokens per second while the model is generating. Higher is better.
Median time until the first token. Lower is better.
Aggregate output tokens per second for the whole run. Higher is better.
Median time to finish a request. Lower is better.
Median prompt processing speed. Higher is better.
Median gap between output tokens. Lower is better.
Click a column title to sort. Click a row to expand p90/p99 spreads. Missing values stay at the bottom.
| 309.1 | 1.10 | 2.09 | 231.5 | n/a | |
| 247.5 | 0.799 | 1.57 | 148.4 | 98.5 | |
| 172.8 | 2.25 | 5.89 | 132.2 | 92.3 | |
| 83.5 | 2.05 | 3.63 | 61.6 | 96.8 | |
| 75.6 | 1.98 | 7.63 | 63.7 | 97.1 | |
| 55.4 | 1.76 | 5.44 | 40.3 | 96.2 |
Click two names. Click again to drop one. A third click replaces the older pick.
Select a second provider to see the spread.
*The recording is contributed by a large US market maker. An agentic coding system (Codex with DeepSeek V4 Pro) iteratively developed, debugged, and evaluated trading models for Bitcoin/USDT market data, producing 300 sequential requests. Input token count is not strictly monotonic across the conversation: it comprises three contiguous segments separated by context-checkpoint compactions, after which accumulated input resets.