STAC benchmark

DeepSeek V4 Pro - providers

Recorded workload

The recorded workload* is a long-context, long-trajectory agentic system in which a DeepSeek LLM iteratively develops, debugs, and evaluates trading models.

Median input tokens
172K
Peak input tokens
334K
Turns
300

Fastest - Higher is better

Median tokens per second while the model is generating. Higher is better.

309.1
247.5
172.8
83.5
75.6
55.4
Lambda 2xB200CrusoeNebiusDeepInfraDeepSeekDigitalOcean

Lowest latency - Lower is better

Median time until the first token. Lower is better.

799
1.10
1.76
1.98
2.05
2.25
CrusoeLambda 2xB200DigitalOceanDeepSeekDeepInfraNebius

Highest throughput - Higher is better

Aggregate output tokens per second for the whole run. Higher is better.

231.5
148.4
132.2
63.7
61.6
40.3
Lambda 2xB200CrusoeNebiusDeepSeekDeepInfraDigitalOcean

Request latency - Lower is better

Median time to finish a request. Lower is better.

1.57
2.09
3.63
5.44
5.89
7.63
CrusoeLambda 2xB200DeepInfraDigitalOceanNebiusDeepSeek

Prefill speed - Higher is better

Median prompt processing speed. Higher is better.

84.5k
3.1k
1.1k
1.0k
901.6
403.3
Lambda 2xB200NebiusCrusoeDigitalOceanDeepSeekDeepInfra

Inter-token latency - Lower is better

Median gap between output tokens. Lower is better.

0.06
18.3
20.7
26.0
28.6
37.0
DeepSeekCrusoeNebiusDigitalOceanLambda 2xB200DeepInfra

Comparison

Click a column title to sort. Click a row to expand p90/p99 spreads. Missing values stay at the bottom.

Columns
Provider comparison for the STAC benchmark

Compare two providers

Click two names. Click again to drop one. A third click replaces the older pick.

Select a second provider to see the spread.

*The recording is contributed by a large US market maker. An agentic coding system (Codex with DeepSeek V4 Pro) iteratively developed, debugged, and evaluated trading models for Bitcoin/USDT market data, producing 300 sequential requests. Input token count is not strictly monotonic across the conversation: it comprises three contiguous segments separated by context-checkpoint compactions, after which accumulated input resets.