You are the inference engineer on a team that serves open-weight language models. Finance wants to know how many GPUs next quarter needs, and product wants 32k-token context on the largest model. Both answers come from the same place: the KV cache. Every request keeps a key and a value vector for every token, in every layer, for as long as it is running, and every decode step reads all of them back from GPU memory along with every weight. You have 16,000 benchmark runs across 48 model configurations and 5 GPU classes. Your job is to compute the cache from the model shape, work out what fits, prove the cache is exact, and build a throughput model the capacity planners can trust.
Count the runs, classify each model's attention layout, and measure how wide the throughput range is.
Each row is one benchmark run: a decoder-only model (n_layers, d_model, n_heads, n_kv_heads, head_dim, params_b), stored at a precision (weight_bytes_per_param and kv_bytes_per_elem are 2 for 16-bit, 1 for 8-bit, 0.5 for 4-bit), served on one GPU (gpu_class, gpu_mem_gb, hbm_bandwidth_tbps, peak_tflops), holding batch_size concurrent requests of context_len tokens each. The label is decode_tokens_per_s: generated tokens per second summed over the whole batch.
Implement explore_runs(train_df) returning a dict with:
The spread is the reason this project grades on relative error. A model that is 50 tokens per second off is excellent on one row and useless on another.
Evaluated server-side against a hidden test set.