Open recipes · NVIDIA DGX Spark · GB10
Frontier models on a stack of Sparks.
Reproducible serving recipes for GB10 clusters: kernels, overlays and boot scripts, with every number measured on real hardware and published alongside the code.
- Recipes
- –
- Fastest decode
- –
- Longest context
- –
Recipes
Loading recipes…
How we measure
Decode
Single-stream generation speed in tokens per second, split by workload. Code and structured output draft well, so speculative decoding helps them most. Prose is the hard case.
Prefill
Prompt processing on a cold request with no prefix-cache hits, at the context length shown. This sets time to first token on long prompts.
Under load
Aggregate tokens per second across concurrent streams. Per-request speed drops as streams are added; total throughput rises until the cluster saturates.
Live from the repo
Each card reads kindling.json from its recipe repository, so the numbers here change when the recipe does. Benchmark scripts live in each repo.