Open recipes · NVIDIA DGX Spark · GB10

Frontier models on a stack of Sparks.

Reproducible serving recipes for GB10 clusters: kernels, overlays and boot scripts, with every number measured on real hardware and published alongside the code.

Recipes
–
Fastest decode
–
Longest context
–

Recipes

Loading recipes…

How we measure

Decode

Single-stream generation speed in tokens per second, split by workload. Code and structured output draft well, so speculative decoding helps them most. Prose is the hard case.

Prefill

Prompt processing on a cold request with no prefix-cache hits, at the context length shown. This sets time to first token on long prompts.

Under load

Aggregate tokens per second across concurrent streams. Per-request speed drops as streams are added; total throughput rises until the cluster saturates.

Live from the repo

Each card reads kindling.json from its recipe repository, so the numbers here change when the recipe does. Benchmark scripts live in each repo.