DeepSeek-V4-Flash-DSpark · vLLM 0.26 SM121 (r0b0tlab)
Full reproducibility package + optimized container for dual NVIDIA GB10 (DGX Spark / SM121) serving of DeepSeek-V4-Flash-DSpark on vLLM 0.26.0+dspark.sm121.2.
Maintained by r0b0tlab.
Not an official DeepSeek or vLLM release. Model weights are not redistributed.
Headline results
| Metric | Result | Prior floor |
|---|
| Dedicated c1 decode median (2048 tok × 5) | 80.58 tok/s | 77.76 (v0.25 K6) |
| Static c16 aggregate best-of-2 | 402.0 tok/s | 342.70 (v0.25) |
| Exact ~300K long request | QUALIFIED | — |
| Staggered c16 | 16/16 | — |
| GSM8K smoke N=50 | 50/50 | smoke only |
Quick links
Container
docker pull ghcr.io/r0b0tlab/deepseek-v4-flash-dspark-v026-sm121:v0.26.0-sm121-optimized
# digest sha256:269a1c6ac581f4fb2fe13110d2468159c3ca65c07ffb09b4080549673f625841
- Local build ID:
sha256:f904b9445aa407420b26f70ee5c2075c7d04ef0fdd38a728b20985ee331303a6
- Details: CONTAINER.md
Runtime identity
vllm 0.26.0+dspark.sm121.2
kv_cache_dtype nvfp4_ds_mla
moe_backend flashinfer_b12x
mtp dspark K=6
max_model_len 327680
model_revision 913f0657a874f76844e2e91cbe706dbcaceeb6d7
upstream_vllm 568afb3a13806beb53bb2e6bd518269357b237c0
Profiles (do not mix rows)
- Production decode —
max_num_seqs=16, max_num_batched_tokens=16384, K=6
- C1 long-context —
max_num_seqs=1, max_num_batched_tokens=8192, K=6
Benchmark results
Raw harness outputs live under evidence/benchmarks/ (c1 decode JSON/logs, concurrent ladder, staggered c16, correctness, GSM8K smoke).
What's included for reproducibility