DeepSeek-V3 on Nvidia H800
Algorithm and systems co-design stretched a restricted H800 fleet, but did not replace the hardware chokepoint.
Capability restored
DemonstratedA 671B-parameter mixture-of-experts model delivered strong author-reported benchmark results; on METR’s autonomy suite it was comparable to Claude 3.5 Sonnet (Old) while trailing newer frontier models.
Repeatable scale
PartialOne 2,048-H800 training run is documented. Weights and inference code are public, but the full training stack and dataset are not.
Extra resources
MaterialThe disclosed 2.788 million GPU-hours exclude prior research, ablations, data, staff, acquisition, energy, and failed work. Low-level co-design was substantial; neither this total nor the engineering effort isolates the extra cost caused by controls.
Outside dependency
HighTraining relied on Nvidia H800 accelerators. Hardware replenishment is the relevant foreign-controlled dependency; using foreign-origin open-source software alone does not establish an enforceable chokepoint.
Enforcement resilience
MixedAlgorithms and weights diffuse readily; installed chips cannot be recalled. Future hardware refresh, HBM, and accelerator replenishment remain enforceable chokepoints.
Next generation
UnclearThis run establishes useful 2024-era capability. It does not by itself demonstrate independent, repeatable access to the hardware needed for later generations.
Still unresolved
- H800 acquisition dates and inventory
- Total program and electricity cost
- Full-run reproducibility
- Hardware used for later DeepSeek models
- Whether the efficiency choices were caused by export controls