DeepSeek reports pre-training a 671B-parameter MoE on 14.8 trillion tokens and using 2.788 million H800 GPU-hours across all disclosed training stages.
DeepSeek-AIDeepSeek-V3 Technical Report (opens in new tab)Source details
Quotation or reported data
“2.788M H800 GPU hours” for the full training run.
- Exact location
- Abstract; section 1; Table 1
- Source accessed
- 22 Sept 2026
Counterevidence
The run and benchmarks are author-reported and have not been independently reproduced end to end.
Remaining uncertainty
The public report does not establish acquisition dates, electricity use, or total program cost.