The research record

The source record.

Inspect the source behind each claim, its strongest qualification, and the questions still open.

21records across 5 cases

5 of 21 evidence records · DeepSeek-V3 efficiency

  1. EV-01supportsHigh confidence

    DeepSeek reports pre-training a 671B-parameter MoE on 14.8 trillion tokens and using 2.788 million H800 GPU-hours across all disclosed training stages.

    DeepSeek-AIDeepSeek-V3 Technical Report (opens in new tab)
    Primary sourceTechnical preprintPublished 27 Dec 2024
    Source details

    Quotation or reported data

    “2.788M H800 GPU hours” for the full training run.
    Exact location
    Abstract; section 1; Table 1
    Source accessed
    22 Sept 2026

    Counterevidence

    The run and benchmarks are author-reported and have not been independently reproduced end to end.

    Remaining uncertainty

    The public report does not establish acquisition dates, electricity use, or total program cost.

  2. EV-02qualifiesHigh confidence

    The widely repeated $5.576 million figure is a rental-equivalent training-compute estimate, not an all-in development cost.

    DeepSeek-AIDeepSeek-V3 Technical Report (opens in new tab)
    Primary sourceTechnical preprintPublished 27 Dec 2024
    Source details

    Quotation or reported data

    The report excludes prior research and ablation experiments from its training-cost estimate.
    Exact location
    Section 1, note immediately below Table 1
    Source accessed
    22 Sept 2026

    Counterevidence

    The narrow figure remains useful for comparing the disclosed final run if its scope is stated.

    Remaining uncertainty

    R&D, acquisition, depreciation, staff, data, networking, and failed-run costs are undisclosed.

  3. EV-03supportsModerate confidence

    On METR’s autonomy suite, DeepSeek-V3 was comparable to Claude 3.5 Sonnet (Old), while trailing newer frontier models.

    METRDetails about METR’s Preliminary Evaluation of DeepSeek-V3 (opens in new tab)
    Secondary sourceIndependent evaluationPublished 12 Feb 2025
    Source details

    Quotation or reported data

    METR reported comparable autonomy performance to Claude 3.5 Sonnet (Old), with weaker results than newer frontier models.
    Exact location
    Opening findings; General Autonomous Capabilities
    Source accessed
    03 Sept 2026

    Counterevidence

    METR used basic elicitation and warned that looping and other failures might understate or distort capability.

    Remaining uncertainty

    Benchmark performance is task- and elicitation-dependent and does not establish frontier parity.

  4. EV-04supportsHigh confidence

    Nvidia identified H800 as a product covered by the October 2023 China licensing restrictions.

    NVIDIA Corporation / U.S. SECNVIDIA Corporation Form 10-K for Fiscal 2024 (opens in new tab)
    Primary sourceRegulatory filingPublished 21 Feb 2024
    Source details

    Quotation or reported data

    Restrictions on H800 shipments took effect on 23 October 2023.
    Exact location
    Item 1, Government Regulations—Global Trade
    Source accessed
    03 Sept 2026

    Counterevidence

    DeepSeek may have acquired its H800 fleet before the restriction took effect; acquisition timing is not public.

    Remaining uncertainty

    The size and provenance of DeepSeek’s remaining controlled-hardware inventory are unknown.

  5. EV-05qualifiesHigh confidence

    The public repository enables inference and inspection of model weights but not reproduction of the full training run.

    DeepSeek-AIDeepSeek-V3 Official Repository (opens in new tab)
    Primary sourceCompany repositoryPublished 26 Dec 2024
    Source details

    Quotation or reported data

    The repository publishes model weights and links to inference support for Nvidia, AMD, and Ascend hardware.
    Exact location
    Repository README, model downloads and inference sections
    Source accessed
    22 Sept 2026

    Counterevidence

    Open weights make broad deployment and independent downstream testing possible.

    Remaining uncertainty

    Training data, the full HAI-LLM stack, and exact full-run configuration are unavailable.

“Primary” identifies an original or official source for the recorded claim. It does not mean that the finding has been independently verified. Full JSON and CSV datasets are available below.