The research record

The source record.

Inspect the source behind each claim, its strongest qualification, and the questions still open.

21records across 5 cases

21 of 21 evidence records

  1. EV-01supportsHigh confidence

    DeepSeek reports pre-training a 671B-parameter MoE on 14.8 trillion tokens and using 2.788 million H800 GPU-hours across all disclosed training stages.

    DeepSeek-AIDeepSeek-V3 Technical Report (opens in new tab)
    Primary sourceTechnical preprintPublished 27 Dec 2024
    Source details

    Quotation or reported data

    “2.788M H800 GPU hours” for the full training run.
    Exact location
    Abstract; section 1; Table 1
    Source accessed
    22 Sept 2026

    Counterevidence

    The run and benchmarks are author-reported and have not been independently reproduced end to end.

    Remaining uncertainty

    The public report does not establish acquisition dates, electricity use, or total program cost.

  2. EV-02qualifiesHigh confidence

    The widely repeated $5.576 million figure is a rental-equivalent training-compute estimate, not an all-in development cost.

    DeepSeek-AIDeepSeek-V3 Technical Report (opens in new tab)
    Primary sourceTechnical preprintPublished 27 Dec 2024
    Source details

    Quotation or reported data

    The report excludes prior research and ablation experiments from its training-cost estimate.
    Exact location
    Section 1, note immediately below Table 1
    Source accessed
    22 Sept 2026

    Counterevidence

    The narrow figure remains useful for comparing the disclosed final run if its scope is stated.

    Remaining uncertainty

    R&D, acquisition, depreciation, staff, data, networking, and failed-run costs are undisclosed.

  3. EV-03supportsModerate confidence

    On METR’s autonomy suite, DeepSeek-V3 was comparable to Claude 3.5 Sonnet (Old), while trailing newer frontier models.

    METRDetails about METR’s Preliminary Evaluation of DeepSeek-V3 (opens in new tab)
    Secondary sourceIndependent evaluationPublished 12 Feb 2025
    Source details

    Quotation or reported data

    METR reported comparable autonomy performance to Claude 3.5 Sonnet (Old), with weaker results than newer frontier models.
    Exact location
    Opening findings; General Autonomous Capabilities
    Source accessed
    03 Sept 2026

    Counterevidence

    METR used basic elicitation and warned that looping and other failures might understate or distort capability.

    Remaining uncertainty

    Benchmark performance is task- and elicitation-dependent and does not establish frontier parity.

  4. EV-04supportsHigh confidence

    Nvidia identified H800 as a product covered by the October 2023 China licensing restrictions.

    NVIDIA Corporation / U.S. SECNVIDIA Corporation Form 10-K for Fiscal 2024 (opens in new tab)
    Primary sourceRegulatory filingPublished 21 Feb 2024
    Source details

    Quotation or reported data

    Restrictions on H800 shipments took effect on 23 October 2023.
    Exact location
    Item 1, Government Regulations—Global Trade
    Source accessed
    03 Sept 2026

    Counterevidence

    DeepSeek may have acquired its H800 fleet before the restriction took effect; acquisition timing is not public.

    Remaining uncertainty

    The size and provenance of DeepSeek’s remaining controlled-hardware inventory are unknown.

  5. EV-05qualifiesHigh confidence

    The public repository enables inference and inspection of model weights but not reproduction of the full training run.

    DeepSeek-AIDeepSeek-V3 Official Repository (opens in new tab)
    Primary sourceCompany repositoryPublished 26 Dec 2024
    Source details

    Quotation or reported data

    The repository publishes model weights and links to inference support for Nvidia, AMD, and Ascend hardware.
    Exact location
    Repository README, model downloads and inference sections
    Source accessed
    22 Sept 2026

    Counterevidence

    Open weights make broad deployment and independent downstream testing possible.

    Remaining uncertainty

    Training data, the full HAI-LLM stack, and exact full-run configuration are unavailable.

  6. EV-06supportsModerate confidence

    Huawei reports a 718B-parameter MoE run on 6,000 Ascend 910B NPUs and says its system supports all training stages.

    HuaweiPangu Ultra MoE: How to Train Your Big MoE on Ascend NPUs (opens in new tab)
    Primary sourceTechnical preprintPublished 07 May 2025
    Source details

    Quotation or reported data

    “MFU of 30.0%” on 6,000 Ascend NPUs.
    Exact location
    Paper abstract; sections 1 and 2.2
    Source accessed
    03 Sept 2026

    Counterevidence

    The run is vendor-reported and has no independent full-training audit or reproduction.

    Remaining uncertainty

    Run duration, energy, capital cost, fleet availability, and yield are not disclosed.

  7. EV-07qualifiesHigh confidence

    The model and execution plan were reshaped around NPU constraints, with cumulative interventions raising reported model FLOPs utilization (MFU) by 58.7 percent relative to the baseline.

    HuaweiPangu Ultra MoE: How to Train Your Big MoE on Ascend NPUs (opens in new tab)
    Primary sourceTechnical preprintPublished 07 May 2025
    Source details

    Quotation or reported data

    Table 5 reports a 58.7% relative MFU improvement, reaching 30.0% MFU on 6,000 NPUs; techniques include hierarchical all-to-all, recomputation, and host swapping.
    Exact location
    Section 4.5, Table 5
    Source accessed
    22 Sept 2026

    Counterevidence

    Thirty percent MFU still implies substantial unused peak compute, and no comparable counterfactual run is public.

    Remaining uncertainty

    The incremental burden caused specifically by controls cannot be separated from normal large-model engineering.

  8. EV-08supportsHigh confidence

    The official model card says Pangu Ultra was trained from scratch; the repository publishes roughly 1.47 TB of weights and documents a minimum 32-card Ascend serving setup.

    Huawei openPanguopenPangu Ultra MoE 718B Model (opens in new tab)
    Primary sourceCompany repositoryPublished date unverified
    Source details

    Quotation or reported data

    62 weight shards; four Atlas 800T A2 nodes; 32 64GB cards.
    Exact location
    Model card overview and sections 1, 4.1, and 4.4; file tree
    Source accessed
    03 Sept 2026

    Counterevidence

    Serving reproducibility does not reproduce pretraining or validate the disclosed training fleet.

    Remaining uncertainty

    The complete training corpus, code, and run configuration are not public.

  9. EV-09qualifiesModerate confidence

    BIS assessed that listed Ascend 910-series chips likely involved U.S. design technology, software, or equipment subject to export rules.

    U.S. Bureau of Industry and SecurityGuidance on Application of General Prohibition 10 (GP10) to People’s Republic of China (PRC) Advanced-Computing Integrated Circuits (ICs) (opens in new tab)
    Primary sourceGovernment guidancePublished 13 May 2025
    Source details

    Quotation or reported data

    The illustrative list names Ascend 910B and 910C.
    Exact location
    Pages 1–2, summary, illustrative list, and applicable regulations
    Source accessed
    03 Sept 2026

    Counterevidence

    This is an enforcement assessment, not a bill-of-materials audit of the Pangu fleet or CloudMatrix installation.

    Remaining uncertainty

    The precise origin and legal status of each chip and upstream component are not public.

  10. EV-10supportsHigh confidence

    CloudMatrix384 connects 384 Ascend NPUs and 192 Kunpeng CPUs through an all-to-all Unified Bus architecture.

    Huawei and SiliconFlowServing Large Language Models on Huawei CloudMatrix384 (opens in new tab)
    Primary sourceTechnical preprintPublished 19 Jun 2025
    Source details

    Quotation or reported data

    384 NPUs, 192 CPUs, disaggregated prefill, decode, and caching.
    Exact location
    Abstract; sections 3.2–3.4
    Source accessed
    03 Sept 2026

    Counterevidence

    The paper is authored by Huawei and SiliconFlow researchers and is not peer reviewed.

    Remaining uncertainty

    Vendor deployment claims now exist. Independent utilization, reliability, metered power and total cost across the reported fleet remain unverified.

  11. EV-11supportsModerate confidence

    The disclosed DeepSeek-R1 test achieved 1,943 decode tokens per second per NPU at 49.4 ms, below one profiled H800 result but above another published H800 baseline.

    Huawei and SiliconFlowServing Large Language Models on Huawei CloudMatrix384 (opens in new tab)
    Primary sourceTechnical preprintPublished 19 Jun 2025
    Source details

    Quotation or reported data

    1,943 tokens/s/NPU at 49.4 ms; reported efficiency 1.29 tokens/s/TFLOPS.
    Exact location
    Section 5.2, Tables 2–4
    Source accessed
    03 Sept 2026

    Counterevidence

    Tests used 256 of 384 NPUs; decode assumed a 70% MTP acceptance rate; batch sizes and implementations differed; and the highest prefill result assumed idealized expert load balancing.

    Remaining uncertainty

    Like-for-like cost, power, reliability, and hardware-count comparisons are unavailable.

  12. EV-12qualifiesHigh confidence

    CloudMatrix-Infer reduced batch size to meet tighter latency targets, lowering reported per-NPU throughput and illustrating the latency–throughput trade-off.

    Huawei and SiliconFlowServing Large Language Models on Huawei CloudMatrix384 (opens in new tab)
    Primary sourceTechnical preprintPublished 19 Jun 2025
    Source details

    Quotation or reported data

    Throughput fell from 1,943 to 538 tokens/s/NPU as the latency target tightened from 50 ms to 15 ms and batch size per NPU fell from 96 to 8.
    Exact location
    Section 5.2, Table 4
    Source accessed
    22 Sept 2026

    Counterevidence

    The system still met the stated tighter latency target and maintained useful throughput.

    Remaining uncertainty

    Real traffic mixes, uptime, and long-run service cost are not independently reported.

  13. EV-13qualifiesLow confidence

    A specialist estimate argues that matching a GB200 NVL72-class system requires substantially more Ascend devices, optical components, and power.

    SemiAnalysisHuawei AI CloudMatrix 384: China’s Answer to Nvidia GB200 NVL72 (opens in new tab)
    Secondary sourceSpecialist analysisPublished 16 Apr 2025
    Source details

    Quotation or reported data

    Estimated 4.1 times the system power of a GB200 NVL72.
    Exact location
    Opening comparison; CloudMatrix 384 System Architecture
    Source accessed
    03 Sept 2026

    Counterevidence

    The estimate is not metered or audited, and the compared systems differ in chip count and configuration.

    Remaining uncertainty

    No public independent energy, cooling, or total-cost measurement validates the exact multiplier.

  14. EV-14supportsHigh confidence

    A company and its owner pleaded guilty in a scheme involving at least $160 million of exported and attempted H100 and H200 shipments.

    U.S. Department of JusticeU.S. Authorities Shut Down Major China-Linked AI Tech Smuggling Network (opens in new tab)
    Primary sourceGovernment enforcementPublished 08 Dec 2025
    Source details

    Quotation or reported data

    Conduct spanned October 2024 to May 2025; more than $50 million was seized.
    Exact location
    Release summary; paragraph beginning ‘According to court documents’
    Source accessed
    03 Sept 2026

    Counterevidence

    The $160 million figure combines completed and attempted exports; as of the 8 December 2025 release, not every person described had been convicted.

    Remaining uncertainty

    The number of chips delivered, ultimate recipients, and resulting AI tasks are not disclosed.

  15. EV-15supportsModerate confidence

    The public record describes straw purchasers, third-country customers, relabeling, false paperwork, and warehousing as elements of the diversion route.

    U.S. Department of JusticeU.S. Authorities Shut Down Major China-Linked AI Tech Smuggling Network (opens in new tab)
    Primary sourceGovernment enforcementPublished 08 Dec 2025
    Source details

    Quotation or reported data

    Labels were removed and replaced with a fictitious company name.
    Exact location
    Paragraphs beginning ‘According to charging documents’ and complaint allegations
    Source accessed
    03 Sept 2026

    Counterevidence

    Several details come from charging documents and, as of the 8 December 2025 release, remained allegations against defendants who had not pleaded guilty.

    Remaining uncertainty

    The release does not quantify how common comparable networks are.

  16. EV-16qualifiesHigh confidence

    BIS identified concrete checks that could expose transshipment, opaque customers, foreign IaaS access, and implausible data-center capacity.

    U.S. Bureau of Industry and SecurityIndustry Guidance to Prevent Diversion of Advanced Computing Integrated Circuits (opens in new tab)
    Primary sourceGovernment guidancePublished 13 May 2025
    Source details

    Quotation or reported data

    Red flags include freight forwarders, undisclosed end users, foreign IaaS, and infrastructure inconsistent with ordered chips.
    Exact location
    Pages 1–5, red flags and due-diligence actions
    Source accessed
    03 Sept 2026

    Counterevidence

    Guidance does not establish that every recommended check is implemented or effective in every jurisdiction.

    Remaining uncertainty

    Detection rates and cross-border enforcement capacity are not public.

  17. EV-17supportsHigh confidence

    China continued national tax-preference eligibility for integrated-circuit production, design, materials, packaging, and major projects in 2024.

    National Development and Reform Commission of China and partner ministries2024 Integrated-Circuit and Software Tax-Preference Eligibility Notice (opens in new tab)
    Primary sourceGovernment policyPublished 22 Mar 2024
    Source details

    Quotation or reported data

    The eligibility notice covers production at 28 nm and below and multiple upstream inputs and services.
    Exact location
    Article 1 and attached eligibility criteria
    Source accessed
    03 Sept 2026

    Counterevidence

    Eligibility rules do not disclose recipients, award value, additionality, or a capability outcome.

    Remaining uncertainty

    No firm-level causal link to a reviewed AI-chip workaround is established.

  18. EV-18qualifiesModerate confidence

    A 2025 Hangzhou policy authorized annual compute vouchers and large-model training support, but no reviewed award record ties that support to DeepSeek-V3.

    Hangzhou Municipal People’s Government杭州市人民政府关于印发杭州市加快建设人工智能创新高地实施方案(2025年版)的通知 (Hangzhou AI innovation highland implementation plan, 2025 edition; translated title) (opens in new tab)
    Primary sourceGovernment policyPublished date unverified
    Source details

    Quotation or reported data

    Annual compute-voucher ceiling: CNY 250 million; qualifying model-training support: up to CNY 50 million.
    Exact location
    Appendix, sections (1) and (3), pages 11–12. Notice dated 13 June 2025; effective 15 July 2025. Exact publication date not verified.
    Source accessed
    22 Sept 2026

    Counterevidence

    The policy postdates the DeepSeek-V3 release and does not prove DeepSeek received a grant.

    Remaining uncertainty

    Recipients, disbursements, hardware origin, and capability produced are not public in the reviewed record.

  19. EV-19qualifiesHigh confidence

    GAO reports that BIS used interim-final rules partly to reduce pre-rule stockpiling, but this does not establish the scale of any Chinese inventory response.

    U.S. Government Accountability OfficeExport Controls: Commerce Implemented Advanced Semiconductor Rules and Took Steps to Address Compliance Challenges (opens in new tab)
    Primary sourceGovernment auditPublished 02 Dec 2024
    Source details

    Quotation or reported data

    BIS cited avoiding stockpiling as one reason for immediate enforcement.
    Exact location
    Highlights; What GAO Found, paragraphs 1–2
    Source accessed
    03 Sept 2026

    Counterevidence

    A policy concern is not evidence of actual inventory, intent, ownership, or use by a named Chinese developer.

    Remaining uncertainty

    Firm-level pre-control and legacy inventories remain largely undisclosed.

  20. EV-20supportsModerate confidence

    Huawei reports commercial deployment of the Atlas 900 A3 platform underlying CloudMatrix384.

    HuaweiGroundbreaking SuperPoD Interconnect: Leading a New Paradigm for AI Infrastructure (opens in new tab)
    Primary sourceCompany disclosurePublished 18 Sept 2025
    Source details

    Quotation or reported data

    Huawei reported more than 300 Atlas 900 A3 SuperPoDs serving over 20 customers and identified CloudMatrix384 as a cloud service built on that platform.
    Exact location
    HUAWEI CONNECT 2025 keynote, Atlas 900 A3 deployment discussion
    Source accessed
    11 Sept 2026

    Counterevidence

    This is a vendor statement. The count covers Atlas systems across customers, not independently audited CloudMatrix instances running DeepSeek-R1.

    Remaining uncertainty

    Utilization, model mix, uptime, delivered capacity, metered energy and total cost are not independently established.

  21. EV-21supportsModerate confidence

    A later Huawei technical disclosure describes production model serving with xDeepServe on CloudMatrix384.

    HuaweiHuawei Cloud Model-as-a-Service on the CloudMatrix384 SuperPod (opens in new tab)
    Primary sourceTechnical preprintPublished 04 Aug 2025
    Source details

    Quotation or reported data

    Version 6 reports 900 ms time to first token and 34.8 ms average time per output token for a representative production workload. Its separate fixed-length peak-decoding test reports 2,400 tokens per second per Ascend 910C chip near 50 ms per output token.
    Exact location
    Version 6 (1 March 2026), sections 7.1–7.2; representative production setup versus peak decoding
    Source accessed
    11 Sept 2026

    Counterevidence

    Vendor-authored and not independently reproduced. Peak throughput and production latency refer to different configurations and workloads; they must not be combined into a single performance claim.

    Remaining uncertainty

    The disclosure does not establish independently metered power, total cost, fleet-wide reliability or like-for-like parity with a competing system.

“Primary” identifies an original or official source for the recorded claim. It does not mean that the finding has been independently verified. Full JSON and CSV datasets are available below.