The research record

The source record.

Inspect the source behind each claim, its strongest qualification, and the questions still open.

21records across 5 cases

7 of 21 evidence records · CloudMatrix systems engineering

  1. EV-09qualifiesModerate confidence

    BIS assessed that listed Ascend 910-series chips likely involved U.S. design technology, software, or equipment subject to export rules.

    U.S. Bureau of Industry and SecurityGuidance on Application of General Prohibition 10 (GP10) to People’s Republic of China (PRC) Advanced-Computing Integrated Circuits (ICs) (opens in new tab)
    Primary sourceGovernment guidancePublished 13 May 2025
    Source details

    Quotation or reported data

    The illustrative list names Ascend 910B and 910C.
    Exact location
    Pages 1–2, summary, illustrative list, and applicable regulations
    Source accessed
    03 Sept 2026

    Counterevidence

    This is an enforcement assessment, not a bill-of-materials audit of the Pangu fleet or CloudMatrix installation.

    Remaining uncertainty

    The precise origin and legal status of each chip and upstream component are not public.

  2. EV-10supportsHigh confidence

    CloudMatrix384 connects 384 Ascend NPUs and 192 Kunpeng CPUs through an all-to-all Unified Bus architecture.

    Huawei and SiliconFlowServing Large Language Models on Huawei CloudMatrix384 (opens in new tab)
    Primary sourceTechnical preprintPublished 19 Jun 2025
    Source details

    Quotation or reported data

    384 NPUs, 192 CPUs, disaggregated prefill, decode, and caching.
    Exact location
    Abstract; sections 3.2–3.4
    Source accessed
    03 Sept 2026

    Counterevidence

    The paper is authored by Huawei and SiliconFlow researchers and is not peer reviewed.

    Remaining uncertainty

    Vendor deployment claims now exist. Independent utilization, reliability, metered power and total cost across the reported fleet remain unverified.

  3. EV-11supportsModerate confidence

    The disclosed DeepSeek-R1 test achieved 1,943 decode tokens per second per NPU at 49.4 ms, below one profiled H800 result but above another published H800 baseline.

    Huawei and SiliconFlowServing Large Language Models on Huawei CloudMatrix384 (opens in new tab)
    Primary sourceTechnical preprintPublished 19 Jun 2025
    Source details

    Quotation or reported data

    1,943 tokens/s/NPU at 49.4 ms; reported efficiency 1.29 tokens/s/TFLOPS.
    Exact location
    Section 5.2, Tables 2–4
    Source accessed
    03 Sept 2026

    Counterevidence

    Tests used 256 of 384 NPUs; decode assumed a 70% MTP acceptance rate; batch sizes and implementations differed; and the highest prefill result assumed idealized expert load balancing.

    Remaining uncertainty

    Like-for-like cost, power, reliability, and hardware-count comparisons are unavailable.

  4. EV-12qualifiesHigh confidence

    CloudMatrix-Infer reduced batch size to meet tighter latency targets, lowering reported per-NPU throughput and illustrating the latency–throughput trade-off.

    Huawei and SiliconFlowServing Large Language Models on Huawei CloudMatrix384 (opens in new tab)
    Primary sourceTechnical preprintPublished 19 Jun 2025
    Source details

    Quotation or reported data

    Throughput fell from 1,943 to 538 tokens/s/NPU as the latency target tightened from 50 ms to 15 ms and batch size per NPU fell from 96 to 8.
    Exact location
    Section 5.2, Table 4
    Source accessed
    22 Sept 2026

    Counterevidence

    The system still met the stated tighter latency target and maintained useful throughput.

    Remaining uncertainty

    Real traffic mixes, uptime, and long-run service cost are not independently reported.

  5. EV-13qualifiesLow confidence

    A specialist estimate argues that matching a GB200 NVL72-class system requires substantially more Ascend devices, optical components, and power.

    SemiAnalysisHuawei AI CloudMatrix 384: China’s Answer to Nvidia GB200 NVL72 (opens in new tab)
    Secondary sourceSpecialist analysisPublished 16 Apr 2025
    Source details

    Quotation or reported data

    Estimated 4.1 times the system power of a GB200 NVL72.
    Exact location
    Opening comparison; CloudMatrix 384 System Architecture
    Source accessed
    03 Sept 2026

    Counterevidence

    The estimate is not metered or audited, and the compared systems differ in chip count and configuration.

    Remaining uncertainty

    No public independent energy, cooling, or total-cost measurement validates the exact multiplier.

  6. EV-20supportsModerate confidence

    Huawei reports commercial deployment of the Atlas 900 A3 platform underlying CloudMatrix384.

    HuaweiGroundbreaking SuperPoD Interconnect: Leading a New Paradigm for AI Infrastructure (opens in new tab)
    Primary sourceCompany disclosurePublished 18 Sept 2025
    Source details

    Quotation or reported data

    Huawei reported more than 300 Atlas 900 A3 SuperPoDs serving over 20 customers and identified CloudMatrix384 as a cloud service built on that platform.
    Exact location
    HUAWEI CONNECT 2025 keynote, Atlas 900 A3 deployment discussion
    Source accessed
    11 Sept 2026

    Counterevidence

    This is a vendor statement. The count covers Atlas systems across customers, not independently audited CloudMatrix instances running DeepSeek-R1.

    Remaining uncertainty

    Utilization, model mix, uptime, delivered capacity, metered energy and total cost are not independently established.

  7. EV-21supportsModerate confidence

    A later Huawei technical disclosure describes production model serving with xDeepServe on CloudMatrix384.

    HuaweiHuawei Cloud Model-as-a-Service on the CloudMatrix384 SuperPod (opens in new tab)
    Primary sourceTechnical preprintPublished 04 Aug 2025
    Source details

    Quotation or reported data

    Version 6 reports 900 ms time to first token and 34.8 ms average time per output token for a representative production workload. Its separate fixed-length peak-decoding test reports 2,400 tokens per second per Ascend 910C chip near 50 ms per output token.
    Exact location
    Version 6 (1 March 2026), sections 7.1–7.2; representative production setup versus peak decoding
    Source accessed
    11 Sept 2026

    Counterevidence

    Vendor-authored and not independently reproduced. Peak throughput and production latency refer to different configurations and workloads; they must not be combined into a single performance claim.

    Remaining uncertainty

    The disclosure does not establish independently metered power, total cost, fleet-wide reliability or like-for-like parity with a competing system.

“Primary” identifies an original or official source for the recorded claim. It does not mean that the finding has been independently verified. Full JSON and CSV datasets are available below.