September 19, 2026
AI Training Moves Toward Million-Scale Clusters: Huawei Builds Peerium Architecture
On 17 September 2026, Huawei introduced Atlas 960E SuperPoD with NPO optical technology in a company blog post. The SuperCluster can scale to 1,000,000 NPUs, while Peerium is designed to unite processors, memory, storage, and networking into one system. In Huawei's simulation, a 100,000-NPU cluster in 4,000-unit blocks raised MFU by 2.75x compared with 8-NPU servers.

In conventional server architecture, data exchange takes more than 40% of training time, Huawei said on 17 September. In its simulation, SuperPoD lowered the cost of that exchange in terms of MFU.
Scale ladder. Atlas 950 SuperCluster can connect up to 256,000 accelerator cards for training and inference. One Ascend 960 supernode is expected to unite up to 4,096 AI processors, while several supernodes form clusters with hundreds of thousands of processors.
Connectivity and timeline. Hi-ONE delivers 7.2 Tb/s per optical interconnect engine. Ascend 960DT is expected in Q1 2027, and Ascend 960PR in Q3 2027. Huawei says it moved both dates forward from the original plan.
Supply market. Huawei has already deployed more than 1,000 systems based on Ascend 910C, and more than 40 AI models have been trained directly on its platform. According to Reuters on 16 September, demand for equipment in China exceeds the company's production capacity, while overseas shipments constrain Huawei.
Huawei's next planned milestones: Ascend 960DT in Q1 2027 and Ascend 960PR in Q3 2027.
Source
