Huawei Introduces UnifiedBus Computing Architecture for SuperPoD and Server Clusters

Huawei's Innovative UnifiedBus Architecture



At the recent HUAWEI CONNECT 2026 conference, Yang Chaobin, Huawei's Executive Director and CEO of ICT Business Group, unveiled pivotal advancements in computing architecture designed to alleviate the inefficiencies inherent in traditional IT frameworks.

In conventional computing layouts, resource utilization tends to diminish significantly as server clusters expand. For instance, only 20% of the computational capacity in a massive neural processing cluster consisting of 100,000 units is effectively utilized, leading to an alarming waste of resources during data exchanges. The rise of complex models, particularly those with billions of parameters requiring substantial amounts of intermediate data during training, further exacerbates these inefficiencies, revealing critical bottlenecks in inter-component communication and data transfer speeds.

To combat these challenges, Huawei has introduced the UnifiedBus architecture, which is poised to establish a new standard in interconnectivity among servers and SuperPoDs. This innovative architecture was developed in response to the increasing demand for computational power tied to the emergence of new agent-based applications. Here are its four main features:

1. Unified Memory and Protocol Semantics



UnifiedBus consolidates over ten interconnect protocols under a singular framework, which escalates interconnect bandwidth from 100 GB/s to multiple TB/s, slashing round-trip latency from 7 microseconds to just 2 microseconds. This protocol also features global unified memory addressing within SuperPoDs, enhancing efficiency.

2. Heterogeneous Computing Collaboration



By directly linking processors, neural processing units, memory, and SSDs, UnifiedBus enables decentralized peer-to-peer access among components. It allows flexible combinations of CPUs and neural processing units (NPUs) and incorporates multi-level hardware acceleration specifically designed for Transformers, enhancing the efficiency of Attention-FFN (AFD).

3. Multi-Level Storage with Global Pooling



UnifiedBus supports hybrid media resource pooling to facilitate the caching of activations. DDR memory is used as an alternative memory source for NPUs, which reduces search latency for services like recommendation algorithms and doubles vector search performance for datasets with up to 100 billion elements. This system significantly decreases the need for high-bandwidth memory capacity per NPU when training models containing billions of parameters, thereby boosting the overall utilization rates of FLOPs (floating-point operations per second).

4. Optoelectronic Interconnectivity and Flexible Networks



UnifiedBus acts as a 'data highway' providing ultra-high bandwidth and ultra-low latency, ensuring scalable computational power.

Latest UnifiedBus Products



During the keynote, Yang also introduced Huawei's UnifiedBus interconnect products which address connectivity within cabinets, between cabinets, and across clusters, paving the way for flexible scalability from a single cabinet to extensive clusters containing millions of processing units. For example, the UnifiedBus LinkBlade eliminates circuit losses with a cable-free design, reducing approximately 196 kilometers of copper cabling in a 4,096 unit neural processing SuperPoD.

Moreover, the UnifiedBus LinkDevice supports 176 ports with a bandwidth of 1.6 Tbit/s each, delivering a total optical interconnect bandwidth of 280 Tbit/s—this sets a new industry benchmark for high-gate protocol devices.

The UnifiedBus UBG switch illustrates remarkable output capabilities with a support capacity of up to 1,024, facilitating superclusters that can manage an impressive total of a million neural processing units while efficiently handling numerous billions of parameters.

Conclusion



Yang Chaobin also showcased Huawei's latest generation of agent-based AI superclusters built on UnifiedBus technology, featuring the TaiShan 950 SuperPoD, Atlas 960, OceanStor M900 memory storage system, and the Xinghe UBG switch. This supercluster is designed for versatile computing collaboration and multi-level resource pooling.

Beyond the SuperPoDs, Huawei has extended UnifiedBus technology to new computational appliances, enabling small to medium enterprises to locally execute models with billions of parameters through its Kunpeng and Ascend modules. As part of its commitment to robust collaborative ecosystems, Huawei continues to partner with open-source frameworks like DeepSeek Harness and OpenCode to enhance intelligent management, inference acceleration, and security features within UnifiedBus technology.

In conclusion, Yang pledges that Huawei's focus will remain on innovating at the systems level, building a diverse array of computing products grounded in open-source technologies, thereby contributing to a transformative ecosystem for computational power globally.

Topics Consumer Technology)

【About Using Articles】

You can freely use the title and article content by linking to the page where the article is posted.
※ Images cannot be used.

【About Links】

Links are free to use.