Huawei Introduces Innovative UnifiedBus Computing Architecture for Next-Gen Clusters
Huawei Introduces UnifiedBus Computing Architecture
At the recent HUAWEI CONNECT 2026 event, Yang Chaobin, Huawei's Executive Director of the Board and CEO of the ICT Business Group, delivered a compelling keynote about the future of computing architectures, specifically focusing on the newly unveiled UnifiedBus interconnect architecture. This innovation is poised to revolutionize how computational resources are utilized in vast clusters and SuperPoDs, addressing critical inefficiencies inherent in traditional systems.
The Challenge of Traditional Architectures
In standard computing environments, scaling usage leads to decreased resource efficiency. For example, in a 100,000-NPU cluster, it's reported that only about 20% of the computing power is applied by various models, resulting in idle capabilities during data transfer times. Moreover, training enormous models, such as a 10 trillion-parameter model, often outstrips the memory capacities of single accelerators, creating notable performance barriers during crucial inter-component communications. The need to solve these inefficiencies is urgent as demand for computational power escalates with emerging applications.
Introducing UnifiedBus Architecture
Yang articulated how Huawei's UnifiedBus architecture directly addresses these challenges through several key innovations:
1. ### Unified Protocol and Memory Semantics
By consolidating over ten interconnect protocols under the UnifiedBus protocol, the architecture boosts interconnect bandwidth from 100 GB/s thresholds to levels in terabits per second. This shift concurrently minimizes round-trip time (RTT) latency, cutting it down from 7 microseconds to just 2 microseconds and enabling unified global memory addressing in SuperPoD configurations.
2. ### Heterogeneous Compute Collaboration
The architecture effectively connects CPUs, NPUs, memory, and SSDs, fostering decentralized peer-to-peer access and flexible combinations of CPUs and NPUs. It features tiered hardware acceleration, facilitating advancements such as Attention-FFN disaggregation, crucial for high-performance computing.
3. ### Tiered Storage with Global Pooling
Leveraging hybrid-media resource pooling, UnifiedBus enhances caching capabilities for computational activations. Overall performance is amplified by utilizing Double Data Rate (DDR) memory to support NPUs, resulting in substantially reduced latency, especially for services like search and recommendation engines. This innovation also has the added benefit of reducing training requirements for models exceeding 10 trillion parameters.
4. ### Optoelectronic Interconnect and Flexible Networking
The UnifiedBus functions as a high-bandwidth, low-latency global