Huawei Introduces Groundbreaking UnifiedBus Computing Architecture
In a bold stride towards innovation, Huawei has unveiled its latest advancement in computing architecture known as UnifiedBus during the HUAWEI CONNECT 2026 event. Yang Chaobin, a notable executive in the company, detailed this new architecture, highlighting its potential to revolutionize the way cluster systems operate and communicate.
The Challenge of Traditional Computing Architectures
Traditional computing architectures have faced significant bottlenecks as the size of clusters increases. It is reported that with 100,000 Neural Processing Units (NPUs), only 20% of the computational capacity is utilized effectively. A large portion of processing power remains idle during data communication. Moreover, models requiring substantial intermediate data for training often exceed the memory capacity of a single accelerator under conventional designs. This inefficiency has made connection and communication between components in a cluster a critical performance bottleneck.
UnifiedBus: A Solution to Bottlenecks
Yang presented the UnifiedBus as a breakthrough solution, designed to facilitate seamless collaboration among clusters and SuperPod systems. This architecture aims to meet the soaring demand for computational power driven by the emergence of innovative agent applications entering the market. The UnifiedBus architecture boasts four fundamental features:
1.
Unified Protocol and Memory Semantics: Over ten connection protocols have been consolidated under the UnifiedBus, boosting connection throughput from 100 GB/s to a remarkable Tier level of TB/s. Furthermore, it reduces Round Trip Time (RTT) latency from 7 microseconds to just 2 microseconds, enabling unified global memory addressing within SuperPod systems.
2.
Heterogeneous Resource Collaboration: The UnifiedBus establishes direct connections between CPU and NPU units, memory, and SSD blocks, creating a decentralized peer-to-peer approach among these components. This allows flexible combinations of CPU and NPU resources, enhancing the efficiency of computational tasks.
3.
Multilevel Storage with Global Aggregation: By employing hybrid media aggregation, UnifiedBus optimizes the caching of activation data. It utilizes DDR memory as an alternative for NPU tasks, effectively reducing latency for services such as search and recommendation algorithms, and doubling the vector search performance when dealing with trillions of data entries.
4.
Optoelectronic Connectivity and Flexible Networking: Acting as a global 'data highway', the UnifiedBus promotes ultra-high bandwidth and ultra-low latency, allowing for flexible scaling of computing resources across varying applications.
Innovative Products Powered by UnifiedBus
At the event, Yang also highlighted Huawei's new line of UnifiedBus connection products that facilitate connections within server enclosures, between cabinets, and across clusters. This new suite enables flexible scaling from single enclosures to expansive clusters containing a million NPUs.
- - LinkBlade within Enclosures: The UnifiedBus LinkBlade features a cable-free design that drastically reduces circuit losses, cutting down copper cable lengths by approximately 196 kilometers in a SuperPod setup containing 4,096 NPUs.
- - LinkDevice Between Cabinets: This device operates with 176 ports, each delivering a throughput of 1.6 Tbit/s, and provides fully optical connectivity with a staggering throughput of 280 Tbit/s. Its extremely low 2 microseconds RTT latency positions it at the forefront of industry-leading high-speed bus protocols.
- - UnifiedBus UBG Switch Across Clusters: This switch supports an impressive branching capacity of up to 1,024, empowering a SuperCluster with a million NPUs and paving the way for advancements towards supporting parameters in the tens of trillions.
AI SuperCluster Unveiling
In addition to the hardware, Yang rolled out Huawei's AI SuperCluster powered by UnifiedBus, which integrates TaiShan 950 systems, Atlas 960 SuperPod, OceanStor M900 storage, and Xinghe UBG switches. This innovative cluster design offers collaborative capabilities among diverse computational resources while employing multilevel resource aggregation.
Huawei is committed to the evolution of its UnifiedBus technology through its computing devices, developing new devices based on this architecture that accommodate one to eight NPUs. These devices, along with Kunpeng and Ascend modules, will empower small and medium-sized enterprises to run models with trillions of parameters locally.
A Commitment to Open Source Collaboration
Yang emphasized that Huawei's collaboration with open source frameworks such as DeepSeek Harness, OpenCode, and openJiuwen is pivotal in creating fully open-source agent plugins for intelligent management, accelerating inference, and enhancing security frameworks. He underscored Huawei's commitment to innovations at the systemic level, constructing a robust portfolio of computational foundations that emphasize open-source and open systems.
"We will continue to collaborate with customers, partners, and developers to allow this ecosystem to thrive and provide the world with a new paradigm in computational power," added Yang, marking a new chapter in Huawei’s ambitious technological journey.