Huawei Unveils OceanStor M900 to Revolutionize AI Inference in Hyperscale Datacenters
Huawei's New Innovation in AI Infrastructure
During the HUAWEI CONNECT 2026 event, David Wang, the Vice Chairman of the Board and rotating Chairman of Huawei, introduced the OceanStor M900 Context Memory Storage solution, designed specifically for AI inference in hyperscale datacenters. This new product signifies a major shift from a compute-centric model to a collaborative approach that integrates computing power, networking, and storage.
The launch coincides with rapid developments in the AI sector, transitioning from breakthrough technologies to large-scale implementations. AI applications have evolved from simple chatbots to advanced agents capable of handling complex tasks independently. These advancements are significantly impacting critical industries, marking the beginning of the era of agent-based AI.
As AI models expand to scales of 10 trillion parameters, the adoption of SuperPODs has become pivotal for AI infrastructure. Many leading AI models now accommodate context windows exceeding one million tokens, with multi-turn inference and intricate tasks becoming the standard. The volume of KV-cache data generated in the inference process is on the rise, straining the limits of on-chip memory and DRAM in terms of both capacity and cost-efficiency.
Industry consensus is shifting towards developing a multi-layered storage architecture that harmonizes on-chip memory, DRAM, and SSDs to create an expansive shared memory space.
Addressing Memory Capacity Constraints
The OceanStor M900 Context Memory Storage product directly addresses the memory capacity bottlenecks posed by extensive context windows and multi-turn inference requirements. Utilizing the UnifiedBus network, it establishes a global KV-cache with PB-scale multi-layer storage through single-hop connections. This advancement fully harnesses the computational capacity of SuperPODs, accelerating AI inference in hyperscale datacenters.
The OceanStor M900's capabilities are transformative for the AI landscape:
1. Breaking Capacity Limits: The architecture supports large-scale AI applications with significant memory capacity. The UnifiedBus interconnect network allows for global bundling and sharing of the KV-cache via multi-layer storage. It extends the KV-cache from on-chip memory and DRAM to the realm of SSDs, with one cluster achieving a staggering 64 PB capacity, subsequently increasing the KV-cache available per NPU from gigabytes to terabytes and dramatically enhancing cache hit rates.
2. Boosting Inference Performance: This product features a pioneering architecture that integrates CPU, network controller, and NAND controller components. By enabling native KV semantics, the OceanStor M900 allows for one-hop connections from the SuperPOD's NPU to SSDs. This innovation eliminates the need for protocol conversion and CPU forwarding, decreasing access latency from milliseconds to 60 microseconds—a striking reduction of 90%. Each cluster delivers an impressive total access bandwidth of 40 TB/s, which is 1.5 times that of competing solutions. In typical AI programming environments, this architecture doubles token throughput for inference clusters, effectively halving time-to-first-token and translating computational power into productivity.
3. Redefining Token Costs: The OceanStor M900 employs the first-ever KV-aware adaptive storage technology in the industry, predicting the lifecycle of the KV-cache based on data value and intelligently distributing data across various storage media. This design supports up to 24 drive writes per day (DWPD), extending the lifespan of SSDs by a factor of 16, ensuring stability for up to three years. By reducing media replacement and operations and maintenance (OM) costs, the technology drives down the long-term expenses of large-scale AI inference infrastructure, facilitating quicker AI deployment.
As AI becomes integrated within essential production systems across diverse sectors, the underlying AI infrastructure is evolving towards a model that fosters closer cooperation between computing power, networking, and storage solutions. The OceanStor M900 Context Memory Storage will play a critical role in continually enhancing the capacity and efficiency of accessing KV caches for large-scale inference needs within hyperscale environments.