Huawei Unveils OceanStor M900 Context Memory Storage Solution to Boost AI Inference in Hyperscale Data Centers
Huawei's OceanStor M900: The New Frontier in AI Inference
At the recent HUAWEI CONNECT 2026 conference, Huawei's vice chairman, David Wang, unveiled the OceanStor M900, a revolutionary memory storage solution designed specifically to optimize AI inference in hyperscale data centers. This innovative product integrates massively shared memory capabilities, supporting petabyte-level capacity and terabyte-per-second performance, thereby transforming the infrastructure landscape from a computing-centric model to a cohesive synergy between computing, networking, and storage. This shift is essential for unleashing the full computational power of SuperPods.
2026 has been a pivotal year for AI, transitioning from technological advancements to large-scale implementations. The evolution of AI applications is evident, moving from simple chatbots to highly sophisticated agents capable of autonomously executing complex tasks in strategic sectors. This signifies the dawn of the agentic AI era. As AI models grow, reaching up to 10 trillion parameters, the demand for efficient and robust infrastructure has never been greater. Current AI models manage context windows exceeding a million tokens, with an increasing amount of data produced during inference processes necessary for multi-turn and complex tasks. This trend has placed significant stress on traditional integrated memory and DRAM systems.
Recognizing these challenges, Huawei's OceanStor M900 is designed to eliminate bottlenecks related to memory capacity in lengthy contexts and multi-iteration inferences. It utilizes the UnifiedBus network to establish a global multi-level KV cache with petabyte-class capacity and single-hop connections. This architecture is crucial for maximizing the computational efficiency of SuperPods, significantly speeding up AI inference in hyperscale environments.
Key Features of OceanStor M900
1. Capacity Expansion for Large-scale AI Optimization
The UnifiedBus high-speed interconnection allows global memory pooling and sharing within a multi-tiered storage system.
The cache from integrated memory and DRAM to SSD can afford each cluster a staggering 64 Petabytes of capacity. This shift allows NPU availability of KV cache to extend from several gigabytes to multiple terabytes, enhancing context storage, sharing, and reuse. This seamless capacity increase leads to a notable rise in KV cache hit rates.
2. Enhanced Inference Performance to Maximize Computational Power
The OceanStor M900 claims the industry's first architecture to incorporate a processor, network controller, and NAND controller within a single unit. The native KV semantics enable direct single-hop connectivity between SuperPods’ NPUs and the SSDs. This eliminates the need for protocol conversion or processor overhead, achieving a remarkable reduction in access latency—from several milliseconds to just 60 microseconds—markedly improving efficiency by up to 90%. A single cluster can now achieve an aggregated access bandwidth of 40 Terabytes per second, outperforming competitive solutions by 1.5 times. In regular AI programming settings, this architecture doubles the token throughput while halving the time to retrieve the first token, translating computational potential directly into enhanced productivity.
3. Token Cost Reduction for Economically Feasible Large-scale AI Adoption
OceanStor M900 is also equipped with the first adaptive KV storage technology in the industry, capable of forecasting KV cache lifecycle based on data value and intelligently distributing it across various storage mediums. This innovation allows for a durability rate of up to 24 daily writes (DWPD), extending SSD lifespan 16-fold, ensuring stability over three years. By minimizing costs linked to hardware replacement and maintenance, this solution dramatically reduces long-term expenses of large-scale AI inference infrastructures, promoting quicker adoption of AI.
As AI continues to pervade critical business systems across various sectors, its requisite infrastructure is evolving. It is shifting from a computation-centric model to a framework emphasizing seamless collaboration between computing resources, networking setups, and storage solutions. The context memory storage systems like OceanStor M900 will be crucial for continuously improving caching capacity and access efficiency for extensive inference caches.
In conclusion, Huawei's OceanStor M900 not only marks a significant step forward in data center technology but also asserts a promising foundation for the expansive future of AI-driven enterprise solutions.