Huawei Unveils OceanStor M900 to Enhance AI Inference in Hyperscale Data Centers
Huawei Introduces OceanStor M900 Context Memory Storage
At the recent HUAWEI CONNECT 2026 event, Huawei's rotating chairman and Board Vice President, David Wang, unveiled the OceanStor M900 Context Memory Storage. This innovative solution is aimed at enhancing artificial intelligence (AI) inference within hyperscale data centers. The OceanStor M900 is designed specifically for SuperPoDs, featuring shared memory capabilities that can scale up to multiple petabytes with remarkable performance metrics.
In 2026, AI has seen rapid progression from groundbreaking technological innovations to large-scale implementations. AI applications have evolved significantly, moving from simple chatbots to advanced agents capable of handling complex tasks autonomously. These developments reflect a burgeoning era of agent-based AI that is now widely adopted across strategic sectors.
With large models now exceeding ten trillion parameters, SuperPoDs are emerging as the optimal choice for AI infrastructure. Traditional large models support impressive context windows of over one million tokens, making multi-turn inference and complex tasks commonplace. The cache key-value (KV) data generated during inference continues to expand exponentially, prompting the need for a coherent storage system that integrates on-chip memory, DRAM, and SSDs to create an extensively shared memory space.
Huawei's OceanStor M900 has been engineered to address memory capacity limitations faced during ultra-long contexts and multi-turn inference. Utilizing the UnifiedBus network, the OceanStor M900 establishes a global, multilayer cache KV system on a petabyte scale, facilitating one-hop connections that maximize the computational power of SuperPoDs and expedite AI inference processes.
Key Features of OceanStor M900
1. Enhanced Capacity for Large-Scale AI:
Leveraging the high-speed UnifiedBus interconnect, the KV cache facilitates global pooling and sharing, providing a robust multi-level storage solution. The extension of KV cache from on-chip memory and DRAM to include SSDs allows a single cluster to deliver an unprecedented 64 petabytes of storage capacity. The cache capacity available for Neural Processing Units (NPU) can advance from gigabytes to terabytes, enhancing the ability to store, share, and reuse contextual information, thus significantly improving cache hit ratios.
2. Optimized Inference Performance:
The OceanStor M900 claims to be the first in its sector to blend CPU, network control units, and NAND control units into its architecture. It offers native KV semantics that enable one-hop connections between SuperPoD NPUs and SSDs. This innovation removes the need for protocol conversion and CPU forwarding, cutting access latency down from milliseconds to 60 microseconds—a staggering 90% reduction. By providing an aggregated bandwidth of 40 TB/s per cluster, this architecture enhances throughput, doubling token processing rates and halving the time taken to deliver the initial token, significantly converting computational power into productivity.
3. Cost Efficiency for Token Expenses:
OceanStor M900 utilizes the industry’s first KV-aware adaptive storage technology that intelligently predicts KV cache lifecycles based on data value. This capability spreads data distribution across multiple storage mediums, facilitating up to 24 drive writes per day (DWPD), expanding SSD lifespan by a factor of 16, and ensuring stability for three years. By lowering media replacement and management costs, OceanStor M900 decreases long-term expenses associated with large-scale AI inference infrastructure, encouraging faster adoption of AI technologies in various industries.
As AI becomes increasingly embedded within production systems across multiple sectors, the paradigm of AI infrastructure is shifting from a computation-heavy model to one that promotes tighter interdependencies between computation, network, and storage. Context memory storage solutions like OceanStor M900 will prove pivotal in continuously enhancing the capacity and efficiency of KV inference caches in hyperscale environments.