Huawei Unveils OceanStor M900 to Accelerate AI Inference in Hyperscale Data Centers

Huawei Unveils OceanStor M900 Context Memory Storage



During the HUAWEI CONNECT 2026 conference, David Wang, Vice Chairman and rotating Chairman of Huawei's board, introduced the cutting-edge OceanStor M900 Context Memory Storage. This groundbreaking solution is designed specifically for AI inference in hyperscale data centers, allowing SuperPoD systems to utilize a fully shared memory space offering petabyte-scale capacity and terabyte-per-second performance. This innovation marks a significant shift in AI infrastructure from a computation-centric approach to one that emphasizes deep collaboration among computing, networking, and storage.

As we progress through 2026, the transition toward comprehensive AI implementations has accelerated, with applications evolving from simple chatbots to complex agents capable of autonomously executing intricate tasks across critical sectors. This advancement signals the dawn of the agent-based AI era.

Given the rapid expansion of large models—scaling up to 10 trillion parameters—SuperPoD systems are becoming the optimal choice for AI infrastructure. Traditional large models are already supporting context windows exceeding one million tokens, and intricate tasks are now the norm, leading to a continuous increase in the volume of KV cache data generated during inference. These trends have stretched the capabilities of on-chip memory and DRAM beyond feasible limits, both in terms of capacity and cost-effectiveness. Industry consensus has emerged on the necessity for a multi-tiered storage system that coordinates on-chip memory, DRAM, and SSD storage, creating a fully shared memory space of enormous capacity.

Huawei's OceanStor M900 Context Memory Storage addresses the memory capacity bottlenecks associated with ultralong contexts and multi-stage inference. Utilizing a UnifiedBus network, the system constructs a global multi-tier KV cache at petabyte scale, facilitating single-hop connections. This optimization fully unleashes the computational potential of SuperPoD systems and significantly accelerates AI inference in hyperscale data centers. Key features of the OceanStor M900 include:

1. Overcoming Capacity Limits and Supporting Large-Scale AI Applications



With its high-speed UnifiedBus interconnect, the OceanStor M900 enables global aggregation and sharing within multi-tier storage architecture. The KV cache of SuperPoD systems expands from on-chip memory and DRAM to SSDs, allowing a single cluster to achieve a capacity of 64 PB. The available KV cache capacity for NPUs has surged from gigabytes to terabytes, thereby enhancing the storage, sharing, and reuse of extensive context data. This advancement greatly increases the success rate of KV cache hits.

2. Enhancing Inference Performance and Unlocking Full Computational Power



The OceanStor M900 introduces a unique architecture that integrates CPU, network control units, and NAND controllers, providing a native KV semantics. This design facilitates immediate connections from NPU SuperPoD to SSD without the need for protocol conversion or CPU redirection, reducing access latency from milliseconds to 60 microseconds—a 90% improvement. Each cluster can deliver a cumulative access bandwidth of 40 TB/s, 1.5 times more than competing solutions. In typical AI programming scenarios, this architecture doubles the token throughput of the inference cluster while halving the time to the first token, translating computational capabilities into productivity gains.

3. Reducing Token Costs for Economic Large-Scale AI Implementation



Employing the first adaptive storage technology that supports KV for predictive lifecycle management, the OceanStor M900 intelligently distributes data across storage media based on data value. This innovation allows for up to 24 disk writes per day (DWPD), extending SSD lifespan by 16 times and ensuring stability over three years. By lowering media replacement and operational costs, long-term expenses for extensive AI inference infrastructures are significantly reduced, facilitating speedier deployments.

As AI continues to permeate the core manufacturing systems across various industries, the infrastructural landscape for AI is transitioning from computation-centric models to tighter integration of computing, networking, and storage. Context memory storage is poised to become essential for continually increasing the capacity and efficiency of hyperscale KV inference caches.

Topics Consumer Technology)

【About Using Articles】

You can freely use the title and article content by linking to the page where the article is posted.
※ Images cannot be used.

【About Links】

Links are free to use.