Huawei Unveils OceanStor M900: Revolutionizing AI Inference in Hyperscale Data Centers

Huawei's OceanStor M900 Context Memory Storage



In a groundbreaking announcement at the HUAWEI CONNECT 2026 conference, David Wang, the Vice Chairman and rotating chairman of Huawei, unveiled the innovative OceanStor M900 Context Memory Storage solution. This state-of-the-art storage system is tailored for AI inference in hyperscale data centers, providing a fully shared memory space measured in petabytes (PB) and boasting extraordinary performance rates in terabytes per second (TB/s). This move marks a significant shift in AI infrastructure, transitioning from a computation-centric model to a heavily collaborative approach between computing power, networking, and storage.

As we step into 2026, the rapid evolution of AI is moving beyond technological breakthroughs towards large-scale deployment. AI applications have transformed from simple chatbots to agents capable of independently completing complex tasks, expanding their use in critical industries and ushering in a new era of agent-based AI.

The burgeoning size of AI models, now reaching up to 10 trillion parameters, has made SuperPoD systems the optimal choice for AI infrastructure. Prominent large models support contextual windows exceeding one million tokens, and multi-round inference has become the standard, further increasing the volume of key-value (KV) cache data generated during inference. These trends push on-chip memory and DRAM to their limits in both capacity and cost efficiency.

Industry consensus has highlighted the need for a multi-tiered storage system that coordinates on-chip memory, DRAM, and SSDs, effectively creating a fully shared memory space with immense capacity. To address the restrictions in memory capacity for extremely long contexts and multi-round inference, Huawei has introduced OceanStor M900. Utilizing the UnifiedBus network, it establishes a global multi-level KV cache measured in petabytes with single-hop connectivity, fully unleashing the computational potential of SuperPoD systems and accelerating AI inference in hyperscale data centers.

Key Capabilities of OceanStor M900


1. Breaking Capacity Limits:
- The impressive memory offered allows for large-scale AI support.
- With the high-speed UnifiedBus, the KV cache achieves global aggregation and sharing with multi-tier storage. The SuperPoD's KV cache expands seamlessly from on-chip memory and DRAM to SSDs, enabling a single cluster to offer up to 64 PB of capacity.
- The available KV cache capacity per NPU has escalated from gigabytes to terabytes, allowing for the storage, sharing, and re-utilization of more context, significantly enhancing cache hit rates.

2. Performance and Latency Reduction:
- OceanStor M900 is the industry's first architecture integrating CPU, network controller, and NAND controller, providing native KV semantics that enable single-hop connection from SuperPoD NPUs to SSDs.
- This architecture eliminates the need for protocol conversion and CPU forwarding, cutting access latency from milliseconds to just 60 microseconds—an impressive reduction of 90%. One cluster achieves an aggregated access throughput of 40 TB/s, surpassing competitor solutions by 1.5 times. In typical AI programming scenarios, this architecture doubles token throughput and halved the time to the first token, translating computational power into tangible productivity.

3. Cost Efficiency in AI Deployment:
- OceanStor M900 incorporates the first adaptive KV-aware storage technology in the industry, intelligently distributing data based on the value of the information and predicting the lifecycle of the KV cache.
- This allows for up to 24 drive writes per day (DWPD), extending SSD lifespan 16-fold and ensuring stability for three years. By reducing media replacement costs and operational expenses, it lowers the long-term infrastructure cost for large-scale AI inference and promotes faster AI deployment.

As AI continues to infiltrate core production systems across various sectors, the infrastructure supporting such advancements is evolving from a pure computational focus to an integrated framework that combines computational power with networking and storage solutions. The introduction of context memory storage by Huawei will be essential for continuously enhancing capacity and efficiency in accessing KV caches for inference in a hyperscale context.

Topics Consumer Technology)

【About Using Articles】

You can freely use the title and article content by linking to the page where the article is posted.
※ Images cannot be used.

【About Links】

Links are free to use.