Huawei Unveils OceanStor M900 Context Memory Storage for Data Center AI Inference

Huawei Unveils OceanStor M900 Context Memory Storage



At the HUAWEI CONNECT 2026 event, David Wang, the Deputy Chairman and rotating Chairman of Huawei, officially introduced the new OceanStor M900 Context Memory Storage. This innovative solution is designed to accelerate AI inference in hyperscale data centers, addressing the growing needs for processing power and storage capacity. The OceanStor M900 provides a completely shared memory area for SuperPoDs with capabilities measured in petabytes (PB) and performance reaching terabytes per second (TB/s).

The emergence of AI technology has rapidly transitioned from mere breakthroughs to large-scale implementations. In 2026, AI applications have evolved beyond simple chatbots to sophisticated agents capable of autonomously performing complex tasks across critical sectors, marking the dawn of an era governed by agent-based AI initiatives.

As models expand up to 10 trillion parameters, SuperPoDs have become the prime choice for robust AI infrastructure. These large-scale models now support context windows exceeding one million tokens, establishing multi-turn inference and complex operations as the new standard. However, the generated key-value (KV) cache data during inference has significantly increased, challenging the limits of on-chip memory and Dynamic Random Access Memory (DRAM) in terms of capacity and cost efficiency. Industry consensus has shifted towards developing a multi-layered memory system that coordinates on-chip memory, DRAM, and solid-state drives (SSDs) to create a vast shared memory environment.

To counter challenges posed by storage limitations in long-context and multi-turn inference, Huawei has developed the OceanStor M900. This system utilizes UnifiedBus networking to establish a global, multi-tier KV cache on a petabyte scale with one-hop connectivity, fully unleashing the compute potential of SuperPoDs and accelerating AI inference in hyperscale environments. The key features of the OceanStor M900 include:

Breaking Capacity Barriers for Large-Scale AI


With its high-speed UnifiedBus connection network, the KV cache achieves global bundling and sharing across multi-tier storage. The SuperPod’s KV-cache is expanded from on-chip memory and DRAM to SSDs, enabling a single cluster to provide a capacity of 64 PB. The available KV-cache capacity per Neural Processing Unit (NPU) scales from gigabytes to terabytes, facilitating greater context storage, sharing, and reuse, dramatically increasing the cache hit rate.

Enhancing Inference Performance to Maximize Compute Capacity


The OceanStor M900 boasts the industry’s first architecture that integrates the CPU, network controllers, and NAND controllers. This design provides native KV semantics, enabling one-hop connections from the SuperPod NPUs directly to SSDs, eliminating protocol conversion and CPU forwarding, which reduces access latency from milliseconds to 60 microseconds—a 90% decrease. Each cluster can deliver an aggregated access bandwidth of 40 TB/s, 1.5 times higher than comparable solutions. In typical AI programming scenarios, this architecture doubles the token throughput of inference clusters and halves the time to the first token, translating computational power into productivity.

Reducing Token Costs for Economical Large-Scale AI Deployment


Huawei's OceanStor M900 deploys the industry’s first KV-aware adaptive storage technology, which predicts the lifecycle of the KV cache based on data value and intelligently distributes data across storage media. This innovation supports up to 24 drive writes per day (DWPD), extending the lifespan of SSDs by 16 times and ensuring stability for over three years. By reducing costs related to media replacement, operations, and maintenance, this technology lowers the long-term expenses of a large-scale AI inference infrastructure, streamlining AI adoption.

As AI penetrates various industries and their critical production systems, the infrastructure shifts from a computation-centric model towards a more integrated collaboration between computing power, networking, and storage. Developing a context memory will be vital to continuously enhance the capacity and access efficiency of hyperscale inference KV caches.

Topics Consumer Technology)

【About Using Articles】

You can freely use the title and article content by linking to the page where the article is posted.
※ Images cannot be used.

【About Links】

Links are free to use.