Huawei Introduces OceanStor M900 Storage Solution for Accelerating AI Extraction in Hyperscale Data Centers

Huawei Unveils OceanStor M900



In a significant announcement during the HUAWEI CONNECT 2026 event, David Wang, Vice Chairman of the Huawei Management Board, introduced the OceanStor M900 Context Memory Storage System. This innovative solution is specifically developed to expedite artificial intelligence (AI) extraction processes in hyperscale data centers. It offers a fully shared memory space capable of PB-scale capacity and TB/s level performance for SuperPod systems, representing a shift from a compute-focused model to one that emphasizes deep collaboration between computing, networking, and storage.

As we move through 2026, we witness accelerated transitions of AI from technological breakthroughs to large-scale applications. AI implementations have evolved dramatically, transitioning from basic chatbots to more complex systems capable of autonomously handling intricate tasks. These changes signify the dawn of an agent-based AI era, widely adopted across critical sectors.

The SuperPods, which are pivotal in this evolution, are becoming preferred choices for AI infrastructure as the parameters of large models expand to trillions. The usage of leading AI models, which now support context windows exceeding one million tokens, is increasing, making multi-stage extraction and complex task handling the norm. Consequently, the volume of KV cache data generated during extraction is rising, pushing the limits of traditional on-chip memory and DRAM regarding capacity and cost-effectiveness.

Addressing Memory Capacity Bottlenecks


To combat the challenges posed by ultra-long-context multi-turn extraction tasks, Huawei has rolled out the OceanStor M900 Context Memory Storage solution. Utilizing the UnifiedBus network, this system provides single-step connections and constructs a global multi-layered KV cache at PB-scale. These innovations unleash the full computational power of SuperPods and significantly expedite AI extraction in hyperscale data environments.

The OceanStor M900 boasts three fundamental characteristics that elevate its capabilities:

1. Breaking Capacity Barriers for Large-Scale AI

The high-speed UnifiedBus connection enables a global pooling and sharing mechanism for KV cache through hierarchical storage. This configuration allows the KV cache of SuperPods to be expanded from on-chip memory and DRAM to SSDs, delivering a single cluster with 64 PB of capacity. This enhancement elevates the usable KV cache capacity per NPU from gigabytes to terabytes, facilitating improved storage, sharing, and reuse of contexts, thereby significantly boosting KV cache hit rates.

2. Maximizing Extraction Performance for Enhanced Computational Power

The OceanStor M900 introduces the industry's first architecture that consolidates the CPU, network controller unit, and NAND controller under a single umbrella. This configuration allows for direct single-step connections from SuperPod NPUs to SSDs while providing local KV semantics. By doing so, it eliminates the need for protocol conversion and transmission through the CPU, reducing access latency from milliseconds to 60 microseconds—an impressive 90% decrease. Compared to similar solutions, a single cluster achieves a total access bandwidth of 40 TB/s, which is 1.5 times higher. In typical AI programming scenarios, this architecture effectively doubles the token processing capacity of the extraction cluster and halves the time required to obtain the first token, transforming computational power into efficiency.

3. Reducing Token Costs for Economic and Scalable AI Adoption

Employing the industry's first KV-sensitive adaptive storage technology, the OceanStor M900 predicts KV cache lifecycle based on data value and intelligently distributes data across storage environments. This innovation offers a daily write capacity of up to 24 drives (DWPD), enhancing SSD durability by 16 times and ensuring stable operation over three years. By reducing media replacement and operational and maintenance costs, it lowers the long-term expenditures associated with building scalable AI extraction infrastructure, fostering quicker adoption of AI technologies.

As AI proliferates across large-scale production systems in various industries, the evolution of AI infrastructure is shifting from a compute-centric model towards tighter collaboration among computing, networking, and storage. Context memory storage will be pivotal in consistently enhancing the capacity and access efficiency of hyperscale extraction KV caches.

Topics Consumer Technology)

【About Using Articles】

You can freely use the title and article content by linking to the page where the article is posted.
※ Images cannot be used.

【About Links】

Links are free to use.