Huawei Unveils OceanStor M900 Context Memory System for AI Advancements

Huawei Introduces OceanStor M900 for Advanced AI Capabilities



During the HUAWEI CONNECT 2026 event, David Wang, Huawei's vice chairman and rotating chairman, officially unveiled the OceanStor M900, a state-of-the-art context memory system aimed at accelerating AI inference in hyperscale data centers. This innovative solution redefines AI infrastructure by enabling SuperPoD systems to utilize fully shared memory space, offering capacities measured in petabytes (PB) and impressive performance rates reaching terabytes per second (TB/s).

As AI technology matures, 2026 marks a significant shift from just breakthrough technological achievements to widespread deployments. AI applications have evolved from simple chatbots to sophisticated agents capable of performing complex tasks autonomously. Such advancements highlight the dawn of a new era defined by agent-based artificial intelligence.

With large models reaching a staggering scale of 10 trillion parameters, SuperPoD systems present themselves as the optimal choice for AI infrastructure. Currently, mainstream large models are already handling context windows that exceed a million tokens, and complex multi-step inferences have become standard practice, leading to a continuous increase in the volume of cached KV data generated during inference. This scenario has exposed the limitations of embedded memory and DRAM in terms of capacity and cost-effectiveness, prompting a consensus on the critical need for a multi-level storage system that allows the integration of chip embedded memory, DRAM, and SSDs into a fully shared memory space boasting substantial capacity.

Overcoming Memory Capacity Challenges



Huawei’s OceanStor M900 emerges as a transformative context memory system that addresses capacity constraints, facilitating the processing of extensive contexts and multi-stage inferences. Utilizing a UnifiedBus network, the M900 generates a global, multi-level KV cache measured in petabytes, ensuring direct data access. This architecture optimizes computational power of SuperPoD systems, thereby speeding up AI inference in hyperscale data centers.

OceanStor M900 offers three pivotal capabilities:

1. Breaking Memory Capacity Boundaries: By leveraging the fast UnifiedBus connections, KV cache resources are pooled globally and shared through multi-level storage. The memory cache in SuperPoD systems is expanded from embedded memory and DRAM to SSDs, enabling a single cluster to achieve an astounding capacity of up to 64 PB. The available KV cache per NPU processor climbs from gigabytes to terabytes, empowering systems to store, share, and reutilize a larger volume of contextual data while significantly enhancing KV cache hit ratios.

2. Boosting Inference Performance: The M900 integrates an industry-first architecture linking the CPU, network controller, and NAND controller. It natively supports KV semantics, allowing direct connections between the NPU processor and SSDs while eliminating the need for protocol conversions and data transfers through the CPU. This innovation reduces access delays by 90%, dropping from milliseconds to 60 microseconds. A single cluster can provide a cumulative access bandwidth of 40 TB/s, which is 1.5 times greater than comparable solutions. For standard AI applications related to programming, this architecture doubles cluster inference bandwidth based on processed tokens and reduces the time taken to generate the first token by half, translating computational power into actual productivity.

3. Cost-Effective Scaling of AI Deployments: OceanStor M900 features the industry-first adaptive storage technology that recognizes KV data. This intelligent system anticipates the lifecycle of KV data in the cache based on its value and strategically distributes them across varied storage mediums. This technology enables up to 24 full disk writes per day (DWPD), amplifying SSD durability by 16 times and ensuring stable operations for three years. By minimizing the necessity for media replacement and lowering operational maintenance costs, it effectively reduces the long-term infrastructure expenditures critical for widespread AI deployment and accelerates the adoption of artificial intelligence.

As AI takes a more central role across major production systems in every sector, AI infrastructure is transitioning from a power-centric model to a cohesive integration of computational resources, networking, and storage. Context memory systems like OceanStor M900 will be crucial in continuously enhancing the capacity of hyperscale KV caches employed in inference, ensuring efficient access to stored data while driving AI innovations into the future.

Topics Consumer Technology)

【About Using Articles】

You can freely use the title and article content by linking to the page where the article is posted.
※ Images cannot be used.

【About Links】

Links are free to use.