DriveNets and AMD Unveil New Architecture for Enhanced AI Cluster Performance

In a significant development for AI infrastructure, DriveNets, renowned for its large-scale networking solutions, has officially released an end-to-end reference architecture designed to enhance the performance and efficiency of AI clusters. This architecture integrates the power of AMD's Instinct™ MI350 series GPUs with DriveNets' AI Fabric, offering a well-validated blueprint for constructing high-performance AI systems.

This innovative approach promotes both scale-out and scale-across architectures, enabling a unified solution for front-end and storage networking. By adopting automated orchestration and full-stack integration services, institutions can achieve faster deployments and seamless end-to-end scaling. The newly published reference architecture underscores the strategic alliance between DriveNets and AMD, particularly after AMD's involvement as a strategic investor in DriveNets' recent $410 million Series D funding round.

Accompanying the reference architecture is a deployment guide. This guide presents a system-level approach to assembling, implementing, and fine-tuning expansive AI GPU clusters. The documentation thoroughly addresses compute and networking design, providing insights into optimizing performance across the AMD ROCm™ software ecosystem, RCCL collective communications, and other integral technologies.

One notable aspect of the reference architecture is the extensive benchmarking it includes. It assesses workloads related to inference, training, network isolation, and fabric resiliency on clusters harnessing the MI355X GPUs. These benchmarks yield verified, reproducible evidence of the platform's readiness for large-scale AI deployments. Various network topologies and detailed technical guidance ensure that users can optimize their performance and scalability effectively.

The highlighted benchmark results reveal that DriveNets' AI fabric enhances throughput by approximately 5% and reduces the time to first token by 10-15% compared to existing industry benchmarks. In multi-node configurations, the platform meets rigorous production service-level objectives, including maintaining inter-token latencies below 20 milliseconds and generating at least 50 output tokens per second per user.

Furthermore, that resiliency testing has shown steady performance under concurrent RDMA operations, revealing no degradation during transient link interruptions. The fabric bandwidth remained reliable throughout, demonstrating stability and performance comparable to leading benchmarks available in the market.

The architecture crafted through this collaboration sets the stage for an effective lifecycle for large language models (LLMs). It focuses on maximizing GPU efficiency while simultaneously lowering the total cost of ownership, thus providing a cost-effective alternative to traditional integrated AI stacks. Ido Susan, co-founder and CEO of DriveNets, emphasized that the industry is witnessing a pivot from single-vendor frameworks toward more diversified, multi-vendor infrastructures, further highlighting the crucial role of networking in making this transition successful.

Both AMD's Arvind Balakumar and DriveNets are focused on assuring end-to-end optimization of the software stack. The collaboration extends to direct engagements with joint customers, including those developing LLMs and NeoCloud service providers. They have established a joint lab where customers can test their own workloads, validating the efficacy of this new architecture.

DriveNets continues to be at the forefront of networking solutions within the realm of AI, powering networks for major players including ATT and Comcast, which together handle a substantial portion of total internet traffic in the U.S. Meanwhile, AMD's commitment to high-performance computing places it as a leader in addressing the escalating demands of modern AI workloads. This partnership signifies a new era of cooperative innovation in the AI infrastructure sector, promising greater performance, efficiency, and cost-effectiveness for AI applications worldwide.

Topics Consumer Technology)

【About Using Articles】

You can freely use the title and article content by linking to the page where the article is posted.
※ Images cannot be used.

【About Links】

Links are free to use.