Clockwork.io Secures $31 Million to Enhance AI Resilience and Minimize GPU Time Waste

Clockwork.io Secures $31 Million to Enhance AI Resilience



On October 5th, 2026, Clockwork.io announced a significant funding round of $31 million, aimed at improving its advanced fault-tolerant software for AI workloads. This funding comes as the company gains traction among industry giants such as LinkedIn, Together AI, and WhiteFiber, who utilize its solutions to prevent the wastage of expensive GPU hours.

Innovation in AI Workload Management



Clockwork.io specializes in ensuring continuous operation for AI training, reinforcement learning, and inference tasks, even during infrastructure failures. The company's innovative software suite offers solutions that route around failures and migrate workloads seamlessly, ensuring that GPU resources are not idly waiting during outages. According to CEO Suresh Vasudevan, “Fault tolerance is a goodput multiplier. It allows GPUs to continue performing useful work instead of waiting for recovery or repeating previously completed tasks.”

The New TorchPass Innovations



Among the notable advancements introduced alongside this funding are two new features in Clockwork.io’s TorchPass solution. These features include:

1. Multi-Node Platform Snapshots: This industry-first capability captures the entire state of a distributed AI job across all nodes without code changes, storing it for quick restoration in case of a significant failure.
2. Quick Asynchronous Application Checkpoints: Created in the background during job execution, these checkpoints facilitate faster updates of model weights to inference replicas, thereby reducing wait time and ensuring that models learn efficiently without being halted by outdated data.

Addressing the Challenge of Large-Scale AI



As AI workloads grow increasingly large, often spanning thousands of GPUs, maintaining synchronization is critical. A single hardware failure or connection issue can stymie the entire job. Meta's experience during the Llama-3 training on 16,384 GPUs highlighted this vulnerability, with an average of one unexpected interruption every three hours. The typical recovery process can squander up to 90 minutes of GPU time, a costly inefficiency that Clockwork.io seeks to eliminate.

The real-time data monitoring capabilities of Clockwork.io’s FleetLens platform allow teams to pinpoint resource bottlenecks and automatically adjust workflows, validating and mobilizing GPU resources more effectively.

Critical Partnerships and User Impact



The integration of Clockwork.io’s technologies has had a transformative impact on its partners:

  • - LinkedIn has stated that the implementation of LinkPass network fault tolerance across its AI infrastructure has prevented tens of thousands of GPU hours of downtime each month. As noted by Raghu Hiremagalur, LinkedIn’s CTO, “Thanks to Clockwork.io, a single network issue no longer takes down functioning GPUs or disrupts ongoing workloads.”
  • - Together AI is now marketing TorchPass as a service within its GPU clusters, presenting innovative demonstrations that showcase job continuity despite simulated failures.
  • - WhiteFiber (NASDAQ: WYFI) is expanding the use of Clockwork.io's software to enhance the reliability of its growing global GPU-as-a-service infrastructure, emphasizing the importance of validating every node and connection before going live. Tom Sanfilippo, CTO at WhiteFiber, expressed that anticipatory error detection leads to faster deployment and more reliable customer experiences.

Future Prospects and Funding Utilization



The recent funding round, led by Premji Invest, Wing Venture Capital, and others, brings the total investment in Clockwork.io to $73 million. The capital will be utilized to accelerate the deployment of their fault tolerance suite and enhance acceptance in enterprise applications, ultimately broadening partnerships with cloud service providers.

Clockwork.io is looking to solidify its role as a decisive player in the AI infrastructure landscape, meeting the challenges of sustainability and efficiency that come with increased hardware demands. As Greg Papadopoulos of NEA noted, “The only thing that can scale perfectly is unreliability. When you pack enough GPUs into a single system, something will inevitably fail. Clockwork.io transforms these errors into manageable maintenance events.”

As their software continues to evolve, Clockwork.io stands at the forefront of enhancing the reliability of AI workloads, ensuring that businesses can harness the full capability of their GPU investments without interruption.

For more information, visit Clockwork.io.

Topics Consumer Technology)

【About Using Articles】

You can freely use the title and article content by linking to the page where the article is posted.
※ Images cannot be used.

【About Links】

Links are free to use.