Clockwork.io raised $31M from Premji Invest, Wing VC and Seligman Ventures to keep AI training and inference running through GPU and network failures.
Key Takeaways
- Clockwork.io raised $31M on October 5, 2026, co-led by Premji Invest, Wing Venture Capital and Seligman Ventures.
- NEA and e& Capital, both existing investors, joined the round. Total funding reaches $73M. Valuation was not disclosed.
- LinkedIn, Together AI and WhiteFiber run the Palo Alto company's fault-tolerance software in production.
Lead
Clockwork.io, a Palo Alto software company, announced a $31M funding round on October 5, 2026. Premji Invest, Wing Venture Capital and Seligman Ventures co-led it, with NEA and e& Capital participating again. The announcement put total funding at $73M, which implies about $42M raised before this round. The company did not disclose a valuation.
The money will go toward rolling out its fault-tolerance software across AI training, inference and reinforcement learning, expanding enterprise adoption, and scaling delivery through cloud partners. The pitch is narrow: when one GPU or network link in a large cluster fails, thousands of healthy GPUs should not sit idle waiting for a restart.
What Does Clockwork.io Actually Sell?
Clockwork sells a software layer between GPUs and AI workloads that detects hardware and network faults and works around them. It has three main products. LinkPass reroutes traffic around failed network links. TorchPass migrates a training job off a failing GPU while the job keeps running. TorchSnap, launched recently, captures a snapshot of an entire multinode distributed inference job without code changes, so it can restart quickly.
The company also offers FleetLens, a telemetry and validation product for clusters. Its technology traces back to a 2018 Stanford research project. Founders include Balaji Prabhakar, Yilong Geng, Deepak Merugu and VMware co-founder Mendel Rosenblum, who serves as chief scientist. Suresh Vasudevan, previously CEO of Sysdig and Nimble Storage, is chief executive.
Why Do GPU Failures Cost So Much?
Large training jobs run in lockstep, so one failed component can stall the whole cluster. Meta reported hardware or infrastructure interruptions roughly every three hours during a 54-day Llama 3 training run on 16,384 GPUs. Recovery from a checkpoint can take up to 90 minutes, according to Clockwork's announcement, during which healthy GPUs produce nothing.
At current GPU rental prices, those idle hours are expensive. Clockwork cites SemiAnalysis benchmarks showing TorchPass cuts training goodput loss from 14% to under 3%. That figure comes from the vendor's own materials and covers a specific benchmark setup, so production results will vary by cluster and failure rate.
Who Is Using It?
LinkedIn has deployed LinkPass across its AI infrastructure and says it prevents tens of thousands of GPU-hours of downtime each month. Together AI offers TorchPass as a service, and WhiteFiber (NASDAQ: WYFI) uses the technology for cluster reliability audits. Other named customers include Wells Fargo, Nebius, Nscale and DCAI.
That customer list is the strongest evidence in the announcement. LinkedIn's monthly figure is a statement about its own fleet, and Clockwork did not publish a dollar value for it. Cloud and neocloud operators, who sell GPU time by the hour, have a direct financial reason to adopt tools that raise usable uptime.
What Does This Round Say About Clockwork's Trajectory?
The round signals investor appetite for AI infrastructure software that raises utilization rather than adds raw capacity. Clockwork began in 2021-era networking and clock-synchronization work, when its NEA-led $21M Series A in March 2022 funded network visibility and time-sync technology. The company has since moved toward failure recovery as GPU clusters grew.
Two details temper the enthusiasm. The announcement gives no valuation and no revenue figures, and press coverage of the round used inconsistent series labels, so the stage is unconfirmed. Three lead investors sharing one round is also common for companies whose growth is still being proven. Investor commentary from NEA frames the product as treating failure as the normal state of large clusters, which matches how operators describe the problem.
What Comes Next for Fault-Tolerance Software?
Competition is likely to come from the cloud providers and framework developers themselves. Checkpointing improvements in training frameworks, and resilience features built into GPU vendors' stacks, could absorb part of what Clockwork sells. A standalone vendor needs to show that its migration and snapshot tools work across mixed hardware and do not add overhead to jobs.
Inference is the second test. Training failures are well documented, but inference snapshots for distributed workloads are newer, and TorchSnap's adoption will show whether the market extends beyond training clusters. Clockwork's plan to scale through cloud partners suggests it expects distribution, not direct enterprise sales, to drive growth.
Outlook
Clockwork.io enters its next phase with $31M in fresh capital, $73M raised in total and production deployments at LinkedIn, Together AI and WhiteFiber. The commercial case rests on a measurable problem: idle GPUs during recovery. Whether the company defends its position will depend on published customer savings, an independent view of its benchmarks, and how quickly larger platforms add similar features.



