San Francisco's Wafer closed a $40M Series A at $200M+ after showing its AI agents push AMD GPUs to 80% of Nvidia B200 throughput at under half the cost.
Key Takeaways
- Wafer's $40M Series A values the company at over $200M, a 50x jump from its $4M seed round closed just five months earlier.
- The startup's autonomous agents tuned AMD's MI355X to ~80% of Nvidia B200 throughput at under half the per-token cost.
- Wafer rejected acquisition offers from multiple cloud and inference providers before committing to the independent raise.
Lead
Wafer, a San Francisco startup that deploys autonomous AI agents to continuously optimize GPU inference workloads, closed a $40 million Series A on September 1, 2026. The round was co-led by Marathon and Chemistry, with Wing, AMD Ventures, Outset Capital, Fifty Years, and Y Combinator also participating. Post-money valuation exceeds $200 million, a 50x step-up from Wafer's $4 million seed round, which closed just five months earlier in April 2026.
What Does Wafer Actually Build?
Wafer's product is a set of autonomous software agents that monitor live inference deployments and re-tune GPU workloads continuously, in real time. Traditional kernel optimization requires manual tuning cycles applied at deployment. Wafer's approach applies ongoing adjustments as traffic patterns, model versions, and hardware conditions shift.
The clearest demonstration came in July 2026, when Wafer tuned Z.AI's GLM-5.2 model to run on AMD's MI355X GPU, reaching approximately 80% of the throughput Nvidia's B200 delivers - at under half the cost per token.
Why Does Closing the AMD-Nvidia Gap Matter?
Nvidia's B200 is the effective benchmark for inference throughput, and for organizations evaluating alternatives, a double-digit performance gap has typically ended the conversation before cost enters it. At 80% of B200 throughput, the MI355X running Wafer-optimized workloads falls within acceptable margins for many production deployments. The cost differential can exceed 50%.
AMD Ventures joined the round as an investor. AMD has produced competitive silicon that repeatedly underperformed at the software layer, and its presence on the cap table suggests Wafer's optimization gains are considered reproducible, not benchmarking artifacts.
How Did the Round Come Together So Quickly?
Wafer turned down acquisition offers from multiple cloud and inference infrastructure companies before closing on the Series A. That decision, combined with a 50x valuation expansion in five months, reflects how contested software-layer inference optimization has become as enterprises treat compute costs as a primary constraint. The company's angel list spans the CEOs of Cloudflare, Vercel, Deepgram, and Bot, the CTO of DoorDash, the COO of Notion, and Jeff Dean, now CEO of DiscoveryLoop.
The April-to-September timeline is aggressive by any measure. A $4 million seed and a $200 million-plus Series A within a single calendar year compresses what has historically been a 12-to-18-month product validation cycle into less than half that.
Competitive Context
Inference optimization has attracted significant capital as AI compute costs remain the largest variable expense for enterprises running large language models at scale. Wafer's autonomous-agent model differs from earlier tools that required manual tuning, but the question of whether that distinction holds as hardware vendors ship their own software stacks - AMD's ROCm roadmap chief among them - will define the company's defensibility over the next 18 months.
The funding will go toward automating more of the optimization loop and expanding hardware support beyond the current AMD lineup.
Outlook
Wafer's case rests on two variables that could shift independently: AMD closing more of the raw performance gap with Nvidia, and Wafer's agents maintaining an edge over vendor-supplied optimization tools as those tools mature. If both hold, the cost argument for AMD-plus-Wafer becomes difficult to dismiss at enterprise scale. The 50x seed-to-Series-A trajectory in five months suggests Marathon, Chemistry, and their co-investors believe both conditions will persist long enough to matter.



