San Francisco's Wafer has secured $40 million in Series A funding, co-led by Marathon and Chemistry, after proving AMD's MI355X can match 80% of Nvidia B200 throughput at less than half the price.
- Wafer's autonomous inference agents hit 2,626 tokens/second aggregate on AMD MI355X running GLM-5.2, undercutting Nvidia Blackwell on cost by more than 2x.
- The $40M round values Wafer above $200M, a 50x jump from its $4M seed just five months ago.
- The company turned down multiple acquisition offers from cloud and inference providers before closing the round.
Lead
Wafer, a San Francisco AI infrastructure startup, announced a $40 million Series A on September 1, 2026, co-led by Marathon and Chemistry. The round drew in Wing, AMD Ventures, Outset Capital, and prior backers Fifty Years and Y Combinator. At a post-money valuation north of $200 million, the raise marks a 50x step up from the company's $4 million seed closed in April 2026 - five months and two hardware generations ago.
What Does Wafer Actually Do?
Wafer deploys AI agents to write low-level hardware code in real time, continuously retuning inference workloads as the underlying hardware and model topology shift. The goal is maximizing intelligence per watt, not just raw throughput. It does not require engineers to manually profile or hand-tune kernels.
Founders Emilio Andere and Steven Arellano, both University of Chicago graduates, built the system through the Y Combinator Summer 2025 batch. Andere previously worked at Argonne National Laboratory and the University of Chicago's SAND lab, as well as AI research startup Elicit. Arellano spent time at Two Sigma, Google, and machine learning infrastructure firm SEI Labs.
Why Is the AMD Benchmark Significant?
That question gets to the core of Wafer's commercial proposition. On AMD's MI355X, Wafer's optimization agents achieved 2,626 tokens per second aggregate and 213 tokens per second single-stream running GLM-5.2. By comparison, Nvidia's B200 delivers roughly 25% more throughput on equivalent tasks but at more than double the cost per unit of compute.
The 80% parity figure matters because it crosses an informal industry threshold. Below roughly 70% of Nvidia performance, non-Nvidia hardware remains a niche option for latency-tolerant workloads. Above 80%, it becomes a credible primary stack for most enterprise inference. Wafer's agents appear to have cleared that bar without any hardware changes on AMD's side.
The AMD Ventures participation in the round adds a layer of strategic interest. AMD has spent years narrowing the gap with Nvidia on raw silicon, and software-layer optimization closing the remaining gap would significantly alter the competitive calculus in AI infrastructure procurement.
What Happens When Cloud Providers Come Calling?
Wafer fielded and declined multiple acquisition offers from cloud providers and inference companies before closing this round. The decision to stay independent after receiving acquisition interest from large infrastructure buyers is a meaningful signal about where the company sees its near-term ceiling.
Cloud providers operate enormous inference fleets and would derive immediate cost reduction from software that can squeeze more throughput from non-Nvidia silicon. The fact that more than one made an approach suggests the technology is considered credible at the enterprise level, not just on benchmarks.
Staying independent allows Wafer to price across the market rather than optimizing for one platform. It also keeps the option open for a higher-priced exit once more performance data accumulates across production workloads.
Investor and Advisor Signal
The cap table includes strategic and financial investors that map directly to Wafer's target market. AMD Ventures' participation is the clearest industrial endorsement. The angel roster - Jeff Dean, Guillermo Rauch, Kyle Vogt, Matthew Prince, Akshay Kothari, and Andy Fang - spans chip-level machine learning, developer infrastructure, robotics, and enterprise software. That breadth suggests the company is being positioned as horizontal infrastructure rather than a point solution.
The 50x valuation step-up in five months will draw scrutiny. At that pace, the next milestone has to be production revenue at scale, not further benchmarks.
Outlook
Wafer enters Q4 2026 with $40 million, a validated benchmark that clears the 80% parity threshold with Nvidia, and strategic backing from AMD's own venture arm. The near-term test is whether benchmark performance holds in messy production environments, where model switching, batching patterns, and hardware heterogeneity complicate the optimization surface.
If it does, the company sits at an increasingly valuable intersection: inference costs are the largest and fastest-growing line item in AI deployment, and hardware competition below Nvidia depends on software closing the gap that silicon alone cannot. Wafer's pitch is that agents can do that closing work continuously and automatically. The round suggests enough early evidence exists to take that pitch seriously.



