Andreessen Horowitz leads Gimlet Labs' $300M Series B at a $3B valuation as the startup scales its multi-silicon inference cloud for faster agentic AI.
- Gimlet Labs raised $300M in Series B on Sept. 4, 2026, bringing its valuation to $3B and total funding to $392M.
- a16z led the round; Sapphire Ventures, M12, Arm, Menlo Ventures, and Factory also participated.
- Gimlet's multi-silicon architecture disaggregates AI model inference across chip types to cut latency in agentic workflows.
Lead
Gimlet Labs, the AI inference startup that emerged from stealth only last October, raised $300 million in a Series B round on September 4, 2026, led by Andreessen Horowitz. The round values the two-year-old company at $3 billion and brings total capital raised to $392 million. Participating investors include Sapphire Ventures, M12 (Microsoft's venture arm), Arm, and returning backers Menlo Ventures and Factory.
The capital will fund expansion of Gimlet's multi-silicon cloud capacity - including securing what the company describes as gigawatts of high-density data center pipeline - alongside hiring across systems architecture, distributed networks, and hardware compiler teams.
What Is Multi-Silicon Inference?
The core claim is that no single chip type handles every phase of AI inference optimally. Gimlet's platform disaggregates model execution, routing each phase - prefill, decode, memory-intensive operations - to whichever silicon handles it best. The company works with NVIDIA, AMD, Intel, Arm, Cerebras, and d-Matrix to integrate their respective chips into a unified scheduling layer.
The pitch is that homogeneous GPU clusters, which dominate today's inference infrastructure, carry latency penalties when running multi-step agentic workflows. In agentic tasks, where a model calls tools, re-reads context, and generates iteratively across many sequential steps, those penalties compound. Gimlet claims heterogeneous silicon scheduling reduces end-to-end latency for these workloads. Independent benchmarks validating those claims remain limited.
Why Does Agentic AI Demand Different Infrastructure?
Single-turn inference and agentic inference are structurally different problems. A chatbot responding to a one-shot prompt is latency-tolerant by comparison. An autonomous agent running a pipeline of interdependent calls - each waiting on the previous output - sees latency errors stack multiplicatively. A 20% per-step improvement translates into a substantially larger reduction in total wall-clock time across a ten-step workflow.
That framing explains why Andreessen Horowitz and co-investors are willing to price a seed-to-$3-billion arc in under 24 months. Inference speed is now a key competitive differentiator for enterprise AI buyers, and whoever controls the routing layer between chips and models sits at an unusually high-leverage point in the stack.
From Stealth to $3 Billion in Under a Year
Gimlet's funding trajectory warrants scrutiny. The company launched publicly in October 2025 with a $12 million seed led by Factory. Five months later, in March 2026, Menlo Ventures led an $80 million Series A. Six months after that, the company is announcing a $300 million Series B at a valuation that implies roughly 37 times the seed-round capital now in play.
The Series A valuation was not disclosed, so the step-up multiple into the Series B is unclear. What is clear is that inference optimization is seeing valuation inflation driven by customer urgency rather than conventional revenue multiples. Whether Gimlet's commercial metrics support a $3 billion price tag is a question the company has not answered publicly.
The co-founders bring credible technical pedigree. CEO Zain Asgar was a GPU architect at NVIDIA and an engineering lead at Google AI before founding Pixie Labs, which New Relic acquired in 2020. The broader founding team spans chip design, compiler engineering, and distributed systems.
Outlook
Gimlet enters the second half of 2026 with enough capital to pursue infrastructure deals at a scale most startups cannot reach. The central tension is whether multi-silicon orchestration becomes a durable moat or a commodity feature absorbed into hyperscaler inference platforms. Microsoft, Google, and Amazon all operate their own inference infrastructure and have no structural incentive to cede that routing layer to a third party.
The company's path to defensibility runs through compiler-level optimization that hyperscalers cannot replicate without years of dedicated partnership. Arm's participation as an investor, not merely a chip vendor, suggests Gimlet is pursuing exactly that kind of integration. Whether it holds will be visible in the terms of the next round.



