The chip startup is gambling that the most expensive part of an AI server can be designed around. Positron, an AI inference chip company, said on Friday it raised $875 million in a Series C round that values it at $5 billion, with NEA, Atreides Management, Valor Equity Partners, Andra Capital, and SemiAnalysis Capital among the co-leads.
The company has scaled quickly. Positron is based in Reno, Nevada, and is led by Mitesh Agrawal, who has described the design goal as giving a single chip nearly ten times the memory capacity of a conventional GPU. In February 2026 it raised a $230 million Series B at a valuation above $1 billion. The new round lifts that mark to $5 billion, a pace that tracks how fast inference has moved from an afterthought to the main event in AI hardware.
The money is aimed at a single architectural bet. Positron’s next chip, Asimov, is built around what the company calls a memory-first design, using consumer-grade LPDDR5X instead of the high-bandwidth memory that every major accelerator now depends on. HBM is in short supply and priced accordingly, and Positron is betting it can sidestep both.
The company claims the approach pays off in efficiency. Positron says its architecture can reach roughly 90 percent bandwidth utilization, against about 30 percent for common GPU designs. If those numbers hold, the chip would extract far more work from the same memory, which is the point of choosing cheaper parts.
The scale ambitions are large. Asimov is planned to tape out on TSMC’s N3P process by the end of 2026 and enter mass production in the second half of 2027. A single chip would support between 288 gigabytes and 2.3 terabytes of memory, enough, the company argues, to run trillion-parameter models and contexts of more than 10 million tokens on a single node.
That is a direct pitch to inference, the part of AI spending that is growing fastest as models move from research to production. Inference favors low cost per token and high memory throughput, and Positron is designed for exactly that workload rather than for training.
The shift toward reasoning models has sharpened the demand. Models that think step by step before answering burn far more tokens per query, which means more time running on inference hardware than on training runs. Every model vendor now prices and plans around tokens served, and that arithmetic favors a chip that is cheap per token and generous with memory, which is exactly the trade Positron is making.
The company already has a commercial foothold. Its first Atlas systems, more than 50 racks, are being deployed on Oracle Cloud, giving Positron a reference customer in one of the hyperscalers most aggressively courting AI workloads.
The backdrop is the HBM shortage itself. The memory that accelerators need is sold out, and prices have risen sharply, which has pushed chip designers to reconsider the assumption that HBM is the only option. A cheaper memory path has moved from a technical preference to a commercial opportunity.
The HBM bottleneck is concentrated in three names. SK Hynix, Samsung, and Micron produce nearly all of the high-bandwidth memory that accelerators rely on, and demand from Nvidia and others has kept the product sold out through successive price increases. That concentration is what makes the alternative attractive: LPDDR5X is made at commodity scale for phones and laptops, so supply is deep and prices are far lower. The catch is that it was never designed for server duty, which is where Positron’s engineering claims will be tested.
Analysts said the bet is not without risk. LPDDR5X is abundant and inexpensive, but it was designed for phones, and building a server chip around it means solving bandwidth, reliability, and software problems that the HBM ecosystem has already absorbed. The 90 percent utilization claim will be tested only when customers run real workloads.
Positron also competes against a crowded field of inference specialists, each claiming to undercut the incumbent GPU economics. The difference here is the memory choice, and the $875 million round suggests investors are willing to fund that specific wager at scale.
The field of inference specialists is crowded and well funded. Groq raised $640 million in 2024 at a $2.8 billion valuation on a bet that a simpler, memory-focused design could undercut GPU economics, and Cerebras has pursued a different route with a wafer-scale chip sized for single-node large models. Others, including d-Matrix and Etched, have pitched purpose-built silicon for specific parts of the inference workload. Positron’s differentiator is narrower and more specific: the memory itself.
The Oracle deployment is the kind of reference customer that matters in this market. Hyperscalers buy chips only after they have run months of real workloads, and a named cloud win shortens that cycle for everyone else. Positron’s pitch to the next customer is simple: the memory you cannot buy is not your bottleneck, and the memory you can buy is now enough.
The next checkpoints are clear: a successful tapeout by the end of 2026 and production in 2027. Between now and then, the company must turn a memory-first thesis and a $5 billion valuation into silicon that customers can measure. The HBM shortage has opened the door; Positron now has to walk through it.


