TechReaderDaily.com
TechReaderDaily
Live
Silicon · ASIC inference

Transformer-only inference ASICs move from demo to delivery

Etched, OpenAI, AMD, and the hyperscalers are betting that inference-specialized ASICs will beat general-purpose GPUs on tokens per watt as transformer-only silicon moves from demo racks to production delivery.

A chart comparing Etched Sohu chip performance metrics against other AI accelerators. dataconomy.com
In this article
  1. The tradeoff no one says aloud
  2. Whose roadmap this slots into

Quantitative trading firm Jane Street took delivery of an Etched inference rack in July and ran production workloads on it. Three weeks later, the startup behind the rack raised $700 million at a $21 billion valuation, led by the same firm that had just become its first customer. In June, TechCrunch reported Etched had $1 billion in contract bookings for Sohu, its transformer-only inference ASIC, before a single production system had left the lab. The August capital raise, SiliconANGLE reported, was the second in a month and the fastest valuation step in the current custom silicon wave.

That trajectory has made inference-specialized ASICs the most concentrated bet in semiconductor design. HPCwire wrote in August that custom AI chips had moved beyond slideware, pointing to Etched's round, AMD's acquisition of Taalas, and frontier labs designing their own processors. The common thread is not more compute. The custom parts are designed to do fewer things, for less power, at lower latency. They win by removing what the workload does not need: training loops, branch prediction, general-purpose control logic, and the software stack required to make a GPU run anything.

Ask what these chips chose not to be good at and the answer is usually everything outside inference. Sohu runs transformer decoder workloads and not much else. Taalas etches a specific model's weights into silicon, so a new architecture means new silicon. OpenAI's Jalapeño, built with Broadcom and announced in June, is fixed to the inference side of the lifecycle. Even Nvidia's Groq 3 LPX, which entered full production in August as a dedicated inference accelerator according to SiliconANGLE, exists because token generation has become a separate market from training.

The bottleneck question has shifted. For the agentic workloads that now dominate inference revenue, the limiting factor is rarely raw floating-point throughput. It is the time spent fetching model weights and waiting on memory. Groq 3 LPX uses SRAM for decode, and TechTimes reports it reaches 3,400 tokens per second in a full production rack. That number matters more than peak FLOPS. A chip that can generate a token in the time a GPU would still be scheduling the next kernel changes what an operator can run.

OpenAI's first chip arrived at the same conclusion from the model side. The company unveiled Jalapeño, designed with Broadcom, in June. SemiAnalysis benchmark results reported by TechCrunch on MSN showed it beating Nvidia Blackwell systems on both output per watt and latency per token. The numbers were specific: up to 1.9 times more output per watt and end-to-end latency reduced by up to 3.6 times, according to SemiAnalysis. OpenAI is already using its own frontier models to design silicon, CFO Sarah Friar confirmed in September, Yahoo Finance reported.

AMD's answer was to buy the weight-fused approach outright. The company agreed to acquire Toronto-based Taalas, Reuters reported on August 6, after the startup built processors custom-made for specific models. Taalas does not stop at the transformer. It bakes the model weights into the chip, making the ASIC inseparable from the model it serves. That removes the memory bottleneck for the served model and makes the part useless for any other model, the sharpest expression of the inference-specialization thesis.

Then there is the heterogeneous route. Inference service provider Parasail said in July it would pair Nvidia Hopper and Blackwell GPUs with d-Matrix Corsair accelerators to hit 10 times faster token generation, as Aspen Daily News reported. d-Matrix followed with an acquisition of Wallaroo.ai to speed deployment of mixed GPU and ASIC inference workloads, Yahoo Finance reported. The strategy presumes GPUs remain useful for prefill and routing while specialized ASICs handle decode. That is less a replacement thesis than a workload partition.

The financialization of inference hardware has also changed. In July, General Compute, an inference cloud operator, secured $400 million in debt financing from Upper90, with the loan secured against inference-specific chips rather than Nvidia GPUs, TechCrunch reported. That is the first time a lender that pioneered GPU-backed AI loans accepted non-Nvidia silicon as collateral. If specialized ASICs can be repossessed as readily as an H100 fleet, then the asset class has matured.

The tradeoff no one says aloud

The first trade is model risk. A transformer-only ASIC is a leveraged bet on one architecture. If the next frontier model moves to a state-space design, a recurrent variant, or a mixture of experts with irregular routing, the chip does not flex. Etched can tape out a new generation, but that is an 18-month turn and another very large check. A general-purpose GPU has no such cliff; its overhead is the premium.

The second trade is software. Nvidia has CUDA, NCCL, TensorRT, and a decade of tooling gravity. A startup ASIC arrives with a compiler, a kernel library, and a serving layer that must be built and tuned concurrently with the silicon. The hardware may be fast, but the deployment is only as good as the compiler's ability to map a model onto fixed dataflows without leaking utilization.

The third trade is packaging and memory. Sohu and its peers still depend on advanced packaging and the same memory supply chains that serve GPUs. Samsung, which is manufacturing Nvidia's Groq 3 LPX at scale, is now close to profitability on its foundry business from that ramp, MSN reported. The custom ASIC buildouts are beginning to pull the same packaging and memory lines that GPUs depend on, which means the new entrants do not escape the old constraints. They just shift the argument from compute to bandwidth.

Whose roadmap this slots into

Nvidia's. The company has chosen to sell into the inference-only trend rather than pretend it does not exist. Groq 3 LPX is a rack-scale decode product aimed at AI agents, the same market startups are attacking. Nvidia can offer GPUs for training and prefill while selling a dedicated accelerator for decode, then bundle software and networking on top. In that stack, a transformer-only ASIC from Etched is not a direct threat; it is another supplier competing for the same marginal token budget.

The hyperscalers' roadmaps also align. OpenAI's Jalapeño, built on Broadcom silicon, exists to cut its own cost per token and reduce dependency on merchant GPU supply. Microsoft-backed d-Matrix is pursuing ultra-low-latency inference in data centers. Broadcom already monetizes Alphabet, Meta, and OpenAI at scale, 24/7 Wall St reported. The specialized inference chip is the piece of the stack where buyers and providers have the same incentive: put the most popular query pattern into the cheapest possible silicon.

The open question is not whether these chips work. Sohu shipped, Jalapeño benchmarked, Groq 3 LPX entered full production, and AMD put money behind weight-fused silicon. The open question is whether they can scale past their first anchor customers without losing the economics that made them attractive. Etched has one disclosed customer and a $1 billion order book, but a shipping program at Jane Street is not the same as a multi-tenant inference cloud running thousands of tenants.

The previous quarter provides a cautionary signal. Cerebras Systems, which sells server systems built around wafer-scale inference chips, reported a $25.4 billion order backlog and 103 percent revenue growth, 24/7 Wall St reported. Its shares still slid after quarterly results failed to impress investors, Reuters reported. A big backlog is real, but the market is pricing whether specialized silicon can convert backlog into gross margin without becoming another box maker.

What to watch over the next 120 days: Etched's second and third production racks, not just the first. If Jane Street signs for more capacity, that tells the market the microsecond latency claim holds at scale. If OpenAI's Jalapeño enters Broadcom production in volume, the custom inference part stops being a lab artifact. And if General Compute's debt deal with Upper90 draws copycat lenders, the capital markets will have underwritten the thesis before any of these chips has proven it across a full fiscal year.

The transformer-only inference ASIC has already changed who can bid for an AI workload. The next milestone is not another fundraising round; it is a quarter of sustained production behind the first anchor customer. If that quarter arrives without a yield stumble or a software delay, the transformer-only ASIC moves from a risky option to a line item in data center procurement.

Read next

Progress 0% ≈ 9 min left
Subscribe Daily Brief

Get the Daily Brief
before your first meeting.

Five stories. Four minutes. Zero hot takes. Sent at 7:00 a.m. local time, every weekday.

No spam. Unsubscribe anytime · Privacy.