At Hot Chips 2026, NVIDIA announced Groq 3 LPX — the inference accelerator from its $20B December acquisition of Groq's assets, the largest purchase in company history — is now in full production. It delivered a record 3,400 output tokens/sec in Artificial Analysis benchmarking, versus 750 tokens/sec promised by OpenAI's Cerebras-powered Ultrafast mode. First racks go live at Nebius later this year, two days before NVIDIA reports earnings.
At Hot Chips 2026, NVIDIA announced Groq 3 LPX — the inference accelerator from its $20B December acquisition of Groq's assets, the largest purchase in company history — is now in full production. It delivered a record 3,400 output tokens/sec in Artificial Analysis benchmarking, versus 750 tokens/sec promised by OpenAI's Cerebras-powered Ultrafast mode. First racks go live at Nebius later this year, two days before NVIDIA reports earnings.
What happened: NVIDIA announced Monday at Hot Chips 2026 that Groq 3 LPX, its interactive AI inference accelerator, is in full production — commercializing technology from the $20B purchase of Groq's assets in December, the company's largest acquisition on record (CNBC). Each LPX rack packages 256 Groq 3 chips (fabbed by Samsung, while TSMC makes NVIDIA's GPUs) and extends the Vera Rubin NVL72 platform by dramatically increasing token-generation rates. First deployment is neocloud Nebius, alongside Vera CPUs and Rubin GPUs, online later this year per NVIDIA senior director Dion Harris.
Key metric: Record 3,400 output tokens/sec in Artificial Analysis benchmarking on Gemma 4 31B with a 100,000-token context — the fastest performance ever recorded for the model. NVIDIA claims 4x faster responsiveness for latency-sensitive workloads than the nearest alternative platform.
Comparison: OpenAI's newly announced Ultrafast mode, powered by Cerebras, promises 750 tokens/sec — the Groq 3 LPX cited benchmark is ~4.5x that. AMD is counter-moving with its own Cerebras rack integrations. At GTC in March, Jensen Huang projected $1T in cumulative Blackwell + Vera Rubin sales through 2027 and said a quarter of his coding-focused data center space would run Groq chips — 'the rest of my data center is all 100% Vera Rubin.'
Why it matters: Agentic AI runs on inference, not training. Agents generate massive token volumes across hundreds or thousands of steps, and token speed determines whether agents feel instant or unusable — NVIDIA explicitly frames LPX as unlocking 'premium tiers of service' for latency-sensitive SLAs. Inference economics are becoming the growth engine of the AI trade, and NVIDIA now owns a dedicated product for the highest-value slice of it, launched two days before earnings.
Ondo Finance launched Ondo Intelligent Portfolios, a new onchain product category delivering portfolios built on strategies developed by BlackRock for Ondo as single transferable tokens. Three tokens went live today — High Income, Diversified Growth, High Growth — with smart-contract-level rebalancing, available to eligible investors outside the US.
US spot bitcoin ETFs have swung to net inflows of roughly $320M for 2026 (Bloomberg tally; ~$349M per Askthetape) — the first positive year-to-date print since April — after investors added about $4.6B since Aug 19, erasing a $5.69B deficit accumulated by July 13. BlackRock's IBIT captured ~$1.02B over four sessions as Bitcoin rallied ~35% from ~$64,100 to an eight-month high of $87,265.
On Monday, Sept 21, the European Central Bank switched on Pontes, a platform linking distributed-ledger markets to the Eurosystem's TARGET Services so eligible banks and market infra providers can settle tokenized bonds and funds in central-bank money. Christine Lagarde announced the go-live at Friday's Eurogroup, calling it 'a digital euro made available for banks.' It lands three days after the SEC's Innovation Exemption opened onchain trading of tokenized US stocks — the institutional tokenization race is now transatlantic.