Chipmaker Nvidia Corp. says its dedicated artificial intelligence inference accelerator Groq 3 LPX has now entered full production as it strives to maintain its dominance in the world of AI compute.
Nvidia's dedicated inference accelerator Groq 3 LPX enters full production to supercharge AI agents
Chipmaker Nvidia Corp. says its dedicated artificial intelligence inference accelerator Groq 3 LPX has now entered full production as it strives to maintain its dominance in the world of AI compute.
SiliconANGLE
Publisher
Aug 24, 2026 at 3:00 PM UTC · 5 min de lectura

The new chip, announced today at Hot Chips 2026, is described as a purpose-built extension to Nvidia’s flagship Vera Rubin data center platform. According to Nvidia, it’s designed to deliver ultra-fast token generation speeds, which are necessary to run highly responsive agentic AI workloads. The chipmaker said the neocloud provider Nebius Group N.V. has already signed on as the first customer to commit to using the new chip.
Inference is the AI industry’s lingo for the process of running fully trained AI models in production, and it’s increasingly focused on autonomous AI agents that can perform tasks on behalf of humans. These agents must be able to do everything from reason and plan, write and execute code, inspect system files and use third-party tools in continuous loops. They can quickly crunch through thousands of tokens across these complex chains, but this sometimes results in massive “decode latency” that can create frustrating delays.
To prevent this from happening, AI data centers need more specialized compute architectures that can disaggregate the enormous context processing from token generation, in order to increase the speed at which AI agents can reason and work.
Article Intelligence
Topics
Sponsored
AdNewsLayer Premium
Unlock deeper intelligence.
Ad-free reading, exclusive research, and real-time onchain insights.
Go Premium
