The technology company Groq has unveiled a new advanced Language Processing Unit (LPU) designed to excel at processing large language models (LLMs), offering a level of performance that outpaces that of traditional GPU-based AI accelerators. This innovation seems to herald a shift in the AI market, which has been dominated by Nvidia’s GPUs, as more specialized alternatives are introduced.
Offering a dedicated approach to language model computing, Groq’s LPU optimizes sequential processing and eschews commonly used DRAM or HBM for onboard SRAM, delivering not only speed but also efficiency in handling LLMs. The company’s background in designing the Tensor Stream Processor (TSP) has evolved into this focus on LLMs, refining the technology for generative AI tasks based on inference.
The introduction of Groq’s LPU comes at a time when the AI industry is expanding, with companies and former Google engineers seeking to provide next-generation AI processors that surpass current industry standards. Samsung has also stepped into the ring with its new AGI Computing Lab, directed by former Google TPU developer Dr. Woo Dong-hyuk.
Groq’s LPU offers a significant advantage over legacy GPU or NPU designs due to its purpose-built nature, enabling faster generation of text sequences, lower latency, and higher sustained throughput. The LPU’s use of 230 MB SRAM per chip with a remarkable 80 TB/s bandwidth not only cuts costs by avoiding pricier memory solutions but also delivers a swift performance superior to that of GPGPU setups.
To showcase the LPU’s speed, Groq has made available a video demonstrating their chatbot’s capability to toggle between Llama 2 / Mixtral LLMs versus the performance of OpenAI’s Chat-GPT, highlighting the speed at which their LPU can generate text.
With the promise of scalability, multiple LPUs can be linked, enabling the tackling of even more demanding LLM tasks. As the AI accelerator market continues to evolve, Groq’s LPU positions itself as a game-changing contender that could redefine expectations for AI-powered language processing.






