AI & ML


From the editor's desk: Groq – the future of AI processing?

28 March 2025 AI & ML


Peter Howells, Editor

For the past few years, the world has been hit by a storm of AI-generated information, mostly using generative pre-trained transformer (GPT) models to perform the AI inferencing. These models are excellent at performing large language model (LLM) requests, but they do have one drawback. The response time or lag is noticeable. This is largely due to the hardware that these models are being processed on, namely GPUs.

Many of the GPUs running AI models in professional data centres are Nvidia’s A100 (or similar) series which contain thousands of CUDA cores, many more than the handful of processing cores in a standard CPU. These CUDA cores work together to answer the language requests directed at them as they are designed for parallel processing and are optimised for tasks like scientific simulations. But are they really optimised?

Groq seems to be the new kid in all this AI hoopla, but they have been around since 2016 when the company was founded by a group of former Google employees led by Jonathan Ross, one of the designers of the Tensor Processing Unit (TPU), and Douglas Wightman, an engineer at Google X. The TPU is an AI accelerator Application-Specific Integrated Circuit (ASIC), a custom-designed chip tailored for a specific task. These ASICs offer optimised performance and efficiency compared to general-purpose processors.

And this is where the story gets exciting. The Groq AI model runs on ASICs as opposed to GPU architecture to deliver similar responses to the current slew of GPT models in use. Groq’s architecture is developed to expedite machine learning workloads, providing unparalleled speed and efficiency. This is a big deal – Groq needs much less energy to answer these same requests, and more importantly, does it with seemingly no lag. This last property is down to the speed at which the ASICs perform these ‘application-specific’ tasks.

Real-world testing by myself bears this out. When asking exactly the same technical question to both chatGPT 3.5 and 4.0 models and also to Groq, and then comparing the response times, I can without a doubt say that Groq certainly has minimal lag compared to the GPT models. The information produced in the responses is presented in a different format, but compares favourably with each other. Groq’s response is almost immediate whereas the other models take a few seconds before beginning to display an answer.

The introduction of Groq’s ASIC-based approach to AI inferencing marks a significant shift in the landscape of LLMs. By prioritising speed and efficiency, Groq is challenging the current dominance of GPU-driven AI, offering near-instantaneous responses while consuming less power. As AI applications continue to expand, this technology could redefine the way we interact with AI systems, setting a new benchmark responsiveness.

Whether this signals a broader industry shift remains to be seen, but one thing is clear – Groq has introduced a compelling alternative that demands attention.


Credit(s)



Share this article:
Share via emailShare via LinkedInPrint this page

Further reading:

Compact direct Time-of-Flight 3D LiDAR module
Altron Arrow AI & ML DSP, Micros & Memory
The VL53L9 from STMicroelectronics is the first direct Time-of-Flight (dToF) 3D LiDAR all-in-one module in ST’s portfolio, offering a resolution of 2,3K zones, wide field of view, on-chip processing, 100 frames per second, and sensing range from 5 centimeters to 9 meters.

Read more...
SoC brings scalable edge AI to life
Altron Arrow AI & ML
The NXP i.MX 937 applications processor is optimised for performance without excessive power draw and bridges the gap between entry-level chips and high-end processors.

Read more...
Wall-mount cooling for Edge IT
AI & ML
Designed to run 24/7 in critical compute environments, Vertiv CoolPhase Wall delivers up to 60% greater airflow than comfort cooling.

Read more...
From the editor's desk: Local can be international
Technews Publishing Editor's Choice News
Welcome to the July 2026 issue of Dataweek. As you can see from this introduction, Dataweek’s regular editor, Peter Howells, is on extended leave and I am filling the void in his editor’s column – hopefully without being too boring.

Read more...
Compact direct Time-of-Flight 3D LiDAR module
Altron Arrow Opto-Electronics Electronics Technology AI & ML
The VL53L9 from STMicroelectronics is the first direct Time-of-Flight (dToF) 3D LiDAR all-in-one module in ST’s portfolio, offering a resolution of 2,3K zones, wide field of view, on-chip processing, 100 frames per second, and sensing range from 5 centimeters to 9 meters.

Read more...
From the editor's desk: The art of measuring the truth
Technews Publishing Editor's Choice News
All electronic measurements are a lie. The trick is making the lie as small as possible.

Read more...
Edge AI improves condition monitoring
EBV Electrolink AI & ML
STMicroelectronics has introduced the IIS3DWB10IS, an intelligent MEMS vibration sensor designed for high-performance industrial condition monitoring applications.

Read more...
Tiny SoM integrates NXP i.MX 91
AI & ML
The QS91 module from Direct Insight utilises the low-cost single core i.MX 91 to deliver secure, energy-efficient Linux capabilities for edge applications.

Read more...
Robotics platform powers intelligent systems
Vepac Electronics AI & ML
The CEXD-INTRBL open robotics development system combines high-performance AI processing, sensor integration, and motion control into a compact embedded platform.

Read more...
From the editor's desk: Pricing surge reshapes engineering reality
Technews Publishing News
The recent and continuing surge in memory prices has become more than a supply-chain story confined to global semiconductor markets. We have watched in disbelief as the ASP of memory has risen by over ...

Read more...









While every effort has been made to ensure the accuracy of the information contained herein, the publisher and its agents cannot be held responsible for any errors contained, or any loss incurred as a result. Articles published do not necessarily reflect the views of the publishers. The editor reserves the right to alter or cut copy. Articles submitted are deemed to have been cleared for publication. Advertisements and company contact details are published as provided by the advertiser. Technews Publishing (Pty) Ltd cannot be held responsible for the accuracy or veracity of supplied material.




© Technews Publishing (Pty) Ltd | All Rights Reserved