Free to read. Sign up to save tools and get alerts when they change. Plus 900+ more AI tool profiles.

Sign up free
5 min read·Updated August 18, 2026

Groq Cloud (GroqCloud) is an ultra-fast AI inference platform powered by custom Language Processing Unit (LPU) chips — purpose-built for real-time AI applications. Following NVIDIA's $20 billion December 2025 deal that absorbed founder Jonathan Ross and senior chip leadership, Groq refocused on its inference neocloud, and in August 2026 closed $350 million at a $3.5 billion valuation while becoming a certified NVIDIA Cloud Partner — adding Nvidia accelerated computing alongside the LPU fleet it still operates.

Share

Listen to this overview

Free preview · first 0:30
0:00 / 0:30

Unlock audio and more

Audio streaming, downloadable PDFs and certificates come with Plus and Pro.

Learning Objectives

  • Understand what Groq Cloud is and how its LPU architecture differs from GPUs
  • Evaluate Groq's pricing, supported models, and use cases
  • Assess the impact of NVIDIA's 2025 licensing deal on Groq's future

What Is Groq Cloud?

Groq Cloud (also called GroqCloud) is an AI inference platform that runs large language models on custom-designed Language Processing Units (LPUs) instead of traditional GPUs. The result: dramatically faster token generation — often 10 to 30 times faster than GPU-based alternatives for single-stream inference.

Founded in 2016 by Jonathan Ross (the architect behind Google's original TPU), Groq designed its chips from scratch to solve one problem: making AI inference as fast as possible. While GPUs excel at training models, Groq's LPU architecture is purpose-built for running them.

💡Key Concept

Language Processing Unit (LPU): Groq's custom chip uses SRAM (not HBM memory like GPUs) and deterministic execution — the compiler pre-computes the entire execution graph down to the clock cycle, eliminating the unpredictable memory bottlenecks that slow down GPU inference.

The NVIDIA Deal and the Inference-Cloud Pivot

In December 2025, NVIDIA entered a non-exclusive licensing agreement for Groq's inference technology, paying approximately $20 billion. The deal brought Groq founder Jonathan Ross and the bulk of senior chip-engineering leadership to NVIDIA, while GroqCloud was explicitly excluded and continues operating independently. At GTC 2026 in March, NVIDIA unveiled the Groq 3 LPU built on the licensed intellectual property — a clean separation of the chip lineage (now at NVIDIA) from the cloud service (still under Groq), with the Groq 3 expected to ship in late 2026.

The remaining Groq is now run as an inference-cloud company under CEO Adam Winter, who joined in 2024 to lead the international business and stepped up to run the company in 2026, with Matt Eng as CFO. The strategic refocus is sharp: Groq is no longer pitching as a frontier chip-design company chasing the next process node, but as an inference-cloud operator running its existing LPU fleet plus collecting royalties on the NVIDIA license.

The pitch is straightforward — inference is now a much larger market than training, and the customer-facing inference category (real-time voice, agentic browser control, low-latency function calling) is where Groq's LPU advantage on tokens-per-second economics remains intact after the chip lineage moved.

Groq closed roughly $650 million in June 2026 to fund the pivot, then a further $350 million in August 2026, led by Disruptive with NVIDIA itself participating. That August round valued the company at $3.5 billion — roughly half the $6.9 billion it carried in September 2025, before the licensing deal. Groq describes the figure as a new valuation for the post-licensing business rather than a down round, which is a fair reading: the company being priced is a cloud operator, not the chip designer that carried the earlier number.

Becoming an NVIDIA Cloud Partner

The sharpest change in Groq's posture landed in August 2026, when it was certified as an NVIDIA Cloud Partner — a designation that validates a provider's ability to deploy and operate NVIDIA accelerated computing to NVIDIA's own reference architecture. In practice it means Groq will bring NVIDIA hardware online alongside its LPUs rather than instead of them, adding capacity and newer inference technology without customers changing a line of code.

This is worth stating plainly, because the obvious reading is wrong: Groq did not abandon its own silicon. It still operates the LPU fleet — by its own account the only team running LPUs in production at scale — across 13 data centers, serving more than 6 million developers, with capacity set to grow from about 54 megawatts to beyond 200 megawatts during 2027.

What it does mean is that NVIDIA now occupies four positions at once in Groq's business: technology licensee, hardware supplier, investor, and beneficiary of Groq's infrastructure spending. For a company founded to compete with NVIDIA, the settled answer is that the fastest way to scale AI infrastructure turned out to be building on NVIDIA rather than against it.

⚠️Warning

The competitive question has shifted twice. Pre-NVIDIA-deal, the worry about building on GroqCloud was that the underlying chip lineage might lose its independent edge. Post-pivot, the question was whether Groq could hold the customer-facing inference category against NVIDIA (with the new Groq 3 lineage), AWS Bedrock, Cerebras, and Together AI. The August 2026 cloud-partner certification answers part of that by changing the terms: Groq now competes as an operator running both its own LPUs and NVIDIA hardware, so the differentiator is latency engineering and operations rather than silicon exclusivity. For a buyer, the practical read is that Groq's speed advantage on single-stream workloads no longer depends on a bet that its chip lineage stays ahead.

Speed Benchmarks

Groq's headline feature is raw inference speed:

ModelGroq Speed (tokens/sec)Typical GPU SpeedSpeedup
Llama 3.1 8B~1,345 tok/sec~100-200 tok/sec7-13x faster
Qwen 3 32B~662 tok/sec~50-100 tok/sec7-13x faster
Llama 2 70B~300 tok/sec (single card)~30-60 tok/sec5-10x faster

These speeds make Groq particularly compelling for real-time chat, agentic AI (where models make many sequential calls), and interactive applications where latency matters more than throughput.

Supported Models

As of March 2026, GroqCloud hosts a curated selection of open-source models:

ToolBest For
Llama 3.3 70BMain workhorse model for general-purpose tasks
GPT-OSS 120BReasoning flagship; newest addition
Qwen 3 32BEfficient multilingual model; replaced Mistral Saba
DeepSeek R1 Distill 70BReasoning-optimized distilled model
Llama 4 Scout 17BLatest Llama generation; mixture-of-experts
Llama 3.1 8B InstantFastest and cheapest option for simple tasks

The API is OpenAI-compatible — switching from OpenAI to Groq requires changing just the base URL and API key. No code rewrite needed.

Pricing

Free$0 (no credit card)
  • Experimentation and prototyping
  • Rate-limited
DeveloperPay-as-you-go
  • Production applications with higher rate limits
EnterpriseCustom pricing
  • Dedicated capacity and SLAs

Per-token pricing (approximate):

Model SizeInput (per 1 million tokens)Output (per 1 million tokens)
Small (8-17 billion params)~$0.11~$0.11
Mid (32-70 billion params)~$0.59-$0.99~$0.79
Large (120 billion+ params)~$1.00Higher

Cost-saving features: Batch processing saves 50% on input costs. Prompt caching gives 50% off cached input tokens.

Groq vs. Competitors

PlatformStrengthLimitation
Groq CloudFastest single-stream latency; purpose-built LPU siliconNo training; model selection limited to hosted open-source
Cerebras InferenceHigher throughput on frontier models (3,000+ tok/sec on 120 billion param models)Smaller model catalog; less developer tooling
Together AIMore model variety; fine-tuning and custom trainingRuns on rented GPUs; higher latency
AWS BedrockBroadest ecosystem; managed services; enterprise trustSlower inference; higher cost per token

Company Details

DetailInfo
Founded2016 by Jonathan Ross (ex-Google TPU architect; moved to NVIDIA December 2025)
CEOAdam Winter (joined 2024; became CEO in 2026)
CFOMatt Eng
HeadquartersMountain View, California
Latest Round$350 million (August 2026) at a $3.5 billion valuation, led by Disruptive with NVIDIA participating
Prior Rounds$650 million (June 2026); $750 million (September 2025) at $6.9 billion valuation
Total RaisedOver $3 billion (cumulative)
DevelopersMore than 6 million
NVIDIA Cloud PartnerCertified August 2026 — NVIDIA hardware deploying alongside the LPU fleet
Data Centers13, with capacity growing from about 54 megawatts toward 200-plus megawatts in 2027
Websitegroq.com

Strengths

  • Fastest inference for real-time applications — 10 to 30 times faster than GPU-based alternatives on single-stream queries
  • Free tier with no credit card required — easy to experiment
  • OpenAI-compatible API — switch from OpenAI with a 3-line code change
  • Purpose-built silicon — LPU architecture designed specifically for inference, not repurposed training hardware
  • Cost-competitive — batch processing and prompt caching reduce costs significantly

Limitations and Considerations

  • Inference only — Groq cannot train or fine-tune models; it only runs pre-trained models
  • Limited model selection — only hosts a curated set of open-source models (no proprietary models like GPT or Claude)
  • Dependence on NVIDIA has replaced independence from it — NVIDIA holds the chip lineage, supplies hardware, and is now an investor, so Groq's roadmap is tied to a company it was founded to compete with
  • Throughput vs. latency — Groq excels at single-stream speed but Cerebras outperforms on high-batch throughput workloads
  • No custom models — you cannot upload or deploy your own fine-tuned models (unlike Together AI or AWS)

Key Takeaways

  • Groq Cloud delivers the fastest single-stream AI inference available, powered by custom LPU chips purpose-built for running language models
  • The free tier and OpenAI-compatible API make it easy to try — switch from OpenAI by changing just the base URL and API key
  • NVIDIA's $20 billion December 2025 licensing deal brought Groq's founder and senior chip leadership to NVIDIA; GroqCloud was explicitly excluded and has refocused as an inference-cloud company under CEO Adam Winter and CFO Matt Eng
  • Groq raised roughly $650 million in June 2026 and a further $350 million in August at a $3.5 billion valuation, led by Disruptive with NVIDIA participating — about half the $6.9 billion it carried before the licensing deal, which Groq frames as a repricing of a different business rather than a down round
  • In August 2026 Groq became a certified NVIDIA Cloud Partner, deploying NVIDIA hardware alongside the LPU fleet it still runs across 13 data centers for more than 6 million developers — it did not abandon its own silicon
  • The strategic bet is that inference is now a much larger market than training and that real-time / agentic workloads (where tokens-per-second economics matter most) remain Groq's strongest competitive ground
  • Best suited for real-time chat, agentic AI workflows, voice and browser-control products, and latency-sensitive applications where speed matters more than throughput

Keep track of the tools you’re evaluating

  • The AI Hub on a phone: a 12-day AI Skill Streak and an expanded Content updates alert listing the saved items that changed.
  • Recommended for you on a phone: nine personalised suggestions labelled Trending in AI news, On your saved list, and Popular.
  • My AI Tools on a phone: saved tools including GitHub Copilot and OpenAI Codex, each with an Updated badge.

Swipe for Recommended for you and My AI Tools

Your AI Hub — sample data.

📰Groq Cloud in the News

Showing the 2 stories where Groq Cloud is tagged in Top AI Stories.

Other tools in Inference & Model Serving

Show 5 more →
🧭Recommended for you

Optional detours — these connect to what you just read, and your next lesson will be waiting.