The short answer
Groq is billed per token. It has a free tier: Free dev tier, no card: 30 RPM / 6K TPM / 14,400 req/day. The paid plans we list start at Per-token, usage-based (Developer / Pay-as-you-go (On-Demand)).
Plans and prices
| Plan | Price | What it covers |
|---|---|---|
| Free (GroqCloud) | $0 | Get started free with low rate limits (e.g. ~30 RPM, capped requests-per-day); shared per-organization limits. |
| Developer / Pay-as-you-go (On-Demand) | Per-token, usage-based | Linear per-token pricing with substantially higher rate limits; no idle infrastructure fees. |
| Batch API | 50% off standard per-token pricing | Asynchronous processing at half the on-demand token cost. |
| Enterprise | Custom (contact sales) | Private/co-cloud instances, SSO/SCIM/MFA, enterprise-only models (e.g. Minimax M2.5, Qwen3-VL 32B), higher capacity. |
How Groq bills you
The pricing model is per token. LPU-based inference host delivering 300-1000+ tokens/sec on open models, with a no-credit-card free dev tier. Best suited to: ultra-low-latency open-model inference.
What to watch for on the bill
- Persistent rate-limit and over-capacity (429) complaints; free tier is tight (~30 RPM)
- Flex/best-effort service tier can return over-capacity errors under load
- SRAM-only LPU architecture (few hundred MB per chip) raises questions about cost-efficiency at very large model sizes
The pricing page we read

Source: https://groq.com/pricing. Prices change, so confirm there before you commit.
Compare Groq with alternatives
Other llm apis: OpenAI API, Anthropic Claude API, Google Gemini API, Mistral La Plateforme, xAI Grok API.
Looking at quality rather than price? See the Groq score on APIbenchmarks. Full entry: Groq pricing.