AssemblyAI vs OpenAI Whisper / GPT-4o Transcribe pricing
Both are in speech-to-text apis. AssemblyAI is billed per minute (usage-based); OpenAI Whisper / GPT-4o Transcribe is billed per minute (per-second billed). Numbers below are copied from each vendor's pricing page.
AssemblyAIFree tier or trial
- Free tier
- $50 one-time credit
- Billed
- Per minute (usage-based)
- Best for
- Accurate STT plus audio intelligence
- Verified
- Jun 27, 2026
OpenAI Whisper / GPT-4o TranscribeNo free tier
- Free tier
- None (pay-per-use)
- Billed
- Per minute (per-second billed)
- Best for
- Transcription inside the OpenAI model API
- Verified
- Jun 27, 2026
AssemblyAI plans
| Plan | Price | What it covers |
|---|---|---|
| Free credits | $50 free | No credit card required; 5 new streams/min concurrency limit. Self-serve signup. |
| Universal-2 (async STT) | $0.15/hr | Pre-recorded transcription, billed per second; lowest base rate tier. |
| Universal-3 Pro (async STT) | $0.21/hr | Higher-accuracy async model (~5.6% mean WER, keyterm prompting). |
| Universal-Streaming | $0.15/hr | Real-time English/multilingual; billed per WebSocket session duration. |
| Voice Agent API | $4.50/hr ($0.075/min) | All-inclusive STT + LLM + TTS + turn detection + tool calling. |
| Enterprise | Custom | 99.9% uptime SLA, custom SLOs, volume pricing, committed-use; 600M+ monthly inference calls capacity. |
- Add-on features (diarization, entity detection, redaction, translation) stack on the base rate, pushing effective cost well above the $0.15/hr headline (~$0.35/hr in independent estimates)
- Cost becomes expensive at high audio volume for usage-heavy apps
- Billing-UX friction reported (e.g. difficulty removing a card or disabling autopay)
OpenAI Whisper / GPT-4o Transcribe plans
| Plan | Price | What it covers |
|---|---|---|
| whisper-1 | $0.006 / min | Original hosted Whisper model; billed per second of audio duration, ~$0.36/hour. |
| gpt-4o-transcribe | $0.006 / min (~$6.00 / 1M audio input tokens) | Flagship GPT-4o speech-to-text; lower WER and better multilingual accuracy than Whisper. |
| gpt-4o-mini-transcribe | $0.003 / min (~$3.00 / 1M audio input tokens) | Cost-efficient smaller model, half the price of the flagship. |
| gpt-4o-transcribe-diarize | Usage-based (audio token pricing) | Adds built-in speaker diarization; Transcription API only, returns diarized_json with A:/B: speaker labels. |
- Token-based GPT-4o billing can produce less predictable costs on long audio than flat per-minute pricing