AssemblyAI vs Together AI Pricing (2026)
How do these two stack up on price? Here's what each one costs, what you get, and where the value sits.
| AssemblyAI | Together AI | |
|---|---|---|
| Starts at | Custom | Custom |
| Number of plans | 2 | 7 |
| Free plan | — | |
| Free trial | — | |
| Pricing model | usage-based | usage-based |
Pay as you go
- Universal-3.5 Pro ($0.21/hr)
- Universal-2 ($0.15/hr)
- Universal-3.5 Pro Realtime ($0.45/hr)
- Universal-Streaming ($0.15/hr)
- Universal-Streaming Multilingual ($0.15/hr)
- Sync API ($0.45/hr)
- Voice Agent API ($4.50/hr)
- Speaker Diarization ($0.02/hr pre-recorded, $0.12/hr realtime)
- Medical Mode ($0.15/hr)
- Keyterms Prompting
- Prompting ($0.05/hr)
- Speaker Identification ($0.02/hr)
- Translation ($0.06/hr)
- Custom Formatting ($0.03/hr)
- Entity Detection ($0.08/hr)
- Sentiment Analysis ($0.02/hr)
- Auto Chapters ($0.08/hr)
- Key Phrases ($0.01/hr)
- Topic Detection ($0.15/hr)
- Summarization ($0.03/hr)
- Profanity Filtering ($0.01/hr)
- PII Audio Redaction ($0.05/hr)
- PII Text Redaction ($0.08/hr)
- Content Moderation ($0.15/hr)
- LLM Gateway
- No minimum commitments
- No credit card required to start
Custom
- Custom rate limits
- Enhanced concurrency
- Enterprise-grade flexibility
- Volume-based pricing
- Custom starting concurrency limits
Serverless Inference
- Chat models (price per 1M tokens)
- Vision models (price per 1M tokens)
- Image generation models (price per image or per mp)
- Audio / TTS models (price per 1M characters)
- Video generation models (price per video)
- Transcription models (price per audio minute)
- Embeddings models (price per 1M tokens)
- Moderation models (price per 1M tokens)
- Batch API pricing available
Provisioned Throughput
- Reserved dedicated capacity in throughput units (PTUs)
- Fixed capacity per PTU with model-dependent tokens-per-minute
- MiniMax M3 support
- GLM-5.2 support
Dedicated Inference
- Guaranteed performance (no sharing)
- Support for custom models
- Autoscaling & traffic spike handling
- NVIDIA HGX H100 on-demand ($5.49/GPU/hr)
- NVIDIA HGX B200 on-demand ($8.99/GPU/hr)
- NVIDIA HGX H200 (contact us)
- NVIDIA HGX B300 (contact us)
- NVIDIA GB200 NVL72 (contact us)
- NVIDIA GB300 NVL72 (contact us)
- Reserved capacity available
GPU Clusters
- On-demand pay-as-you-go GPU capacity (hourly)
- Reserved capacity (7-30 days, 31-90 days, 91-180 days, 181+ days)
- NVIDIA HGX H100 on-demand ($3.99/GPU/hr)
- NVIDIA HGX H200 on-demand ($5.99/GPU/hr)
- NVIDIA HGX B200 on-demand ($8.19/GPU/hr)
- NVIDIA GB200 NVL72 (contact us)
- NVIDIA GB300 NVL72 (contact us)
Sandbox
- Code Sandbox: VM sandboxes for large development environments ($0.0446/vCPU/hr, $0.0149/GiB RAM/hr)
- Code Interpreter: Execute LLM-generated code securely via API ($0.03/session)
Storage
- Shared Filesystem ($0.16/GiB/month)
- High-bandwidth, parallel filesystem colocated with compute
Fine-Tuning
- Supervised Fine-Tuning (LoRA)
- Supervised Fine-Tuning (Full Fine-Tuning)
- Direct Preference Optimization (LoRA)
- Direct Preference Optimization (Full Fine-Tuning)
- Standard pricing for models up to 100B parameters
- Specialized pricing for DeepSeek, GLM, Kimi, Llama 4, Qwen3, and other large models
- Minimum charge per job: $4.00 (standard)
AssemblyAI vs Together AI FAQ
- Which one is cheaper?
- One or both use custom pricing, so it depends on your specific needs.
- Can I use either one for free?
- AssemblyAI has a free plan. Together AI doesn't — you'll need to pay from day one.
- How do they charge?
- Both use a usage-based model, so the comparison is straightforward — it comes down to features and limits at each price point.
- Which one is a better deal?
- Depends on what you need. AssemblyAI: They're positioning as the developer-friendly, API-first alternative to Deepgram and Rev AI — competitive on price at scale but differentiated by the breadth of AI features (LLM Gateway, multichannel, etc.). The AWS Marketplace listing signals they're actively chasing enterprise procurement budgets. Together AI: Together AI is pitching itself as the serious infrastructure layer for AI builders — not the cheapest (Replicate or self-hosting can undercut on specific models), but more flexible and enterprise-ready than most managed inference APIs. At $5.49/GPU/hr for H100 dedicated, they're competitive with CoreWeave and Lambda Labs while bundling more managed services around it.
Keep tabs on both.
We'll monitor pricing changes for AssemblyAI and Together AI and let you know when something moves.