Cohere vs Together AI Pricing (2026)
How do these two stack up on price? Here's what each one costs, what you get, and where the value sits.
| Cohere | Together AI | |
|---|---|---|
| Starts at | Custom | Custom |
| Number of plans | 3 | 7 |
| Free plan | — | — |
| Free trial | — | |
| Pricing model | custom | usage-based |
North
- Intuitive interface
- Purpose-built generative models
- Intelligent search
- AI agents for routine tasks and complex workflows
Compass
- Pre-built data connectors
- Intelligent search
- Document parsing
- Managed index
Model Vault
- Fully managed model deployment
- No shared resources or multi-tenancy overhead
- Seamless integration with Cohere North
- Simple startup and self-serve model access
- Fixed or Flex pricing plans available
Serverless Inference
- Chat models (price per 1M tokens)
- Vision models (price per 1M tokens)
- Image generation models (price per image or per mp)
- Audio / TTS models (price per 1M characters)
- Video generation models (price per video)
- Transcription models (price per audio minute)
- Embeddings models (price per 1M tokens)
- Moderation models (price per 1M tokens)
- Batch API pricing available
Provisioned Throughput
- Reserved dedicated capacity in throughput units (PTUs)
- Fixed capacity per PTU with model-dependent tokens-per-minute
- MiniMax M3 support
- GLM-5.2 support
Dedicated Inference
- Guaranteed performance (no sharing)
- Support for custom models
- Autoscaling & traffic spike handling
- NVIDIA HGX H100 on-demand ($5.49/GPU/hr)
- NVIDIA HGX B200 on-demand ($8.99/GPU/hr)
- NVIDIA HGX H200 (contact us)
- NVIDIA HGX B300 (contact us)
- NVIDIA GB200 NVL72 (contact us)
- NVIDIA GB300 NVL72 (contact us)
- Reserved capacity available
GPU Clusters
- On-demand pay-as-you-go GPU capacity (hourly)
- Reserved capacity (7-30 days, 31-90 days, 91-180 days, 181+ days)
- NVIDIA HGX H100 on-demand ($3.99/GPU/hr)
- NVIDIA HGX H200 on-demand ($5.99/GPU/hr)
- NVIDIA HGX B200 on-demand ($8.19/GPU/hr)
- NVIDIA GB200 NVL72 (contact us)
- NVIDIA GB300 NVL72 (contact us)
Sandbox
- Code Sandbox: VM sandboxes for large development environments ($0.0446/vCPU/hr, $0.0149/GiB RAM/hr)
- Code Interpreter: Execute LLM-generated code securely via API ($0.03/session)
Storage
- Shared Filesystem ($0.16/GiB/month)
- High-bandwidth, parallel filesystem colocated with compute
Fine-Tuning
- Supervised Fine-Tuning (LoRA)
- Supervised Fine-Tuning (Full Fine-Tuning)
- Direct Preference Optimization (LoRA)
- Direct Preference Optimization (Full Fine-Tuning)
- Standard pricing for models up to 100B parameters
- Specialized pricing for DeepSeek, GLM, Kimi, Llama 4, Qwen3, and other large models
- Minimum charge per job: $4.00 (standard)
Cohere vs Together AI FAQ
- Which one is cheaper?
- One or both use custom pricing, so it depends on your specific needs.
- Can I use either one for free?
- Neither has a free plan. But Cohere offers a free trial.
- How do they charge?
- Different approach here. Cohere uses custom pricing, while Together AI goes with usage-based. That changes the math depending on your team size and usage.
- Which one is a better deal?
- Depends on what you need. Cohere: Cohere is squarely targeting enterprise and developer teams that can't or won't send data to OpenAI or Anthropic — data residency, security, and deployment flexibility are the pitch. They're not the cheapest option, but they're positioning as the serious infrastructure play for regulated industries and large orgs that need control. Together AI: Together AI is pitching itself as the serious infrastructure layer for AI builders — not the cheapest (Replicate or self-hosting can undercut on specific models), but more flexible and enterprise-ready than most managed inference APIs. At $5.49/GPU/hr for H100 dedicated, they're competitive with CoreWeave and Lambda Labs while bundling more managed services around it.
Still deciding? See the best Cohere alternatives or the best Together AI alternatives, ranked with verified pricing.
Keep tabs on both.
We'll monitor pricing changes for Cohere and Together AI and let you know when something moves.