Together AI Pricing (2026)
Pure usage-based across every surface — you're paying per token, per GPU-hour, per session, per GiB, depending on what you're touching. The value metric shifts by product: inference is token/output-based, compute is time-based, storage is volume-based, which means your bill is a direct function of how hard you're running the platform.
- Serverless Inference
- Contact Sales
- Provisioned Throughput
- Contact Sales
- Dedicated Inference
- Contact Sales
- GPU Clusters
- Contact Sales
- Sandbox
- Contact Sales
- Storage
- Contact Sales
- Fine-Tuning
- Contact Sales
Keep up with your competitors, without the manual work.
Outmano tracks pricing, features, roadmaps and reviews across your market, then sends one weekly brief: what changed, and what it means for you.
Serverless Inference
- Chat models (price per 1M tokens)
- Vision models (price per 1M tokens)
- Image generation models (price per image or per mp)
- Audio / TTS models (price per 1M characters)
- Video generation models (price per video)
- Transcription models (price per audio minute)
- Embeddings models (price per 1M tokens)
- Moderation models (price per 1M tokens)
- Batch API pricing available
Provisioned Throughput
- Reserved dedicated capacity in throughput units (PTUs)
- Fixed capacity per PTU with model-dependent tokens-per-minute
- MiniMax M3 support
- GLM-5.2 support
Dedicated Inference
- Guaranteed performance (no sharing)
- Support for custom models
- Autoscaling & traffic spike handling
- NVIDIA HGX H100 on-demand ($5.49/GPU/hr)
- NVIDIA HGX B200 on-demand ($8.99/GPU/hr)
- NVIDIA HGX H200 (contact us)
- NVIDIA HGX B300 (contact us)
- NVIDIA GB200 NVL72 (contact us)
- NVIDIA GB300 NVL72 (contact us)
- Reserved capacity available
GPU Clusters
- On-demand pay-as-you-go GPU capacity (hourly)
- Reserved capacity (7-30 days, 31-90 days, 91-180 days, 181+ days)
- NVIDIA HGX H100 on-demand ($3.99/GPU/hr)
- NVIDIA HGX H200 on-demand ($5.99/GPU/hr)
- NVIDIA HGX B200 on-demand ($8.19/GPU/hr)
- NVIDIA GB200 NVL72 (contact us)
- NVIDIA GB300 NVL72 (contact us)
Sandbox
- Code Sandbox: VM sandboxes for large development environments ($0.0446/vCPU/hr, $0.0149/GiB RAM/hr)
- Code Interpreter: Execute LLM-generated code securely via API ($0.03/session)
Storage
- Shared Filesystem ($0.16/GiB/month)
- High-bandwidth, parallel filesystem colocated with compute
Fine-Tuning
- Supervised Fine-Tuning (LoRA)
- Supervised Fine-Tuning (Full Fine-Tuning)
- Direct Preference Optimization (LoRA)
- Direct Preference Optimization (Full Fine-Tuning)
- Standard pricing for models up to 100B parameters
- Specialized pricing for DeepSeek, GLM, Kimi, Llama 4, Qwen3, and other large models
- Minimum charge per job: $4.00 (standard)
AI Pricing Analysis
Pricing Model
Pure usage-based across every surface — you're paying per token, per GPU-hour, per session, per GiB, depending on what you're touching. The value metric shifts by product: inference is token/output-based, compute is time-based, storage is volume-based, which means your bill is a direct function of how hard you're running the platform.
Tier Strategy
There are no tiers in the traditional sense — it's more of a maturity ladder. Serverless is where you start (low commitment, pay-as-you-go), Provisioned Throughput is the step up when latency and rate limits start hurting, and Dedicated Inference or GPU Clusters are for teams that need guaranteed capacity or want to run custom workloads at scale. The upgrade trigger is almost always hitting a performance or cost-efficiency ceiling on serverless.
Competitive Positioning
Together AI is pitching itself as the serious infrastructure layer for AI builders — not the cheapest (Replicate or self-hosting can undercut on specific models), but more flexible and enterprise-ready than most managed inference APIs. At $5.49/GPU/hr for H100 dedicated, they're competitive with CoreWeave and Lambda Labs while bundling more managed services around it.
Growth Lever
Expansion is entirely consumption-driven — no seats, no feature gates, just more usage. The real pull toward higher spend is the shift from serverless to provisioned or dedicated compute, where teams lock in capacity and the per-unit economics improve but the baseline commitment goes up significantly.
Together AI Pricing FAQ
- How much does Together AI cost?
- Pricing is custom — you'll need to talk to their sales team for a quote.
- Is there a free plan?
- No free plan. You'll need to commit to a paid plan to get started.
- Can I try it before paying?
- Not at the moment. There's no free trial or free plan listed on their pricing page.
- How does the pricing work?
- Pure usage-based across every surface — you're paying per token, per GPU-hour, per session, per GiB, depending on what you're touching. The value metric shifts by product: inference is token/output-based, compute is time-based, storage is volume-based, which means your bill is a direct function of how hard you're running the platform.
- Which plan makes sense for me?
- There are no tiers in the traditional sense — it's more of a maturity ladder. Serverless is where you start (low commitment, pay-as-you-go), Provisioned Throughput is the step up when latency and rate limits start hurting, and Dedicated Inference or GPU Clusters are for teams that need guaranteed capacity or want to run custom workloads at scale. The upgrade trigger is almost always hitting a performance or cost-efficiency ceiling on serverless.
More AI pricing
Browse all tools →-
Anthropic Pricing
AI safety company behind Claude.
From $20/mo -
AssemblyAI Pricing
Speech-to-text and audio intelligence API.
Free plan Free trial -
Avoma Pricing
AI meeting assistant and revenue intelligence.
From $25/mo Free trial -
Baseten Pricing
ML model deployment infrastructure.
Free plan Free trial -
Braintrust Pricing
Evals and observability for LLM apps.
From $249/mo -
Cohere Pricing
Enterprise AI models and RAG.
See pricing Free trial -
Comet Pricing
ML experiment tracking and evals.
From $19/mo Free trial -
Copy.ai Pricing
AI-powered copywriting.
From $29/mo
Comparing Together AI to something specific? Try Together AI vs Anthropic, Together AI vs AssemblyAI, or Together AI vs Avoma.
Set it up once. Stay ahead all year.
Add the competitors you care about and Outmano does the watching — then hands you a weekly action plan with what to do next.