Together AI Pricing (2026)

Pure usage-based across every surface — you're paying per token, per GPU-hour, per session, per GiB, depending on what you're touching. The value metric shifts by product: inference is token/output-based, compute is time-based, storage is volume-based, which means your bill is a direct function of how hard you're running the platform.

7 plans usage-based
Serverless Inference
Contact Sales
Provisioned Throughput
Contact Sales
Dedicated Inference
Contact Sales
GPU Clusters
Contact Sales
Sandbox
Contact Sales
Storage
Contact Sales
Fine-Tuning
Contact Sales
Verified Jul 26, 2026 Official pricing page

Keep up with your competitors, without the manual work.

Outmano tracks pricing, features, roadmaps and reviews across your market, then sends one weekly brief: what changed, and what it means for you.

Try Outmano free »

Serverless Inference

Contact Sales
  • Chat models (price per 1M tokens)
  • Vision models (price per 1M tokens)
  • Image generation models (price per image or per mp)
  • Audio / TTS models (price per 1M characters)
  • Video generation models (price per video)
  • Transcription models (price per audio minute)
  • Embeddings models (price per 1M tokens)
  • Moderation models (price per 1M tokens)
  • Batch API pricing available

Provisioned Throughput

Contact Sales
  • Reserved dedicated capacity in throughput units (PTUs)
  • Fixed capacity per PTU with model-dependent tokens-per-minute
  • MiniMax M3 support
  • GLM-5.2 support

Dedicated Inference

Contact Sales
  • Guaranteed performance (no sharing)
  • Support for custom models
  • Autoscaling & traffic spike handling
  • NVIDIA HGX H100 on-demand ($5.49/GPU/hr)
  • NVIDIA HGX B200 on-demand ($8.99/GPU/hr)
  • NVIDIA HGX H200 (contact us)
  • NVIDIA HGX B300 (contact us)
  • NVIDIA GB200 NVL72 (contact us)
  • NVIDIA GB300 NVL72 (contact us)
  • Reserved capacity available

GPU Clusters

Contact Sales
  • On-demand pay-as-you-go GPU capacity (hourly)
  • Reserved capacity (7-30 days, 31-90 days, 91-180 days, 181+ days)
  • NVIDIA HGX H100 on-demand ($3.99/GPU/hr)
  • NVIDIA HGX H200 on-demand ($5.99/GPU/hr)
  • NVIDIA HGX B200 on-demand ($8.19/GPU/hr)
  • NVIDIA GB200 NVL72 (contact us)
  • NVIDIA GB300 NVL72 (contact us)

Sandbox

Contact Sales
  • Code Sandbox: VM sandboxes for large development environments ($0.0446/vCPU/hr, $0.0149/GiB RAM/hr)
  • Code Interpreter: Execute LLM-generated code securely via API ($0.03/session)

Storage

Contact Sales
  • Shared Filesystem ($0.16/GiB/month)
  • High-bandwidth, parallel filesystem colocated with compute

Fine-Tuning

Contact Sales
  • Supervised Fine-Tuning (LoRA)
  • Supervised Fine-Tuning (Full Fine-Tuning)
  • Direct Preference Optimization (LoRA)
  • Direct Preference Optimization (Full Fine-Tuning)
  • Standard pricing for models up to 100B parameters
  • Specialized pricing for DeepSeek, GLM, Kimi, Llama 4, Qwen3, and other large models
  • Minimum charge per job: $4.00 (standard)

AI Pricing Analysis

Pricing Model

Pure usage-based across every surface — you're paying per token, per GPU-hour, per session, per GiB, depending on what you're touching. The value metric shifts by product: inference is token/output-based, compute is time-based, storage is volume-based, which means your bill is a direct function of how hard you're running the platform.

Tier Strategy

There are no tiers in the traditional sense — it's more of a maturity ladder. Serverless is where you start (low commitment, pay-as-you-go), Provisioned Throughput is the step up when latency and rate limits start hurting, and Dedicated Inference or GPU Clusters are for teams that need guaranteed capacity or want to run custom workloads at scale. The upgrade trigger is almost always hitting a performance or cost-efficiency ceiling on serverless.

Competitive Positioning

Together AI is pitching itself as the serious infrastructure layer for AI builders — not the cheapest (Replicate or self-hosting can undercut on specific models), but more flexible and enterprise-ready than most managed inference APIs. At $5.49/GPU/hr for H100 dedicated, they're competitive with CoreWeave and Lambda Labs while bundling more managed services around it.

Growth Lever

Expansion is entirely consumption-driven — no seats, no feature gates, just more usage. The real pull toward higher spend is the shift from serverless to provisioned or dedicated compute, where teams lock in capacity and the per-unit economics improve but the baseline commitment goes up significantly.

Together AI Pricing FAQ

How much does Together AI cost?
Pricing is custom — you'll need to talk to their sales team for a quote.
Is there a free plan?
No free plan. You'll need to commit to a paid plan to get started.
Can I try it before paying?
Not at the moment. There's no free trial or free plan listed on their pricing page.
How does the pricing work?
Pure usage-based across every surface — you're paying per token, per GPU-hour, per session, per GiB, depending on what you're touching. The value metric shifts by product: inference is token/output-based, compute is time-based, storage is volume-based, which means your bill is a direct function of how hard you're running the platform.
Which plan makes sense for me?
There are no tiers in the traditional sense — it's more of a maturity ladder. Serverless is where you start (low commitment, pay-as-you-go), Provisioned Throughput is the step up when latency and rate limits start hurting, and Dedicated Inference or GPU Clusters are for teams that need guaranteed capacity or want to run custom workloads at scale. The upgrade trigger is almost always hitting a performance or cost-efficiency ceiling on serverless.

Set it up once. Stay ahead all year.

Add the competitors you care about and Outmano does the watching — then hands you a weekly action plan with what to do next.

Start tracking free »