serverless inference · api · Layer 2.5
Together AI
Teams serving open-weight models who do not want to operate a GPU at all.
low portability Sticky. Leaving means rewriting against a different interface, not changing a hostname.
What you are buying
- Business model
- You never see a GPU. You send a request and pay per token or per second of execution.
- How you reach it
- You call an endpoint. There is no infrastructure to see, and no infrastructure to move.
- Accelerators
- H100, H200, B200 behind the API; dedicated clusters available
- Regions
- US
- Pricing model
- Per token for serverless inference; per GPU-hour for dedicated endpoints and training clusters.
- Getting started
- API key.
- Capacity
- Good on popular open-weight models.
- Ownership
- Private.
Verified 2026-09-09. Fields reading “not published” are exactly that — we do not estimate a figure a vendor withholds.
The counterparty
Low for serverless — you are buying tokens, not capacity. High if you build a product on one provider's fine-tune tooling and its model catalogue.
A multi-year GPU commitment is a credit decision wearing a cloud contract. This is the section no benchmark covers and the one that decides what happens to your workload in 2028.
The catch
Per-token pricing hides utilisation. It is cheaper than a dedicated GPU right up until it is dramatically more expensive, and the crossover is a real calculation nobody does before signing up.
The economics
The correct comparison is not against other token prices — it is against your own GPU-hour cost at your actual utilisation. Below roughly 40% duty cycle serverless usually wins; above it, rarely.
No rate is quoted on this page on purpose. Published list prices at this layer move weekly and essentially nobody signing a real contract pays them. Dated, sourced figures live in the GPU rental price index, where the spread between the cheapest and dearest seller of the identical chip runs to roughly 9x.
The layers underneath this one
Whatever you rent here is a chip in a building that needs power. In 2026 megawatts, not silicon, are the binding constraint on the whole industry — a frontier rack draws 120–200 kW against a 2026 average near 27 kW, and the US interconnection queue exceeds 2,600 GW.
Layer 2 — Silicon · Layer 1 — Energy · The interconnection queue · Tokens per watt
Compared against
- Together AI vs Fireworks AI — Per-token serving where the benchmark that matters is your prompt shape, not theirs.Compare with Fireworks AI
- Together AI vs RunPod — Tokens against GPU-hours: the crossover calculation almost nobody does.Compare with RunPod
Related providers
- Baseten — serverless inference, low portability
- Fireworks AI — serverless inference, low portability
- Replicate — serverless inference, low portability
- Modal — serverless inference, low portability
- Akash Network — marketplace, medium portability
- AWS (P5, P6, G6) — hyperscaler, medium portability