Layer 2.5 · head to head
Together AI vs Fireworks AI
Per-token serving where the benchmark that matters is your prompt shape, not theirs.
| Together AI | Fireworks AI | |
|---|---|---|
| Can you leave | Sticky. Leaving means rewriting against a different interface, not changing a hostname. | Sticky. Leaving means rewriting against a different interface, not changing a hostname. |
| Counterparty | Low for serverless — you are buying tokens, not capacity. High if you build a product on one provider's fine-tune tooling and its model catalogue. | Low for serverless use. The lock-in is the tuned-model artefact, not the compute. |
| What it is | You never see a GPU. You send a request and pay per token or per second of execution. | You never see a GPU. You send a request and pay per token or per second of execution. |
| How you reach it | You call an endpoint. There is no infrastructure to see, and no infrastructure to move. | You call an endpoint. There is no infrastructure to see, and no infrastructure to move. |
| Accelerators | H100, H200, B200 behind the API; dedicated clusters available | Not disclosed per endpoint; NVIDIA datacentre class |
| Regions | US | US |
| Pricing model | Per token for serverless inference; per GPU-hour for dedicated endpoints and training clusters. | Per token, with dedicated deployments available. |
| Getting started | API key. | API key. |
| Capacity | Good on popular open-weight models. | Good on popular open-weight models. |
| Ownership | Private. | Private. |
Together AI
For: Teams serving open-weight models who do not want to operate a GPU at all.
The catch: Per-token pricing hides utilisation. It is cheaper than a dedicated GPU right up until it is dramatically more expensive, and the crossover is a real calculation nobody does before signing up.
Economics: The correct comparison is not against other token prices — it is against your own GPU-hour cost at your actual utilisation. Below roughly 40% duty cycle serverless usually wins; above it, rarely.
Fireworks AI
For: Latency-sensitive open-weight inference where throughput per dollar matters more than owning the stack.
The catch: Serving optimisations are the product and they are proprietary. Benchmarks against a self-hosted baseline are not portable to your own hardware.
Economics: Compete on tokens per second per dollar rather than on raw GPU price. Measure with your prompt shape, not theirs.
Neither table row is a price
Deliberately. Published on-demand rates at this layer move weekly, and essentially nobody signing a real contract pays them — every serious buyer pays less than every list figure either of these companies publishes. Quoting one here would date this page within a month.
The GPU rental price index carries dated, sourced figures instead, and the durable finding there is the spread: the identical H100 rents from roughly $1.38 to $12.29 an hour depending only on who you rent it from.
The layers underneath both
Whichever you pick is renting you chips in a building that needs power. In 2026 that is the constraint that binds: Microsoft has disclosed an Azure backlog it cannot fill for want of megawatts rather than accelerators, and the US interconnection queue exceeds 2,600 GW with roughly 80% of projects withdrawing before they energise.
Layer 2 — Silicon · Layer 1 — Energy · The interconnection queue