macrostack

Is OpenAI Embeddings API free? What it actually costs in 2026

The short answer

OpenAI embeddings are not free: text-embedding-3-small is $0.02 per million tokens and text-embedding-3-large $0.13, billed on input tokens with no free allowance.

The model pages list a price per million tokens and nothing else — no free tier, no monthly credit. Embedding a corpus is a one-off cost (a million tokens is roughly 750,000 words, so a small documentation set costs cents); re-embedding on every query or every re-index is where the meter runs.

What the free tier actually includes

  • No free tier or credits
  • text-embedding-3-small: $0.02 per 1M tokens
  • text-embedding-3-large: $0.13 per 1M tokens

Where the wall is

Re-indexing. The first pass over a corpus is cheap; the bill comes from embedding user queries at scale and from re-embedding the whole corpus whenever the chunking or model changes — which is also the moment a locally-run open model starts to pay for itself.

After that: text-embedding-3-small $0.02 per 1M tokens; text-embedding-3-large $0.13 per 1M tokens; usage is prepaid or billed monthly like the rest of the API.

Can you self-host OpenAI Embeddings API? No — cloud only

OpenAI's embedding models run only on OpenAI's servers. Nomic Embed and BGE-M3 are open-weight embedding models that run locally; Voyage AI is the hosted alternative compared here.

What it actually takes to run one yourself

Real RAM against the documented minimum, monthly cost, setup time, and the one thing that catches people out.

When paying is still the right call

Under roughly 10–15 million embeddings a month, the API is genuinely cheaper and simpler than self-hosting — $0.02/1M is a fair price and text-embedding-3-small is good. If your whole stack is OpenAI and your data isn't sensitive, staying is the pragmatic call.

What you are locked into

The sneaky kind: embeddings are model-specific, so switching means re-embedding the entire corpus into a new vector space. The API isn't holding your data hostage — the vectors are. Versioning your embedding model from day one is the escape hatch.

If you would rather not pay: Nomic Embed

92

Nomic Embed

The sovereign default — Apache-2.0, one command in Ollama, up to 8k context with rotary scaling.

Free — runs locally via Ollama or sentence-transformers; your hardware is the cost

See all 5 OpenAI Embeddings API alternatives compared

Compare the free options head-to-head

Read from OpenAI's own pricing and documentation on 2026-09-27 — 3 days ago. Vendors change prices without notice; confirm before you commit.

The Macrostack brief

New swaps, worth your inbox.

A short, occasional email when we add a high-intent alternative or ship a new head-to-head. No spam, no selling your address — unsubscribe in one click.