macrostack

Layer 4 · self-hosting reality check

What it actually takes to self-host BGE-M3

The docs say Not stated by the project. About 567M parameters: the pytorch_model.bin is 2.27 GB in fp32 and Ollama's build is 1.2 GB; it runs on CPU.. In practice you want Roughly 3–4 GB of RAM on CPU, or 2–3 GB of VRAM with use_fp16=True, for one request at a time (estimate; no figure is published). Encoding at the full 8,192-token max_length in large batches needs far more, which is why the card tells you to lower max_length if you can.. Here is the honest version — real requirements, real monthly cost, what you will be maintaining, and the one thing that catches people out.

Usually reached from OpenAI Embeddings API alternatives, where BGE-M3 is one of the picks.

Wondering whether you need to at all? Are OpenAI embeddings free? — what the free tier actually allows, and where the wall is.

RAM — documented minimumNot stated by the project. About 567M parameters: the pytorch_model.bin is 2.27 GB in fp32 and Ollama's build is 1.2 GB; it runs on CPU.
RAM — what it really needsRoughly 3–4 GB of RAM on CPU, or 2–3 GB of VRAM with use_fp16=True, for one request at a time (estimate; no figure is published). Encoding at the full 8,192-token max_length in large batches needs far more, which is why the card tells you to lower max_length if you can.
CPUCPU works for a modest corpus; for bulk indexing of long documents, any 8 GB+ NVIDIA card with use_fp16=True is the practical choice
Disk2.27 GB fp32 (the Hugging Face repo totals 4.59 GB with the ONNX copy), 1.2 GB via Ollama. Vector storage is 1,024 dimensions per chunk for dense output — and far more if you keep the multi-vector (ColBERT) output, which stores many vectors per passage.
Monthly cost$0 — it runs on the machine you already own for a small corpus; a 24 GB RTX 3090/4090 on RunPod at $0.22–0.74/hr (read 2026-09-30) for a one-off bulk index
Setup time10 minutes with pip install -U FlagEmbedding, or ollama pull bge-m3 for dense-only
How you install itpip install -U FlagEmbedding, then BGEM3FlagModel('BAAI/bge-m3', use_fp16=True) with return_dense / return_sparse / return_colbert_vecs; ollama pull bge-m3 serves dense vectors only
Ongoing maintenanceStable weights and little to update; FlagEmbedding is still maintained (last pushed August 2026). Changing max_length or retrieval modes later means re-embedding the corpus.
Where it stops scalingStateless, so it scales by adding replicas. One GPU bulk-indexes a large corpus; the bottleneck then moves to the vector store, especially if you keep multi-vector output.

The thing that catches people out

Most people pick BGE-M3 for its three retrieval modes, then serve it in a way that returns only one. Ollama's embed API returns plain dense vectors — its response has no field for sparse or multi-vector output — and OpenAI-style embedding endpoints are the same. The sparse and ColBERT outputs come from FlagEmbedding's BGEM3FlagModel, and you need a store that can hold them (the project points to Milvus and Vespa). Decide on hybrid search before you pick the serving path.

When not to self-host BGE-M3

Your content is English, your chunks are short, and you only need a dense vector on a small machine — nomic-embed-text is about a quarter of the size in Ollama and lighter on CPU. Also skip it if your vector database cannot store sparse vectors and hybrid search was the reason you chose it.

Every guide here carries this section. A site that only ever tells you to self-host is selling something — the useful answer is sometimes no.

Other Layer 4 self-hosting guides

Common questions

How much RAM does BGE-M3 actually need?
Roughly 3–4 GB of RAM on CPU, or 2–3 GB of VRAM with use_fp16=True, for one request at a time (estimate; no figure is published). Encoding at the full 8,192-token max_length in large batches needs far more, which is why the card tells you to lower max_length if you can. in practice. The documented minimum is Not stated by the project. About 567M parameters: the pytorch_model.bin is 2.27 GB in fp32 and Ollama's build is 1.2 GB; it runs on CPU., which is the figure at which the process starts rather than the figure at which it works under real use. CPU works for a modest corpus; for bulk indexing of long documents, any 8 GB+ NVIDIA card with use_fp16=True is the practical choice alongside it.
What does self-hosting BGE-M3 cost per month?
$0 — it runs on the machine you already own for a small corpus; a 24 GB RTX 3090/4090 on RunPod at $0.22–0.74/hr (read 2026-09-30) for a one-off bulk index This is commodity VPS pricing and excludes your time, which is the larger cost for most people — budget for stable weights and little to update; FlagEmbedding is still maintained (last pushed August 2026). Changing max_length or retrieval modes later means re-embedding the corpus.
How long does it take to set up BGE-M3?
10 minutes with pip install -U FlagEmbedding, or ollama pull bge-m3 for dense-only, via pip install -U FlagEmbedding, then BGEM3FlagModel('BAAI/bge-m3', use_fp16=True) with return_dense / return_sparse / return_colbert_vecs; ollama pull bge-m3 serves dense vectors only.
When should I NOT self-host BGE-M3?
Your content is English, your chunks are short, and you only need a dense vector on a small machine — nomic-embed-text is about a quarter of the size in Ollama and lighter on CPU. Also skip it if your vector database cannot store sparse vectors and hybrid search was the reason you chose it.
What is the most common mistake when self-hosting BGE-M3?
Most people pick BGE-M3 for its three retrieval modes, then serve it in a way that returns only one. Ollama's embed API returns plain dense vectors — its response has no field for sparse or multi-vector output — and OpenAI-style embedding endpoints are the same. The sparse and ColBERT outputs come from FlagEmbedding's BGEM3FlagModel, and you need a store that can hold them (the project points to Milvus and Vespa). Decide on hybrid search before you pick the serving path.
The Macrostack brief

New swaps, worth your inbox.

A short, occasional email when we add a high-intent alternative or ship a new head-to-head. No spam, no selling your address — unsubscribe in one click.