Layer 4 · self-hosting reality check
What it actually takes to self-host BGE-M3
The docs say Not stated by the project. About 567M parameters: the pytorch_model.bin is 2.27 GB in fp32 and Ollama's build is 1.2 GB; it runs on CPU.. In practice you want Roughly 3–4 GB of RAM on CPU, or 2–3 GB of VRAM with use_fp16=True, for one request at a time (estimate; no figure is published). Encoding at the full 8,192-token max_length in large batches needs far more, which is why the card tells you to lower max_length if you can.. Here is the honest version — real requirements, real monthly cost, what you will be maintaining, and the one thing that catches people out.
Usually reached from OpenAI Embeddings API alternatives, where BGE-M3 is one of the picks.
Wondering whether you need to at all? Are OpenAI embeddings free? — what the free tier actually allows, and where the wall is.
| RAM — documented minimum | Not stated by the project. About 567M parameters: the pytorch_model.bin is 2.27 GB in fp32 and Ollama's build is 1.2 GB; it runs on CPU. |
|---|---|
| RAM — what it really needs | Roughly 3–4 GB of RAM on CPU, or 2–3 GB of VRAM with use_fp16=True, for one request at a time (estimate; no figure is published). Encoding at the full 8,192-token max_length in large batches needs far more, which is why the card tells you to lower max_length if you can. |
| CPU | CPU works for a modest corpus; for bulk indexing of long documents, any 8 GB+ NVIDIA card with use_fp16=True is the practical choice |
| Disk | 2.27 GB fp32 (the Hugging Face repo totals 4.59 GB with the ONNX copy), 1.2 GB via Ollama. Vector storage is 1,024 dimensions per chunk for dense output — and far more if you keep the multi-vector (ColBERT) output, which stores many vectors per passage. |
| Monthly cost | $0 — it runs on the machine you already own for a small corpus; a 24 GB RTX 3090/4090 on RunPod at $0.22–0.74/hr (read 2026-09-30) for a one-off bulk index |
| Setup time | 10 minutes with pip install -U FlagEmbedding, or ollama pull bge-m3 for dense-only |
| How you install it | pip install -U FlagEmbedding, then BGEM3FlagModel('BAAI/bge-m3', use_fp16=True) with return_dense / return_sparse / return_colbert_vecs; ollama pull bge-m3 serves dense vectors only |
| Ongoing maintenance | Stable weights and little to update; FlagEmbedding is still maintained (last pushed August 2026). Changing max_length or retrieval modes later means re-embedding the corpus. |
| Where it stops scaling | Stateless, so it scales by adding replicas. One GPU bulk-indexes a large corpus; the bottleneck then moves to the vector store, especially if you keep multi-vector output. |
The thing that catches people out
Most people pick BGE-M3 for its three retrieval modes, then serve it in a way that returns only one. Ollama's embed API returns plain dense vectors — its response has no field for sparse or multi-vector output — and OpenAI-style embedding endpoints are the same. The sparse and ColBERT outputs come from FlagEmbedding's BGEM3FlagModel, and you need a store that can hold them (the project points to Milvus and Vespa). Decide on hybrid search before you pick the serving path.
When not to self-host BGE-M3
Your content is English, your chunks are short, and you only need a dense vector on a small machine — nomic-embed-text is about a quarter of the size in Ollama and lighter on CPU. Also skip it if your vector database cannot store sparse vectors and hybrid search was the reason you chose it.
Every guide here carries this section. A site that only ever tells you to self-host is selling something — the useful answer is sometimes no.
Other Layer 4 self-hosting guides
- Self-hosting Ollama8 GB VRAM for a 7B model at usable speed; 24 GB for 30B-class
- Self-hosting vLLM24 GB VRAM minimum for useful production serving
- Self-hosting QdrantVectors × dimensions × 4 bytes, in RAM, plus overhead — 1M × 768d is roughly 3 GB
- Self-hosting pgvector8 GB — the HNSW index wants to be resident
- Self-hosting LlamaIndex4 GB for the app; your vector store is the real cost
- Self-hosting faster-whisper5 GB VRAM for large-v3 in float16; 2 GB with int8
Common questions
- How much RAM does BGE-M3 actually need?
- Roughly 3–4 GB of RAM on CPU, or 2–3 GB of VRAM with use_fp16=True, for one request at a time (estimate; no figure is published). Encoding at the full 8,192-token max_length in large batches needs far more, which is why the card tells you to lower max_length if you can. in practice. The documented minimum is Not stated by the project. About 567M parameters: the pytorch_model.bin is 2.27 GB in fp32 and Ollama's build is 1.2 GB; it runs on CPU., which is the figure at which the process starts rather than the figure at which it works under real use. CPU works for a modest corpus; for bulk indexing of long documents, any 8 GB+ NVIDIA card with use_fp16=True is the practical choice alongside it.
- What does self-hosting BGE-M3 cost per month?
- $0 — it runs on the machine you already own for a small corpus; a 24 GB RTX 3090/4090 on RunPod at $0.22–0.74/hr (read 2026-09-30) for a one-off bulk index This is commodity VPS pricing and excludes your time, which is the larger cost for most people — budget for stable weights and little to update; FlagEmbedding is still maintained (last pushed August 2026). Changing max_length or retrieval modes later means re-embedding the corpus.
- How long does it take to set up BGE-M3?
- 10 minutes with pip install -U FlagEmbedding, or ollama pull bge-m3 for dense-only, via pip install -U FlagEmbedding, then BGEM3FlagModel('BAAI/bge-m3', use_fp16=True) with return_dense / return_sparse / return_colbert_vecs; ollama pull bge-m3 serves dense vectors only.
- When should I NOT self-host BGE-M3?
- Your content is English, your chunks are short, and you only need a dense vector on a small machine — nomic-embed-text is about a quarter of the size in Ollama and lighter on CPU. Also skip it if your vector database cannot store sparse vectors and hybrid search was the reason you chose it.
- What is the most common mistake when self-hosting BGE-M3?
- Most people pick BGE-M3 for its three retrieval modes, then serve it in a way that returns only one. Ollama's embed API returns plain dense vectors — its response has no field for sparse or multi-vector output — and OpenAI-style embedding endpoints are the same. The sparse and ColBERT outputs come from FlagEmbedding's BGEM3FlagModel, and you need a store that can hold them (the project points to Milvus and Vespa). Decide on hybrid search before you pick the serving path.