Layer 4 · self-hosting reality check
What it actually takes to self-host Axolotl
The docs say 11 GB of VRAM — Axolotl's Llama-3 examples list LoRA on an 8B model (8-bit base) as "Single GPU @ 11GB VRAM"; their Qwen3 14B QLoRA Colab runs on a 16 GB T4, at 9.26 GB after loading. In practice you want 24 GB of VRAM for comfortable LoRA or QLoRA on 7–14B models. Full fine-tuning an 8B is listed at 48 GB on one GPU, and 70B QLoRA+FSDP as "Dual GPU @ 21GB VRAM". Plenty of system RAM too — the FAQ says exit code -9 means you ran out of system RAM, not VRAM.. Here is the honest version — real requirements, real monthly cost, what you will be maintaining, and the one thing that catches people out.
Usually reached from OpenAI Fine-Tuning alternatives, where Axolotl is one of the picks.
Wondering whether you need to at all? Is OpenAI fine-tuning free? — what the free tier actually allows, and where the wall is.
| RAM — documented minimum | 11 GB of VRAM — Axolotl's Llama-3 examples list LoRA on an 8B model (8-bit base) as "Single GPU @ 11GB VRAM"; their Qwen3 14B QLoRA Colab runs on a 16 GB T4, at 9.26 GB after loading |
|---|---|
| RAM — what it really needs | 24 GB of VRAM for comfortable LoRA or QLoRA on 7–14B models. Full fine-tuning an 8B is listed at 48 GB on one GPU, and 70B QLoRA+FSDP as "Dual GPU @ 21GB VRAM". Plenty of system RAM too — the FAQ says exit code -9 means you ran out of system RAM, not VRAM. |
| CPU | NVIDIA Ampere or newer (RTX 30/40-series, A100, L4, H100) for bf16 and Flash Attention, or AMD. Blackwell cards need CUDA 13.0. Linux only — Windows via WSL2 or Docker. |
| Disk | The base model in bf16 (about 16 GB for an 8B), the tokenised dataset cache, a checkpoint per save, and the merged model at the end — budget three to four times the base model's size |
| Monthly cost | Paid per run, not per month: a 24 GB RTX 4090 on RunPod is $0.34–0.74/hr and an 80 GB A100 $1.19–1.59/hr (read 2026-09-30). A small LoRA run is measured in hours, so the bill is usually tens of dollars; $0 on a 24 GB card you already own. |
| Setup time | An hour with the Docker image; an afternoon if you are matching CUDA, PyTorch and flash-attention versions by hand |
| How you install it | uv pip install --no-build-isolation axolotl[deepspeed] on Python 3.12 with PyTorch 2.13+, or the axolotlai/axolotl-uv Docker image; then axolotl fetch examples and axolotl train <config>.yml |
| Ongoing maintenance | Very active — last pushed 2026-09-30. Minimum Python and PyTorch versions move with it (now Python 3.12, PyTorch 2.13), so pin the Docker tag per project and re-test old YAML configs after upgrading. |
| Where it stops scaling | From a single 16 GB card (QLoRA) to multi-GPU and multi-node training via FSDP2 and DeepSpeed, with sequence parallelism for long contexts — the same YAML file scales up. |
The thing that catches people out
A run can finish cleanly and still produce a model that does not know when to stop. Axolotl's chat-template docs warn that your tokenizer's eos_token must match the EOS token in the chat template — otherwise set eos_token under special_tokens — and a mismatch means the model never learns where its turn ends, so it runs on past its answer. Before a long run, inspect tokenised examples with axolotl preprocess config.yml --debug.
When not to self-host Axolotl
You have one small job and no GPU: OpenAI's fine-tuning API or a hosted trainer is far less work. If you will only ever train on one consumer GPU, Unsloth's single-GPU notebooks are the gentler start.
Every guide here carries this section. A site that only ever tells you to self-host is selling something — the useful answer is sometimes no.
Other Layer 4 self-hosting guides
- Self-hosting Ollama8 GB VRAM for a 7B model at usable speed; 24 GB for 30B-class
- Self-hosting vLLM24 GB VRAM minimum for useful production serving
- Self-hosting QdrantVectors × dimensions × 4 bytes, in RAM, plus overhead — 1M × 768d is roughly 3 GB
- Self-hosting pgvector8 GB — the HNSW index wants to be resident
- Self-hosting LlamaIndex4 GB for the app; your vector store is the real cost
- Self-hosting faster-whisper5 GB VRAM for large-v3 in float16; 2 GB with int8
Common questions
- How much RAM does Axolotl actually need?
- 24 GB of VRAM for comfortable LoRA or QLoRA on 7–14B models. Full fine-tuning an 8B is listed at 48 GB on one GPU, and 70B QLoRA+FSDP as "Dual GPU @ 21GB VRAM". Plenty of system RAM too — the FAQ says exit code -9 means you ran out of system RAM, not VRAM. in practice. The documented minimum is 11 GB of VRAM — Axolotl's Llama-3 examples list LoRA on an 8B model (8-bit base) as "Single GPU @ 11GB VRAM"; their Qwen3 14B QLoRA Colab runs on a 16 GB T4, at 9.26 GB after loading, which is the figure at which the process starts rather than the figure at which it works under real use. NVIDIA Ampere or newer (RTX 30/40-series, A100, L4, H100) for bf16 and Flash Attention, or AMD. Blackwell cards need CUDA 13.0. Linux only — Windows via WSL2 or Docker. alongside it.
- What does self-hosting Axolotl cost per month?
- Paid per run, not per month: a 24 GB RTX 4090 on RunPod is $0.34–0.74/hr and an 80 GB A100 $1.19–1.59/hr (read 2026-09-30). A small LoRA run is measured in hours, so the bill is usually tens of dollars; $0 on a 24 GB card you already own. This is commodity VPS pricing and excludes your time, which is the larger cost for most people — budget for very active — last pushed 2026-09-30. Minimum Python and PyTorch versions move with it (now Python 3.12, PyTorch 2.13), so pin the Docker tag per project and re-test old YAML configs after upgrading.
- How long does it take to set up Axolotl?
- An hour with the Docker image; an afternoon if you are matching CUDA, PyTorch and flash-attention versions by hand, via uv pip install --no-build-isolation axolotl[deepspeed] on Python 3.12 with PyTorch 2.13+, or the axolotlai/axolotl-uv Docker image; then axolotl fetch examples and axolotl train <config>.yml.
- When should I NOT self-host Axolotl?
- You have one small job and no GPU: OpenAI's fine-tuning API or a hosted trainer is far less work. If you will only ever train on one consumer GPU, Unsloth's single-GPU notebooks are the gentler start.
- What is the most common mistake when self-hosting Axolotl?
- A run can finish cleanly and still produce a model that does not know when to stop. Axolotl's chat-template docs warn that your tokenizer's eos_token must match the EOS token in the chat template — otherwise set eos_token under special_tokens — and a mismatch means the model never learns where its turn ends, so it runs on past its answer. Before a long run, inspect tokenised examples with axolotl preprocess config.yml --debug.