macrostack

Layer 4 · self-hosting reality check

What it actually takes to self-host Axolotl

The docs say 11 GB of VRAM — Axolotl's Llama-3 examples list LoRA on an 8B model (8-bit base) as "Single GPU @ 11GB VRAM"; their Qwen3 14B QLoRA Colab runs on a 16 GB T4, at 9.26 GB after loading. In practice you want 24 GB of VRAM for comfortable LoRA or QLoRA on 7–14B models. Full fine-tuning an 8B is listed at 48 GB on one GPU, and 70B QLoRA+FSDP as "Dual GPU @ 21GB VRAM". Plenty of system RAM too — the FAQ says exit code -9 means you ran out of system RAM, not VRAM.. Here is the honest version — real requirements, real monthly cost, what you will be maintaining, and the one thing that catches people out.

Usually reached from OpenAI Fine-Tuning alternatives, where Axolotl is one of the picks.

Wondering whether you need to at all? Is OpenAI fine-tuning free? — what the free tier actually allows, and where the wall is.

RAM — documented minimum11 GB of VRAM — Axolotl's Llama-3 examples list LoRA on an 8B model (8-bit base) as "Single GPU @ 11GB VRAM"; their Qwen3 14B QLoRA Colab runs on a 16 GB T4, at 9.26 GB after loading
RAM — what it really needs24 GB of VRAM for comfortable LoRA or QLoRA on 7–14B models. Full fine-tuning an 8B is listed at 48 GB on one GPU, and 70B QLoRA+FSDP as "Dual GPU @ 21GB VRAM". Plenty of system RAM too — the FAQ says exit code -9 means you ran out of system RAM, not VRAM.
CPUNVIDIA Ampere or newer (RTX 30/40-series, A100, L4, H100) for bf16 and Flash Attention, or AMD. Blackwell cards need CUDA 13.0. Linux only — Windows via WSL2 or Docker.
DiskThe base model in bf16 (about 16 GB for an 8B), the tokenised dataset cache, a checkpoint per save, and the merged model at the end — budget three to four times the base model's size
Monthly costPaid per run, not per month: a 24 GB RTX 4090 on RunPod is $0.34–0.74/hr and an 80 GB A100 $1.19–1.59/hr (read 2026-09-30). A small LoRA run is measured in hours, so the bill is usually tens of dollars; $0 on a 24 GB card you already own.
Setup timeAn hour with the Docker image; an afternoon if you are matching CUDA, PyTorch and flash-attention versions by hand
How you install ituv pip install --no-build-isolation axolotl[deepspeed] on Python 3.12 with PyTorch 2.13+, or the axolotlai/axolotl-uv Docker image; then axolotl fetch examples and axolotl train <config>.yml
Ongoing maintenanceVery active — last pushed 2026-09-30. Minimum Python and PyTorch versions move with it (now Python 3.12, PyTorch 2.13), so pin the Docker tag per project and re-test old YAML configs after upgrading.
Where it stops scalingFrom a single 16 GB card (QLoRA) to multi-GPU and multi-node training via FSDP2 and DeepSpeed, with sequence parallelism for long contexts — the same YAML file scales up.

The thing that catches people out

A run can finish cleanly and still produce a model that does not know when to stop. Axolotl's chat-template docs warn that your tokenizer's eos_token must match the EOS token in the chat template — otherwise set eos_token under special_tokens — and a mismatch means the model never learns where its turn ends, so it runs on past its answer. Before a long run, inspect tokenised examples with axolotl preprocess config.yml --debug.

When not to self-host Axolotl

You have one small job and no GPU: OpenAI's fine-tuning API or a hosted trainer is far less work. If you will only ever train on one consumer GPU, Unsloth's single-GPU notebooks are the gentler start.

Every guide here carries this section. A site that only ever tells you to self-host is selling something — the useful answer is sometimes no.

Other Layer 4 self-hosting guides

Common questions

How much RAM does Axolotl actually need?
24 GB of VRAM for comfortable LoRA or QLoRA on 7–14B models. Full fine-tuning an 8B is listed at 48 GB on one GPU, and 70B QLoRA+FSDP as "Dual GPU @ 21GB VRAM". Plenty of system RAM too — the FAQ says exit code -9 means you ran out of system RAM, not VRAM. in practice. The documented minimum is 11 GB of VRAM — Axolotl's Llama-3 examples list LoRA on an 8B model (8-bit base) as "Single GPU @ 11GB VRAM"; their Qwen3 14B QLoRA Colab runs on a 16 GB T4, at 9.26 GB after loading, which is the figure at which the process starts rather than the figure at which it works under real use. NVIDIA Ampere or newer (RTX 30/40-series, A100, L4, H100) for bf16 and Flash Attention, or AMD. Blackwell cards need CUDA 13.0. Linux only — Windows via WSL2 or Docker. alongside it.
What does self-hosting Axolotl cost per month?
Paid per run, not per month: a 24 GB RTX 4090 on RunPod is $0.34–0.74/hr and an 80 GB A100 $1.19–1.59/hr (read 2026-09-30). A small LoRA run is measured in hours, so the bill is usually tens of dollars; $0 on a 24 GB card you already own. This is commodity VPS pricing and excludes your time, which is the larger cost for most people — budget for very active — last pushed 2026-09-30. Minimum Python and PyTorch versions move with it (now Python 3.12, PyTorch 2.13), so pin the Docker tag per project and re-test old YAML configs after upgrading.
How long does it take to set up Axolotl?
An hour with the Docker image; an afternoon if you are matching CUDA, PyTorch and flash-attention versions by hand, via uv pip install --no-build-isolation axolotl[deepspeed] on Python 3.12 with PyTorch 2.13+, or the axolotlai/axolotl-uv Docker image; then axolotl fetch examples and axolotl train <config>.yml.
When should I NOT self-host Axolotl?
You have one small job and no GPU: OpenAI's fine-tuning API or a hosted trainer is far less work. If you will only ever train on one consumer GPU, Unsloth's single-GPU notebooks are the gentler start.
What is the most common mistake when self-hosting Axolotl?
A run can finish cleanly and still produce a model that does not know when to stop. Axolotl's chat-template docs warn that your tokenizer's eos_token must match the EOS token in the chat template — otherwise set eos_token under special_tokens — and a mismatch means the model never learns where its turn ends, so it runs on past its answer. Before a long run, inspect tokenised examples with axolotl preprocess config.yml --debug.
The Macrostack brief

New swaps, worth your inbox.

A short, occasional email when we add a high-intent alternative or ship a new head-to-head. No spam, no selling your address — unsubscribe in one click.