No data center required to start. A single consumer graphics card is enough to run a compact model - and a 24 GB card is enough to finetune your own. Scale up to multiple GPUs or a cluster when you need a larger model.
If it has enough memory and a supported driver stack, it works. The two big GPU families are covered.
RTX 3090, 4090 and 5090 lead the way, with the rest of the RTX 30/40/50 series close behind - up to data-center cards (A100, H100, H200) for large models. The most widely tested path.
Train & runRadeon RX 7900 XTX / 7900 XT, with the 24 GB 7900 XTX as the sweet spot, up to Instinct MI-series accelerators. A full first-class training path.
Train & runVRAM is what matters most. Running a compact model is light; finetuning one needs a bit more headroom.
QLoRA keeps memory low by loading the base in 4-bit. Exact limits shift with model size and adapter rank - these are practical guidelines, not hard cutoffs. Larger models scale across multiple GPUs or a cluster.
Linux is the smoothest path for both NVIDIA and AMD, and Windows works through WSL. The models themselves are plain files - they don't care what OS made them.
If you can play modern games on it, you can probably finetune a model on it - and when you need more, the same pipeline runs on a cluster.
See how it works →