Supported hardware

From the GPU you have
to a full cluster.

No data center required to start. A single consumer graphics card is enough to run a compact model - and a 24 GB card is enough to finetune your own. Scale up to multiple GPUs or a cluster when you need a larger model.

Which cards work

If it has enough memory and a supported driver stack, it works. The two big GPU families are covered.

🟩

NVIDIA

CUDA

RTX 3090, 4090 and 5090 lead the way, with the rest of the RTX 30/40/50 series close behind - up to data-center cards (A100, H100, H200) for large models. The most widely tested path.

Train & run
🟥

AMD

ROCm

Radeon RX 7900 XTX / 7900 XT, with the 24 GB 7900 XTX as the sweet spot, up to Instinct MI-series accelerators. A full first-class training path.

Train & run

What you can do, by memory

VRAM is what matters most. Running a compact model is light; finetuning one needs a bit more headroom.

GPU memory
Run a model
Finetune a model
Typical cards
8 GBentry
smaller models
-
RTX 3060, RX 7600
12-16 GBmainstream
comfortable
smaller models
RTX 4070, RX 7800 XT
24 GBrecommended
most models
mid-size models
RTX 4090 / 3090, RX 7900 XTX
48 GB+workstation / cluster
large models
large models, multi-GPU
RTX 5090 (x N), A100/H100/H200

QLoRA keeps memory low by loading the base in 4-bit. Exact limits shift with model size and adapter rank - these are practical guidelines, not hard cutoffs. Larger models scale across multiple GPUs or a cluster.

🐧 Operating system

Linux is the smoothest path for both NVIDIA and AMD, and Windows works through WSL. The models themselves are plain files - they don't care what OS made them.

One good card is your whole AI lab.

If you can play modern games on it, you can probably finetune a model on it - and when you need more, the same pipeline runs on a cluster.

See how it works →
Quantinia
© 2026 Quantinia