← World Of AI
Generative AI
QLoRA
Fine-tuning a quantised model through small adapter matrices, so it fits on one GPU.
Low-rank adaptation freezes the original weights and trains a pair of small matrices alongside them. QLoRA adds four-bit quantisation of the frozen base, cutting memory again.
The practical effect is that fine-tuning a substantial open model becomes a consumer-hardware task rather than a cluster task, and the resulting adapter is small enough to ship as a file.
Apply
Begin the first module
Become AI native, it is the real deal today, and if it is not for you, you have lost nothing but learnt a new skill.