← World Of AI

Generative AI

QLoRA

Fine-tuning a quantised model through small adapter matrices, so it fits on one GPU.

Low-rank adaptation freezes the original weights and trains a pair of small matrices alongside them. QLoRA adds four-bit quantisation of the frozen base, cutting memory again.

The practical effect is that fine-tuning a substantial open model becomes a consumer-hardware task rather than a cluster task, and the resulting adapter is small enough to ship as a file.

Also in Generative AI

Apply

Begin the first module

Become AI native, it is the real deal today, and if it is not for you, you have lost nothing but learnt a new skill.