AI Glossary term

Quantization

Storing a model's numbers at lower precision so it needs less memory and runs faster.

Also called: 4-bit, GGUF

Storing a model's numbers at lower precision so it needs less memory and runs faster.

Why it mattersHow a 70B model ends up running on a decent laptop with only a small quality hit.

See also Open-Weights Model, Distillation

Explore more AI terms

Browse the full glossary for plain-English definitions across models, agents, data, and safety.

Back to all terms