Quantization
Converting weights to lower precision to save memory, with little quality loss if done well.
Converting weights to lower precision to save memory, with little quality loss if done well. (M08)
Storing a model's numbers in fewer bits, trading a little accuracy for big memory savings and often more speed.