Can I read Optimal 4-Bit Precision for LLMs on EtoBox?
Optimal 4-Bit Precision for LLMs by Przemysław Maksymowicz is a document available to read on EtoBox.
What is Optimal 4-Bit Precision for LLMs about?
The document discusses quantization methods for reducing the number of bits required to represent model parameters, trading accuracy for smaller size and faster inference. The authors study how zero-shot performance of language models scales with total model bits and quantization precision from 3 to 16 bits. Their main finding is that 4-bit precision optimizes zero-shot accuracy for a fixed number of model bits across different model families and sizes from 19M to 176B parameters.
- Author
- Przemysław Maksymowicz
- Language
- EN