Skip to content

Opening book details…

Can I read FPGA-Based QLlama for Efficient Llama2 Inference on EtoBox?

FPGA-Based QLlama for Efficient Llama2 Inference by yifei.star.wang is a document available to read on EtoBox.

What is FPGA-Based QLlama for Efficient Llama2 Inference about?

The document presents QLlama, an FPGA-based accelerator designed for energy-efficient inference of the Llama2 large model, utilizing a novel microscaling quantization method that significantly reduces hardware complexity. It achieves energy efficiency improvements of 2.13 to 10.66 times with minimal accuracy loss by implementing a mixed precision configuration and dedicated computational units for quantized data. The proposed architecture and optimizations enable effective deployment of large models in edge

Author
yifei.star.wang
Language
EN

More by yifei.star.wang

Browse all works by yifei.star.wang