Skip to content

Opening book details…

About this document

LPDDR-based CXL-PNM for LLM Inference by yanliang6156 is a document available to read on EtoBox.

The document discusses a new CXL-PNM platform for efficiently accelerating inference of large transformer-based language models. The platform uses an LPDDR5X memory architecture connected via the Compute eXpress Link to overcome the limitations of GPUs for models that are too large to fit in GPU memory, such as GPT-3.5 which requires over 300 GB of memory and 1.4 TFLOPs of computation for inference.

Author
yanliang6156
Language
EN