Can I read SparseServe: Efficient LLM Serving System on EtoBox?
SparseServe: Efficient LLM Serving System by lambda shi is a document available to read on EtoBox.
What is SparseServe: Efficient LLM Serving System about?
SparseServe is a long-context LLM serving system designed to enhance the efficiency of dynamic sparse attention algorithms (DSAs) by implementing hierarchical HBM-DRAM management. It introduces innovations such as fragmentation-aware KV cache transfer, working-set-aware batch size control, and layer-segmented prefill to address challenges in serving long-context LLMs. Experimental results demonstrate that SparseServe significantly reduces latency and improves token generation throughput compared to existing
- Author
- lambda shi
- Language
- EN