About this document
Tokencake: A KV-Cache-centric Serving Framework For LLM-based Multi-Agent Applications by a791300878 is a document available to read on EtoBox.
Tokencake is a KV-Cache-centric serving framework designed to optimize the performance of large language models (LLMs) in multi-agent applications by addressing challenges of space contention and time underutilization. It employs a Space Scheduler for dynamic memory partitioning to protect critical agents and a Time Scheduler for proactive offloading and predictive uploading of KV Cache during function call stalls. Evaluation results indicate that Tokencake can significantly reduce end-to-end latency by ove
- Author
- a791300878
- Language
- EN