Skip to content

Opening book details…

About this document

Tokencake: A KV-Cache-centric Serving Framework For LLM-based Multi-Agent Applications by a791300878 is a document available to read on EtoBox.

Tokencake is a KV-Cache-centric serving framework designed to optimize the performance of large language models (LLMs) in multi-agent applications by addressing challenges of space contention and time underutilization. It employs a Space Scheduler for dynamic memory partitioning to protect critical agents and a Time Scheduler for proactive offloading and predictive uploading of KV Cache during function call stalls. Evaluation results indicate that Tokencake can significantly reduce end-to-end latency by ove

Author
a791300878
Language
EN