About this document
LLM Inference Latency Metrics Explained by vineet.theodore is a document available to read on EtoBox.
The document discusses key metrics for evaluating LLM (Large Language Model) inference performance, focusing on latency, throughput, and goodput. It highlights various latency metrics such as Time to First Token (TTFT) and Total Latency, as well as throughput measures like Requests per Second (RPS) and Tokens per Second (TPS). Additionally, it emphasizes the importance of balancing latency and throughput to optimize user experience while meeting service-level objectives.
- Author
- vineet.theodore
- Language
- EN