LLaMCAT: Optimizing Large Language Model Inference with Cache Arbitration and Throttling

Publication
ICPP 2025