Context-Aware CodeLLM Eviction (CACE)

Kishanthan Thangarajah1, Boyuan Chen1, Shi Chang1, Ahmed E. Hassan1

1Centre for Software Excellence, Huawei Canada

Paper · Slides · Code

CACE keeps the right CodeLLMs loaded to avoid cold starts.

Abstract

Self-hosted CodeLLMs are increasingly used inside enterprises to support code completion, refactoring, and reasoning across multiple languages and teams. However, GPU memory is limited, and organizations often host many specialized models (per language and task), which cannot all stay loaded at once. Existing systems typically rely on recency-based (LRU) eviction, which ignores model load cost, task criticality, and future demand, leading to frequent cold starts and poor developer experience.

We present CACE (Context-Aware CodeLLM Eviction), a multi-factor eviction policy that scores models using four signals: request recency, reload cost, predicted future demand, and task criticality. CACE keeps latency-critical and soon-needed models resident, reducing unnecessary reloads. Across multilingual completion (TTFT-sensitive) and reasoning (E2E-sensitive) workloads, CACE cuts TTFT by up to 70%, reduces E2E latency by up to 37%, increases cache hit rate to about 0.86 (≈ +55%), and roughly halves reload overhead compared to LRU.

Motivation: Self-hosting many CodeLLMs on limited GPUs

Language diversity

Enterprise teams span front-end, back-end, DevOps, and automation, and rely on different programming languages and tasks. To protect proprietary code and meet compliance requirements (e.g., GDPR) while avoiding network latency from external APIs, many organizations choose to self-host CodeLLMs.

Privacy and latency needs

Self-hosting keeps code and intellectual property within organizational boundaries, simplifies compliance, and removes external round-trip latency, which is critical for responsive IDE experiences. Developers expect sub-200 ms feedback for code completions, not multi-second delays.

GPU limitations → cold starts

In practice, enterprises need to host many specialized models (per language and task) on a limited GPU budget. The serving system must continually load and unload models. Naïve LRU-based eviction often evicts exactly the models that will be needed next. Cold starts can take 2–3 seconds, disrupting IDE flow and harming developer productivity.

Why LRU is not enough

Method: Context-Aware CodeLLM Eviction (CACE)

CACE replaces recency-only eviction with a multi-factor scoring function. Instead of blindly evicting the least recently used model, CACE computes an eviction score per model and evicts the least valuable one.

Four factors (P1–P4)

Models with low “keep” value are selected as eviction candidates. CACE preferentially keeps models that are expensive to reload, frequently used, or critical for latency-sensitive IDE interactions.

Fewer evictions, more stable residency

In timeline-style workloads, CACE:

This reduces cold starts and smooths tail latency.

Results: Consistent gains across workloads

We evaluate CACE on multilingual completion (TTFT-sensitive) and reasoning (E2E-sensitive) tasks across eight popular programming languages and three realistic workload patterns: a uniform mix, an IDE-heavy completion skew, and a popularity-skewed distribution.

These improvements hold across all evaluated workload patterns, indicating that CACE is robust to different mixes of completion vs reasoning requests and language popularity skews in real developer workflows.

Ablation: Which eviction factors matter most?

We ablate each component of CACE to understand its contribution to performance.

Overall, CACE’s gains are primarily driven by protecting latency-critical requests and anticipating upcoming demand, with recency and reload cost providing additional fine-tuning.

Resources

BibTeX

@inproceedings{thangarajah2025cace,
  title     = {Context-Aware CodeLLM Eviction (CACE): Reducing Cold-Start Latency for Self-Hosted CodeLLMs},
  author    = {Thangarajah, Kishanthan and Chen, Boyuan and Chang, Shi and Hassan, Ahmed E.},
  booktitle = {Proceedings of the ASE 2025 Industry Track},
  year      = {2025},
  note      = {To appear}
}

This page was built using the Academic Project Page Template which was adopted from the Nerfies project page. You are free to borrow the source code of this website, we just ask that you link back to this page in the footer.
This website is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.