Sign Up to like & get
recommendations!
0
Published in 2025 at "IEEE Computer Architecture Letters"
DOI: 10.1109/lca.2025.3567844
Abstract: The key-value (KV) cache in large language models (LLMs) now necessitates a substantial amount of memory capacity as its size proportionally grows with the context’s size. Recently, Compute-Express Link (CXL) memory becomes a promising method…
read more here.
Keywords:
memory;
cxl memory;
llm inference;
oasis outlier ... See more keywords