Fathom: per-query read depth for sparse decoding over offloaded KV caches
Read the original at arxiv.org→arXiv:2609.17652v1 Announce Type: new Abstract: When agentic sessions run to a million tokens with many sessions resident at once, the KV cache and the index that ranks it live in host memory, and the scan that...
Original headline: "Fathom: Per-Query Read Depth for Sparse Decoding over Offloaded KV Caches"
Coverage timeline
- Sep 17, 04:00 UTC arXiv cs.LG lead source Fathom: Per-Query Read Depth for Sparse Decoding over Offloaded KV Caches