Nexus proposes depth-adaptive KV-Cache splicing and retrieval-decoupled tool routing for agentic LLMs on unified memory
Read the original at arxiv.org→arXiv:2608.20397v1 Announce Type: new Abstract: Agentic large language models (LLMs) on the Model Context Protocol (MCP) re-encode verbose tool schemas every turn, so prefill - quadratic in sequence length -...
Original headline: "Nexus: Depth-Adaptive KV-Cache Splicing and Retrieval-Decoupled Tool Routing for Agentic LLMs on Unified Memory"
Coverage timeline
- Aug 24, 04:00 UTC arXiv cs.AI lead source Nexus: Depth-Adaptive KV-Cache Splicing and Retrieval-Decoupled Tool Routing for Agentic LLMs on Unified Memory