SkillScriptBench benchmarks self-evolution of executable agent skill packages across documentation repair, script repair, and preservation
Read the original at arxiv.org→arXiv:2610.04008v1 Announce Type: new Abstract: Executable Agent Skills combine natural-language instructions and scripts into reusable packages for LLM agents, and revising them requires fixing errors without...
Original headline: "SkillScriptBench: Benchmarking Self-Evolution of Executable Agent Skill Packages Beyond Markdown"
Coverage timeline
- Oct 6, 04:00 UTC arXiv cs.AI lead source SkillScriptBench: Benchmarking Self-Evolution of Executable Agent Skill Packages Beyond Markdown