Reward-aware population scaling of evolutionary strategies in LLM fine-tuning
Read the original at arxiv.org→arXiv:2607.19408v1 Announce Type: new Abstract: Using Evolutionary Strategies (ES) for fine-tuning large language models is attractive because it is memory-efficient, parallel, and compatible with black-box or...
Original headline: "Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning"
Coverage timeline
- Jul 23, 04:00 UTC arXiv cs.LG lead source Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning