Top-k Pareto bandits: hypervolume regret for multi-objective slate selection
Read the original at arxiv.org→arXiv:2607.26273v1 Announce Type: new Abstract: We consider a stochastic multi-objective bandit problem where, at each round, the agent selects a slate of $k$ arms and observes their $d$-dimensional reward vectors...
Original headline: "Top-$k$ Pareto Bandits: Hypervolume Regret for Multi-Objective Slate Selection"
Coverage timeline
- Jul 30, 04:00 UTC arXiv cs.LG lead source Top-$k$ Pareto Bandits: Hypervolume Regret for Multi-Objective Slate Selection