Online Learning with LLM Experts from Limited Feedback
Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.
Abstract
Adaptive routing of prompts to LLM experts is formulated as a contextual bandit problem with limited feedback, yielding algorithms with sublinear regret bounds and effective routing strategies.
We study adaptive routing of prompts to large language model (LLM) experts to maximize response quality in an online setting with limited feedback. We formulate it as a bandit problem with K actions that represent experts and d features that encode prompts, over a horizon of T rounds. We propose algorithms that strategically select and observe rewards to minimize regret. In the full-information setting, we achieve a regret of O(d T / m), while in the bandit setting we achieve O(d T K / m), where m ll T is a budget on feedback. Our experiments show that we efficiently learn high-quality routing strategies across diverse LLMs from limited feedback.
Community
Get this paper in your agent:
hf papers read 2609.05820 curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper
No model linking this paper
Datasets citing this paper
No dataset linking this paper
Spaces citing this paper
No Space linking this paper
Collections including this paper
No Collection including this paper
More from Hugging Face Daily Papers
-
CERA-MoA: Co-Evolving Routing Mechanisms with Continually Learning LLM Agents
Sep 17
-
Fathom: Per-Query Read Depth for Sparse Decoding over Offloaded KV Caches
Sep 17
-
In-Context Robot Learning with VLM Agents
Sep 17
-
Assessing nnU-Net Generalization across Brain Tumor Populations in BraTS-GoAT 2026
Sep 17
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.