Most retrieval and recommendation pipelines assume items are chosen by how well they satisfy user requirements. This paper shows a different, consequential mechanism: LLM-based agents often rely on item source as a shortcut, causing systematically higher selection rates for some sources even when their items satisfy requirements less well. That mismatch can steer users to worse options and distort which services get traffic.
Key Findings
- Consistent source preferences across models and domains: across 12 agent models and three domains (shopping, accommodation, scholarly search), each model prefers some sources and avoids others, and models largely agree on which sources are preferred. This means source bias is not isolated to a single model or domain.
- Source preference can beat item quality: when a preferred-source item is one requirement short of a dispreferred-source item, agents still select the preferred-source item about two-thirds of the time. The reverse rarely occurs. So what: selection behavior can degrade user outcomes and unfairly penalize dispreferred sources.
- The source label itself matters: hiding the source weakens preferences; relabeling a dispreferred item as coming from a preferred source raises its selection rate. So what: attribution signals directly influence agent choices independent of item content.
- Mitigations work but differ in effect: supplying missing item information and prompts that counter preconceptions reduce source preference; training dynamics that reward better items can inadvertently strengthen source shortcuts. So what: both interface-level fixes (hide/relabel or supply info) and careful training objectives are needed.
Who it's for and trade-offs
Great fit if you design or evaluate LLM-driven retrieval, recommendation, or agent systems and need to understand selection biases that go beyond ranking metrics. Use the paper's controlled comparisons and mitigation tests to audit agent behavior before deploying selection agents commercially. Look elsewhere if you only need low-level model fine-tuning recipes or deployment tooling; the paper focuses on behavioral analysis, cross-model benchmarking, and high-level mitigation, not on production integration details.
Methods (brief)
The study uses realistic request benchmarks across three domains (WebShop, HotelQuEST, ScholarGym), compares items that satisfy identical requirements at the same result positions, and evaluates 12 widely used agent models. Analyses include controlled source hiding/relabeling experiments and tests of training- and information-based mitigation strategies.