Contexts, Conversations & Connections: Spotify Research at RecSys 2026

Feature Image

At Spotify, we have long believed that great recommendations are about more than predicting the next track. It is about understanding listeners deeply, guiding them through a vast and growing world of art and ideas while surprising them with new discoveries. New advances in language models and agentic systems let us build experiences that give people more interactivity and control than ever before, and that is at the center of what Spotify is bringing to RecSys 2026.

For RecSys 2026 in Minneapolis, we are proud to share our largest presence to date: ten papers across the main conference and workshops, two co-organized workshops, keynotes and invited talks at USRW, CARS, the RecSys Challenge on Conversational Music Recommendation, and DaQuaMRec, and two tutorials. As a Silver sponsor, Spotify researchers are also contributing to the conference as members of the executive committee, program committees, and reviewers.

Our contributions center around three themes:

  • Agents, Conversations & the Future of Recommendation. With agentic and conversational technologies, we can build new forms of interactivity into the user experience. We explore what this means for recommendation, both philosophically and practically, and how Spotify is building conversational recommendation at scale.

  • Scaling Personalization with LLMs. We explore how new modeling approaches can power personalized experiences, from hypothesis-driven content shelves and robust search reranking to natural-language user profiles and causal recommendation policies.

  • Rethinking Evaluation. We challenge established assumptions about how we evaluate recommender systems, uncovering biases in LLM-as-a-judge approaches, revisiting what level of sequential model complexity current benchmarks actually require, and examining the gap between reasoning quality and recommendation effectiveness.

Keep reading to learn more about our contributions.

Agents, Conversations & the Future of Recommendation

With agentic technologies, we can rethink how people interact with content, but this demands deep research into how we build conversational systems that recommend effectively through dialogue. These two papers address both the conceptual and the practical sides of this space.

  • Who Are We Recommending To? Recommender Systems in the Agentic Web As we build AI agents that mediate how users discover and consume content, this paper examines the implications for recommender systems. When the "user" is an agent acting on behalf of a human, our assumptions about preferences, engagement signals, and evaluation metrics all need rethinking. We outline the research challenges this creates and propose directions for the field. 

Main Conference - Past, Present and Future Track

  • Bootstrapping Conversational Recommendation Agents At Spotify: Synthetic Data Generation and Self-Improvement Loops  Building a conversational recommendation agent requires training data that does not naturally exist at scale. This paper introduces Spotify's approach to bootstrapping such agents using synthetic data generation and self-improvement loops that enable the system to learn from its own interactions and continuously refine its conversational recommendations. 

Main Conference - Industry Track

Scaling Personalization with LLMs

With new modeling approaches, we can personalize recommendations in ways that were not previously possible. From large language models that let us understand user intent and represent user taste in new ways, to causal methods that help us update recommendation policies more precisely, each of these papers reflects a different facet of how we scale personalization at Spotify.

  • Hypothesis-Driven Shelf Generation for Personalised Recommendation The shelves on Spotify's home page are one of the primary surfaces through which users discover content. This paper explores a hypothesis-driven approach to generating personalized shelves that dynamically creates experiences tailored to individual users. 

Main Conference - Industry Track

  • Textual User Taste: Natural-Language User Context for Foundation-Model Recommender System at Scale Rather than encoding user preferences as opaque embeddings, this work represents user taste as natural-language descriptions, enabling foundation recommendation models to reason about preferences in an interpretable and transferable way. By making user context legible to LLMs, this approach opens new possibilities for personalized content recommendation at scale.  

CARS Workshop (Context-Aware Recommender Systems)

  • Robust Behavioural Feature Injection for LLM Cross-Encoders in Personalised Search LLM-based search rerankers risk over-relying on historical behavioral signals, degrading performance on new or sparse queries. This paper shows how dual-sample feature-dropout training effectively balances the exploitation of strong behavioral priors with robustness on unseen queries, making LLM rerankers more reliable for personalized search at scale. 

USRW Workshop (Unified Search & Recommendation)

  • Incremental Recommendation via Causal Models Recommendation policies typically require expensive full retraining to incorporate new evidence. This work brings causal reasoning to the recommendation setting, showing how causal models can support targeted, incremental policy updates, adjusting recommendations based on new data without rebuilding the entire model.

CONSEQUENCES Workshop

Rethinking Evaluation

As LLM-as-a-judge becomes a common evaluation approach and sequential models grow ever more complex, we think it is important to interrogate our assumptions. Are LLM judges reliable? What level of sequence modeling do current benchmarks actually need? Does better reasoning produce better recommendations? This cluster of papers examines established evaluation practices and explores new methodological directions.

  • Intent-Description Anchoring Bias in LLM-as-a-Judge Evaluation of Recommendation Systems LLM-as-a-judge is increasingly used to evaluate recommendations, but how the intent or task is described to the LLM can systematically bias its judgments. This paper identifies and characterizes this anchoring bias, demonstrating that seemingly minor changes in intent descriptions can significantly shift evaluation outcomes.

Main Conference - Research & Practice Notes

  • From IR to RecSys: Evaluating LLM-based Judges in Cranfield-style Recommendation Collections The Cranfield evaluation paradigm has been foundational in information retrieval for decades. This paper adapts it for recommender systems evaluation using LLM-based judges, examining whether the methodology transfers effectively and what adjustments are needed to produce reliable offline evaluations of recommendation quality.

USRW Workshop (Unified Search & Recommendation)

  • Do Sequential Recommendation Benchmarks Really Require Higher-Order Sequence Modelling? This paper takes a closer look at the role of higher-order sequence modeling in current recommendation benchmarks, asking when added complexity helps and when simpler approaches may perform comparably. The findings offer useful perspective for both research direction and practical system design.

Main Conference - Research & Practice Notes

  • The Disconnect Between Better Descriptive Reasoning Trace Quality and Recommendation Effectiveness This paper investigates the relationship between the quality of an LLM's reasoning traces and the effectiveness of its recommendations, finding that improvements in one do not always lead to improvements in the other. The result raises interesting open questions about how LLMs use reasoning in recommendation contexts. 

GenAIECommerce Workshop

Closing

This is Spotify's most comprehensive RecSys presence to date. We are excited to share these contributions, learn from the whole community, and continue the conversation about where recommender systems are headed next. If you are at RecSys 2026 in Minneapolis, we would love to connect. Come find us at the Spotify booth, at one of our talks, or throughout the main conference and workshops.

Stay tuned for deeper dives into many of these papers on the Spotify blog.

Acknowledgements

This work is the result of a collaborative effort across multiple teams at Spotify. We would like to extend our sincere gratitude to all the co-authors and contributors who made these projects possible:

Abenezer Abebe, Himan Abdollahpouri, Karen Banzon, Shubham Bansal, Paul Bennett, Marcus Better, Anton Blomberg, Hugues Bouchard, Adrià Casas, Tarun Chillara, Ann Clifton, Melissa Crawford, Edoardo D'Amico, Andreas Damianou, Seda Davtyan, Lucas de Haas, Anurag Deshpande, Christine Doig Cardet, Jackie Doremus, Dani Doro, Juan Elenter Litwin, Francesco Fabbri, Ghazal Fazelnia, Erik Franco, Hugo Galvão, Peng Ge, Paul Gigioli, Aloïs Gruson, David Gustafsson, Claudia Hauff, Jeremy Hopple, Maya Hristakeva, Binal Jhaveri, Eliza Klyce, Kyle Kretschman, Ben Lacker, Mounia Lalmas, Daniel Lazarovski, Ciarán Lee, James Leoni, Henrik Lindström, Erik Lybecker, Roberto Mirizzi, Matthew Moellman, David Murgatroyd, Gabriel Negash, Victor Ode, Michael O'Riordan, Enrico Palumbo, Gustavo Penha, Aleksandr V. Petrov, Tasnim Rahman, Yves Raimond, Praveen Ravichandran, Sai Srivatsa Ravindranath, José Luis Redondo García, Kate Remeika, Emma Schüldt, Nandini Singh, Yabai Song, Nathan Stein, Alina Susoykina, Alexandre Tamborrino, Ye Myat Thein, Ali Vardasbi, Athanasios Vlontzos, Alice Wang, Jacqueline Wood, Katie Zelvin, Sharon Zheng