Intro0:00
Better on time. Uh, hi everyone. I am Shivam, I'm from Spotify, and my talk is going to be about how at Spotify we do personalization, um, especially in the era of LLMs. So, uh, for those of you who use Spotify, um, I guess, do we have any Spotify users in the room?
Raise your hand. Nice. Uh, yeah, thanks for using Spotify. And, uh, a bit like, in this talk, so this is going to be less about context engineering from the conventional, like, agentic sense. It's going to be more about how we do context engineering on the modeling side.
So if you're interested at all on how your Spotify app works, how we recommend you songs, tracks, episodes, etc., this talk is going to be really useful for you to contextualize how we use your data for, for just recommending stuff to you that you like.
Um, so a bit about me. Um, I am the tech lead of, uh, the user representations team in Spotify's AI foundation org. So, uh, the AI foundation team builds all of the frontier foundational models that, that are used across the entire stack of recommendations at Spotify.
Um, so we do things like user representations, content representations, as well as, like, adapting open-weight LLMs. Um, we do CPT, SFT, like, all of the stuff that a lot of the frontier labs are doing, as well as, like, other sort of, uh, competitors are doing.
And we try to kind of make sure that, that our, our we're building the best, uh, music recommender system possible, uh, for you guys. So my, my background is as a machine learning engineer. I used to work at Twitter, and I live in London.
So if you're around for a coffee chat, like, I'm very happy to, uh, after this talk or in general, like, talk about this stuff, uh, because it's really cool. Um, so three things I'm going to talk about today.
User Embeddings1:59
Uh, the first thing is going to be about foundational user modeling. Uh, so what is what, what is that? That's essentially us trying to understand you, the users. Um, these are, like, the three key aspects that we think are, like, very key aspects of the future of personalization in the era of LLMs.
So there's the, the user modeling component. Uh, then there's the content side. So how can you kind of teach LLMs about the content that you have, the catalog that you have on, on Spotify or whatever your sort of, um, your, uh, platform is?
And then the last thing, once you have these two, like, pieces of the puzzle, are building these, both of these together into, like, something which is steerable and personalized. So as, as sort of steerable as possible. And I'm, I'm going to talk a bit more about that.
So generally, like, you can think about it as going from, like, sequences of actions to, to vectors. Uh, and then once you have the vectors, you can go to tokens. And then once you have the tokens, you can actually combine the vectors and the tokens with the LLM that you have to get, like, like, a pipeline where you, you have something which is as personalized as possible.
Um, so a bit about Spotify. Uh, for those of you who have, haven't used it, uh, we have, like, about 750 million usersright now, uh, MAUs. Uh, we have, like, a hund catalog of, like, 100 million-plus tracks. Uh, we have, I think, about 350, I think it's more like 400,000 audiobooks now, uh, millions of podcasts, and, like, a lot of video episodes as well.
So we're, we're kind of, uh, more and more creators are sort of switching to video as, as a modality. And we, we're definitely kind of making sure that we support that. Um, and we're in about 184 markets. Um, so as you can see, like, we, we have a lot of users.
We have a lot of data and content. How can we combine all of that to build something which is, which is as useful for our users as possible? Um, the way we do that is we obviously, we've been using machine learning models for at least, like, a decade, if not more.
Um, you might have heard of Discover Weekly, or you're probably a user of that. Um, that's been around since, I think, 2015, uh, back when I was in grad school. And that was one at that time, that was, like, for me, like, one of the coolest products, uh, with, like, my interactions across, like, the tech sort of stack that I was using at the time.
Just because, like, it's something which, which is, like, unique to you. It keeps changing. Um, and over the last decade or so, we've kind of added more and more personalization and products, uh, to the, to the app. Um, so we have, like, you know, we have a bunch of verticals.
We have a bunch of new product surfaces. Uh, we have this thing called the AI DJ, where you can kind of talk to it, and it kind of recommends you stuff, or it plays stuff for you. Uh, we also have, like, a prompted playlist where you can actually prompt the model, uh, like, Spotify's model, and it kind of generates, like, a custom playlist for you based on your prompts.
And as of this week, it also supports podcasts. So if you want, you can kind of just prompt it, and it'll create, like, a, like, a playlist of episodes for you based on whatever you're looking for. Um, so that is sort of the future that we're heading towards, where users have steerability.
Uh, users can kind of talk to Spotify in natural language. And, uh, now we also have something called the taste profile. So this is, this is only supported, like, in a few markets, but this is going to be expanded, uh, later this year.
And the idea is that we want to expose what we know about you. Um, and then we want to kind of let you choose which part of that you want us to kind of keep, which part of that you want us to forget, and, you know, just allow you to kind of have as much control as possible.
Um, as we kind of work on this, like, just sort of some context for those of you who've not worked in the space of, like, recommender systems and machine learning. So, uh, what we call TradRex, which used to be, like, the sort of predominant, uh, sort of paradigm of building recommender systems up until a few years ago, uh, is essentially, like, this multi-step pipeline where you have, like, a massive catalog of items.
Um, you have this candidate generation step, which kind of reduces that item space to from, like, millions to, like, a few hundred. And then you have, like, a ranking stage, and sometimes, like, you have multiple rankers, which essentially bring that further down and then give you, like, the final list of whatever, like, you know, your top songs that we want to recommend to you.
Uh, and we use this, like, across different products. So we have, like, home shelf ranking, and we have personalized playlists and search and podcasts and, like, ads and, like, a lot of other stuff. And this is generally every team, like, every product has its own team that has its own model.
So it's kind of, like, spread across different groups of people. And some models are better than others. Some have different features. So we're kind of moving away from that sort of siloed model of working towards this single unified model, uh, which is which supports, uh, similar to how LLMs work, which supports, like, an LLM backbone, and which allows you to kind of, uh, allows you, the user, to kind of steer it towards the sort of recommendations that you want.
Um, and one of the key components of that, so, uh, is, is the user modeling part. So that is, like, the team that I work that I work with. And what we do is we build user embeddings, uh, which are essentially representations of vectors that, that tell Spotify about the user's taste across, like, all of the sort of history that we have on, on you, the user, like, across the different sort of, uh, sessions that you've had with us over the years.
Um, and that becomes, like, the foundation of all of the models that are downstream, uh, that, that actually, um, recommend stuff or allow you to search for stuff. Um, and these models are, like, very complex. So they, they the embedding model is we generate embeddings for, like, a billion-plus users because we have, we have a lot of users in general, uh, that are also, like, MAUs or that are not MAUs.
Um, so we do that every day. So it's, it's like a massive, uh, pipeline. It's very expensive. Um, and over the years, like, we've kind of moved away from having these sort of generalized user representations, which was, like, the predominant paradigm in machine learning, where you had these, uh, these models.
Like, this in this case, like, this was a paper from our team last year, uh, where we, we kind of publicly spoke about the user embedding model that we have, which is generally, um, like, at that time, it was, like, an autoencoder model.
If you're familiar with that, what that does is it kind of takes all of your features, it compresses it down to, like, a small vector, and then recreates your features from that. And that sort of compression, decompression process allows the model to sort of learn about you, the user, and just, um, represent, represent user interactions in the form of a vector.
So this is fairly standard stuff, like, that, that a lot of, like, folks do in the NLP, computer vision space as well. That is also kind of aligned with the way things work and recommendations. Uh, we're kind of moving away from that towards, uh, the foundation modeling side.
So now we kind of have this single sequential model, which, as, as you can see, like, with the, the so the whole industry moving towards transformers and the whole shift that's happening across, uh, not just, like, the regular tech industry, but also the, the companies that, uh, that use recommender systems and build recommender systems is, is sort of their main, uh, bread and butter.
Um, so we're also a part of that, sort of one of those companies. And we, we we're also moving towards using transformers for, for the for that kind of stuff. And the idea is that you want to have, uh, the, the user's interactions as a part of the prompt.
So it's, it's kind of, like, the context. Like, I guess when we talk about context engineering, this is the context that we're talking about. Um, and then there's obviously the request con-context. There is, like, the query. There's the product surface, etc., like, all of that stuff.
And then there's the item that we that you're recommending. So when you add all of this stuff, and then you put, like, a transformer layer and, like, a bunch of heads and all that stuff, and you, you train it over, like, millions or hundreds of millions of users' data, what you get is something really cool.
Um, what you get is something like this. Uh, so this is an image which is from one of our, like, newer models, which, uh, what it what it essentially shows you, this is, like, a compressed version of what the model is learning.
Um, in this case, we have tracks in the blue, and we have, like, episodes, which are, like, podcast episodes in pink. And then users, so that, that the one in green is me, and some of the other green ones are, like, other folks in my team.
Um, so that kind of shows you that we're kind of doing this cross-content modeling, where embedding users, tracks, and episodes in the same sort of content space, or, or I guess embedding space. Um, and you're able to kind of visualize, like, I guess from, from this image of, like, how on the hypersphere where you live alongside, like, different pieces of content, and what is close to you, what is not close to you, and how can you kind of explore that, that neighborhood region, and how you kind of, uh, where you live, you know, contextualized by where your friends live and things like that.
So this is sort of, like, a visualization of, like, what the model is learning. And as you can see, for me, like, I'm, I'm a machine learning engineer. I care a lot about keeping up with, with what Anthropic is doing and what's happening in the tech industry and all that stuff.
So my specific embedding is really close to this big, big tech podcast. And on theright, you can see that, like, that the, the point where you see, like, those lines spreading, that is me. And then the ones in pink are, like, streams for tracks, and the, the blue ones are streams for episodes.
Uh, and or sorry, it's, it's the reverse. Um, and you can kind of visualize, like, you can contextualize whatever you're listening, whatever you're, like, not listening to, and how, how, how does the embedding space look like for users.
So these models are, like, really smart. Um, and the moment you give them, uh, information about the user, uh, they can, they can kind of learn to put everything together in, in, in, in a single space, which is really cool.
Semantic IDs11:32
Um, this next part is more about catalog understanding. So now that we have the users, we, we understand the users. We have a model for them. How do we understand the catalog,right? Uh, so catalog understanding generally, uh, there's, like, a number of ways that you can do you can understand the catalog.
Like, the most common way is you train, like, similar to, to the user side, like, you, you train a vector to understand the content. So you have a vector that represents the item, uh, whether that's a song, it's an artist, or it's a podcast or an episode.
Um, and you have that vector. Um, and then alongside that, you have the user vector. So that is, like, the Spotify knowledge. So that's what we know about the, the content or the sort of the different entities that we're dealing with.
Um, and then you have, uh, the world knowledge. So that's coming not from Spotify, but it's coming from these open-weight LLMs that we're working with. Um, so models like Llama or Qwen or, like, other sort of open-source models.
Um, what we do is we fine-tune those models, and then we kind of embed Spotify's knowledge into those models through, uh, something that I'm going to talk about. And, uh, what that does is it gives you steerability. It gives you better re-recommendations.
It gives you explainability. So there's a lot of stuff that you get for free when you use language models for recommendations. Um, now that there are, like, trade-offs here, so that the model does end up, like, forgetting stuff.
Like, catastrophic forgetting is an issue. Um, but generally, from what we've seen, like, these models are really good at combining world knowledge with whatever sort of the knowledge that you have from your platform and building something which you can kind of holistically use for, for recommendations.
Um, now, the stuff that I was talking about where on the previous slide, um, how do we actually teach these LLMs about the content? Um, let's start with that. So the way that we do that, uh, is using something called Semantic IDs.
Um, Semantic IDs is, like, a fairly new concept. It, it I think there's, there was a paper from, uh, Google, uh, a few years ago, which sort of introduced this concept in the context of YouTube. Um, and what it does is, like, when you have a vector that represents a piece of content, so that's, like, a track or an episode in our case, what we do is we tokenize it, similar to how LLMs, like, kind of we, we tokenize words.
Um, and what that does is it kind of compresses that massive, like, let's say, 1,000-dimensional vector into, like, four or six tokens. And that allows us to, uh, really use those tokens to train the, the LLM in the way that, like, LLMs are usually trained.
And it allows the LLM to kind of autoregressively generate the next token or in, in this case, the next token is not a word, but it's the next song or it's the next episode that you're going to be listening to.
Um, so that's, that's kind of what we're doing, uh, where post-training or where, I guess, we're, like, uh, continually training these LLMs with Spotify's sort of data that we have about the catalog. Uh, we use Semantic IDs to compress the vec like, the vectors into Semantic oh, sorry, the vectors into Semantic IDs.
And at the bottom, you can see that we have examples of, like, Ariana Grande and Bruno Mars. So, uh, we represent them as six tokens. So those numbers are actually, like, token IDs. Um, and the first two tokens for both of them are shared because they're both, like, pop artists, and they're both, like they both share something, uh, between them.
But then the other tokens are different because those tokens sort of represent, like, more niches. So it's kind of like a hierarchical, uh, structure where you're, you're kind of compressing the embedding into these six tokens, and there's, there's a hierarchy to it.
And that allows the model to autoregressively generate, like, the next artist or the next song that you're going to be listening to. Um, so this kind of just overall, like, this kind of emphasizes how we do this. Uh, we kind of we use the user context.
In this case, we have, like, a user who's Italian, uh, their listening history, which is tokenized. Uh, we send that listening history. Uh, we use that in the training data. So we, we kind of teach the LLM how to talk with Semantic IDs.
And, uh, that is, like, the domain adaptation that I was referring to earlier in the slides. And then the final output is essentially generating, uh, the next item. So whether that's going to be an episode or it's going to be a track, like, what is this person, like, listening to?
Um, this is sort of, like, one e-example, like, taking the example of the Italian person who maybe listens to an episode, like, an Italian podcast. Um, this is an example of a prompt that we give that model. And you can see, like, the prompt on the left has the Spotify URI, which is, like, our representation of the of, of the, uh, the item, in this case, the episode.
We convert that into a Semantic ID, which is, like, the actual tokens that the, the model is, um, attending to. And that is used to finally predict what is the next episode that the mo that the user is going to be listening to.
Soft Tokenization16:14
Um, so this is the catalog understanding part of it. So now that we have the user modeling part, we have the catalog understanding part, the next step is to essentially assemble all, all of these components, uh, to form, like, a single steerable, personalized, generative recommender system.
So this is sort of moving away from the traditional rec-recommender system model to, to this generative model. Um, this is the sort of the product that I was talking about that we launched, uh, just a few weeks ago.
This is called the Taste Profile. The idea is that you have some piece of text that represents what the who the user is. Um, we, we expose it to you, the user, and then you're allowed to kind of tell Spotify by chatting or by kind of sending messages or, I guess, adding some test text.
So maybe you want to start listening to, like, Justin Bieber more, or you, you don't like this specific podcast that's being recommended to you. And what this does is it allows the model to assen like, that, that data, that edit or is going to come back into the generative model, and it's going to up-upgrade, like, the model's kind of ability to, to sort of understand you and recommend stuff that's, that's better for you.
Um, so essentially, like, we have the content piece of the puzzle, but we don't have the user piece yet. Uh, and the user piece doesn't come because ultimately these models are trained on, like, a limited amount of training data.
You cannot train them on every, like, 750 million-plus users that we have. Um, so there is going to be some sort of level of, like, collaborative filtering. So the, the, the model is going to kind of generalize, hopefully, but it also needs to be personalized.
Um, the way that we do that is, again, coming back to the idea of user models. Um, so, uh, you have the LLM. Uh, you have the user representation. What you do is you project the user representation into the space of the LLM.
Uh, and what that does is it kind of allows you to create what's called a soft token within the model. And it's essentially a token that represents the user, which is obviously contextually changed depending on, on the user that, that we're kind of generating this response for.
Um, and that kind of allows the model to be personalized. So that is kind of, like, the final piece of the puzzle where you have this vector projection, which is, like, you have the regular LLM, and then you have a user vector, which is projected to the space of the LLM.
And that gets kind of inserted into the prompt. And then finally, when the model is actually, like, generating a recommendation, um, the model is personalized because it kind of has the context on you, like, whoever, uh, we're kind of generating the, the recommendation for.
Um, these are kind of, like, some early results that we've had. Like, we've seen some pretty positive results on our internal metrics. Um, if you use Spotify, like, if you use the next episode sort of if you listen to, like, podcasts on Spotify, like, this is something which is actually productionized now.
So you if you're getting a recommendation, it's, it's coming from, from something like this. Um, and that's, that's kind of, like, the final piece. So, uh, we have the embeddings, which represent the users. We have the Semantic IDs, which represent, uh, like, a compressed version of the content.
And we have this soft tokenization approach, which allows you to project, like, users into the token space of the model. And this kind of moves away from the traditional recommender system model, and it kind of moves towards this sort of sequential modeling, uh, framework, uh, which is something that we're kind of very excited about.
And, uh, that's yeah, I think that's, that's definitely, uh, going to be, like, really exciting as we go forward and we build this more, more and more into all of the recommender systems that we have. Uh, and we're very excited to kind of put this out there, uh, pretty soon.
Outro19:31
And yeah, uh, that's, that's my time. I'm, I'm really happy to connect or, like, if you have any questions, feel free to reach out after the talk.





