Blog · Build log · October 1, 2026 · James Forrester
How does Spotify's algorithm work, and could Jev help an artist?
Spotify's algorithm works by turning every listener and every track into a point in a shared space and finding the points that sit nearest each other. It then recommends across that small gap, with human editors and the listener's own actions shaping the final result. Spotify has published enough about this in engineering posts, research papers and its help pages to explain it to artists without guessing. Jev is a model from TypeSafe that answers typed questions with a probability and a confidence. It could help an artist decide which nearby group of listeners a new release should reach first. We plan to test that with independent artist Amy Hef, and whether Jev's confidence deserves trust there is the part we have not tested yet.
This post sets out what Spotify says about its recommendations and what Spotify for Artists shows an artist. It then covers the difference between a vote count and a confidence that is claimed to be calibrated, and where our test stands. It is part of the threndle.ai build log, and it follows our post on Jev, which promised this test.
Key takeaways
- Spotify says its recommendations draw on what a listener searches, plays, skips and saves. It also uses general details about the listener, wider trends and content attributes, blended with human editorial judgment.
- Listeners and tracks are represented as vectors, and Spotify's Voyager library finds their nearest neighbours. Its README says it is "queried hundreds of millions of times per day".
- An artist cannot see that space directly. Audience segments, source of streams and Fans Also Like in Spotify for Artists are the closest outside view.
- A five-neighbour vote with equal weights can only return 0, 0.2, 0.4, 0.6, 0.8 or 1, and nobody claims those shares are calibrated. TypeSafe claims Jev's confidence is calibrated, and that claim can be tested against real outcomes.
- We have used Jev since September 26, 2026, including on a 30-claim audit check. The Spotify test with Amy Hef has not run, and this post reports no result from it.
How does Spotify decide what to recommend?
Spotify decides what to recommend from four kinds of input. They are the listener's own actions, general details about the listener, wider trends, and the attributes of the content. Its page on understanding recommendations, dated March 12, 2026, lists actions such as "searching, listening, skipping, or saving to Your Library". It also names "general (non-precise) location, your device, your language, your age", and content attributes such as "genre, release date, podcast category".
The clearest public description of the similar-listener idea comes from Oskar Stål, Spotify's vice president of personalization. In a 2021 Newsroom interview, he described two people with overlapping taste. "You have four of the same top artists, but your fifth artists are different." Each person's fifth artist becomes a reasonable recommendation for the other. He added that a single recommendation rests on "thousands of inputs, all laddering up to one song recommendation".
Human editors also shape the result. In a 2023 Newsroom post, Spotify says it combines "human editorial expertise with a multitude of signals and systems", a blend it calls "Algotorial".
Listeners are also getting more direct controls. In January 2026 Spotify expanded Prompted Playlists, a Premium beta that can draw on "your entire listening history from when you first joined Spotify" (Spotify Newsroom). In September 2026 it began testing a Taste Profile, "Spotify's brief interpretation of your taste across music, podcasts, and audiobooks" (Spotify Newsroom). Spotify says the Taste Profile shapes the Home feed only and is limited to Premium listeners aged 18 and over in the US and New Zealand. It adds that "Changes typically appear within a few hours".
What is a nearest neighbour, and what is Voyager?
A nearest neighbour is the stored point that sits closest to the point you are asking about. Voyager is the library Spotify built to find those points quickly across a very large catalogue.
Spotify's Annoy README explains where those points come from: "every user/item can be represented as a vector in f-dimensional space." A vector is a list of numbers, and once every listener and track is a list of numbers, similarity becomes a question of distance. Recommending then becomes a search for whatever sits closest.
In October 2023 Spotify's engineering team, in a post by Peter Sobot, introduced Voyager and compared it with Annoy, Spotify's older open-source library.
| Measure | What Spotify reports for Voyager |
|---|---|
| Speed | "More than 10 times the speed of Annoy (at the same recall)" |
| Memory | "Up to 4 times less memory usage than Annoy" |
| In production | "production traffic since 2022" |
| What it powers | "Discover Weekly, Home, and countless others" |
| Query volume | "queried hundreds of millions of times per day" |
The first four rows come from the Voyager engineering post, and the last comes from the Voyager README. Voyager is built on HNSW, a graph-based search method whose authors report "logarithmic complexity scaling" (Malkov and Yashunin), so it can stay fast as the catalogue grows.
Spotify's recent research keeps the same shape. A September 2025 paper describes each listener as a "high-dimensional embedding within a stable vector space" used for "nearest-neighbor retrieval". Its audio encoder "learns track embeddings directly from audio features".
What does Spotify for Artists show about an artist's listeners?
Spotify for Artists shows an artist how often listeners return, how they arrived, and which other artists those fans play. None of this exposes Spotify's vectors, but together these views are the closest an artist gets to seeing their own neighbourhood.
The first view is audience segments. Active listeners "intentionally streamed your music in the past 28 days", and programmed listeners "haven't streamed your music from active sources in at least 2 years". Spotify then sorts listeners by how often they played the artist in those 28 days.
The second view is source of streams. Active streams happen "after intentionally seeking it out", while programmed streams happen "because Spotify or another listener selects it for them". The personalized sources include Discover Weekly, Release Radar, Radio, Autoplay, Daily Mix, daylist and AI DJ. A high share of programmed streams has a plain meaning. Spotify's programmed features, or other listeners' playlists, are placing the music in front of people who did not search for it.
The third view is Fans Also Like. Spotify says the list is "determined by looking at your fans' listening habits", and "You can't edit what shows in Fans Also Like." Read beside the Stål quote above, this list is the nearest thing an artist has to a published list of neighbours.
How does a new release reach listeners who have never heard the artist?
Spotify points to its personalized playlists as the main route to new listeners. Its research on new releases shows why a track starts with little to go on.
Followers are the most direct route. Spotify's Release Radar help page says a pitched song will be included "in your followers' Release Radar". Each listener gets "one song per artist per week", and the song can stay "for up to 4 weeks if a listener hasn't heard it". The pitching rules say "Deliver your music at least 7 days before release so our editors have time to listen." They add that "You can only pitch one song at a time" and "Pitching doesn't guarantee playlist placement."
Artists can also opt songs into Discovery Mode, which charges a commission. Spotify says it "does not guarantee it" and that the commission is "deducted from future Spotify statements". It also says "We take note when a listener isn't engaging with a song" (Spotify for Artists). Discovery Mode applies only in listed contexts such as Radio, Autoplay and the Mixes. Spotify states that "All other streams of the same songs in other areas of Spotify are commission-free" (Discovery Mode contexts). The Discovery Mode page does not state a commission percentage.
Spotify's own figures show the scale of its personalized playlists. Spotify reported in June 2025 that Discover Weekly drives "56 million new artist discoveries" each week, with "77% coming from emerging artists". In July 2026 it described Release Radar as "Reaching nearly 9 million listeners every single week". The same update gave listeners controls that narrow it to a genre or focus it "exclusively on new-to-you artists".
New tracks face a cold start: a recommender knows little about a track until listeners react to it. A 2024 Spotify Research post describes centralized exploration, from a RecSys 2023 paper. In its experiment it reports "an increase in the number of listeners by a factor of 10 on the explored content". A May 2024 study of one million users and 282,000 tracks found that "a user's new release taste is distinct from their overall music taste". It also found that "genre and artist popularity are most predictive" of how many people stream a new song. For a small artist, the listeners most likely to play the back catalogue are not automatically the next release's first listeners.
Why does a calibrated confidence matter more than a vote share?
A confidence is only useful if it comes with a promise you can test. A nearest-neighbour vote produces a share with no such promise attached.
The textbook version of this method, as described by scikit-learn, classifies a point by "simple majority vote of the nearest neighbors". Its classifier defaults to five neighbours with equal ("uniform") weights. With those settings, the share of neighbours that agree can only be 0, 0.2, 0.4, 0.6, 0.8 or 1.
Calibration is a separate property. scikit-learn's calibration guide defines it this way: among predictions near 0.8, "approximately 80% actually belong to the positive class". The guide recommends a separate step, CalibratedClassifierCV, when calibrated probabilities are needed. Nobody claims that a raw vote share is calibrated in that sense. That reading is our own, not a quote from the documentation.
Jev is built around that stronger promise. TypeSafe's announcement (September 15, 2026, updated September 28) says "Calibrated: higher confidence means higher accuracy." That is a vendor claim, and the model is still in early access. It is also a claim that can be tested, which is the point of the plan below. Jev reads text and answers through three primitives, Choice, Score and Noul (TypeSafe docs), and "Every question is evaluated in parallel and in isolation." A Choice question can rank up to 255 options, and "A flat shape... means low confidence." A Score question returns "a probability-weighted mean of the level numbers". The docs warn that "Different distributions can produce the same score", so the full distribution is worth keeping.
TypeSafe is also open about the limits. Its known-issues page for Jev 1.13 says "Jev is not a calculator" and that it "reads dates as text". It also warns that "Accuracy falls as the state grows with content unrelated to the decision." Jev also has a 32k-token document limit, so the input has to be short and on topic.
What has been done so far, and what has not?
We have used Jev on our own work since September 26, 2026. The Spotify test has not run, and none of Amy Hef's data has been loaded into Jev.
Our first real use was an accuracy gate for our website audits, run on September 26 through the direct TypeSafe API with the model jev-latest. On a set of 30 audit claims, each paired with its evidence, Jev marked all 30 correctly. It caught all 10 planted errors, raised no false alarms and made no confident wrong calls. Nine of the 30 came back under 0.8 confidence, which our rule sends to a person for review. The caveat matters. The evidence was assembled for Jev in advance, so the gate is only as good as the pairing of each claim with the right evidence. Since September 27 Jev has also reranked the search over our own notes, which our post on giving AI a memory of your business describes.
The Spotify test will be run with Amy Hef on an upcoming release. We are also building Amy Hef's new website, which this post will show once it goes live. Amy Hef's current site was built before our work. The test will use Spotify for Artists exports rather than the Web API. Spotify removed Related Artists, Recommendations and Audio Features for new apps on November 27, 2024. From February and March 2026, Development Mode also requires a Premium account and allows "up to five authorized users". The steps are these:
- Describe candidate listener cohorts using the figures in the Spotify for Artists exports. These include the audience segment and source-of-streams breakdowns, alongside the Fans Also Like list.
- Compute every count, share and date difference in code first, because Jev is not a calculator. Jev sees plain-language descriptions with the numbers already worked out.
- Ask Jev to rank the cohorts, and record each probability and confidence before the release goes out.
- Compare that ranking with the release's real numbers in Spotify for Artists.
If Jev reports high confidence and the real numbers often disagree, that is a finding we will publish. If the confidence holds, that is the more useful finding. It would mean an artist can tell which ranked suggestions to act on and which to question.
Could this work for a business that is not a musician?
The same method could fit any business that has existing customers and has to choose where to look for the next ones. A business's regular customers play the role of an artist's active listeners. The groups around them play the role of nearby cohorts. Examples include customers of a related business, people who bought once and stopped, or a neighbouring town with a few buyers already.
The owner faces the same question as the artist. Which group is most likely to respond to the next offer, and how sure can the owner be before spending money to reach it? A ranked shortlist with a confidence on each line that has been checked against past outcomes would let the owner act on the confident calls. The uncertain ones would go to a person, which is the same rule our audit gate already uses. Amy Hef's release comes first because a release produces measurable numbers within weeks, which makes the confidence easy to check.
If you want to see where AI could take manual work off your business, start with the 3-minute diagnostic. The Spotify results will appear in the threndle.ai build log once the release has real numbers to compare against. The ranking is easy to produce. The confidence is what has to earn trust.