Research · Part I
Why AI Still Can’t Predict Audience Taste
ChronicleOctober 20267 min read
LLMs today are very good at generating content. They are very bad at predicting who will like it. We ran 15,000 tests to work out why, and to see where audiences are hiding.
AI devours content. Feed a big model with text, video, audio or images and it will identify the subject, summarize the main themes, detect the genre and the tone and reliably retrieve other content that resembles it.
What these models cannot do is tell you who might be interested in this content, because they don’t understand how people consume media in the real world and how it resonates emotionally and spreads culturally. LLMs are not built to discern what defines taste and culture.
So today creators and IP owners often have to release their content into the world essentially blindfolded, as they wait for platform algorithms to decide who sees it. It is true that existing recommendation systems are great at predicting what an individual user may watch next. They have years of behavioral data on what that person has watched, what they skipped and what they came back to, which is very useful for boosting their platform’s viewership. But while the algorithms are designed to maximize the time viewers watch familiar, often repetitive material on a specific platform, they in fact hinder rather than help creators of new types of content to find audiences for their work.
Our vision at Chronicle is to help content owners, who need to approach this problem from the opposite direction. Instead of asking “Which content should we show a particular viewer?” we ask “Which of thousands of niche audience segments are most likely to respond to a certain piece of content, and how can I produce better content to reach millions more?”
Those niche cohorts are rarely defined by broad demographics alone. They are smaller segments shaped by distinct tastes, behaviors, and diverse reasons for watching. Understanding which of them has an affinity for a particular piece of IP tells us where to begin, which niches to test, and where a larger fanbase may be waiting. The result we are working toward is an audience map, grounded in observed response, that creators and IP owners can use to plan how they release their content.
To build that map, we first need to understand what actually explains affinity between content and an audience.
Semantic similarity, a measure of how closely two pieces of content resemble one another in meaning or characteristics, would seem an obvious place to start. For example, if someone regularly watches Formula 1 race analysis, it might seem reasonable to expect they would also respond well to a documentary about a Formula 1 driver. That assumption has largely driven content discovery up to now.
Semantic similarity is used extensively in closely-related problems including search, content recommendation and multimodal retrieval. Modern AI has sophisticated tools to measure similarity. Embedding models can represent the meaning of text, images, audio and video in strings of numbers that allow us to compare it with other similar pieces of content.
So armed with these tools, we decided to examine the question: If a piece of content is semantically similar to what an audience already watches, will that audience be more likely to watch it? In other words, does semantic similarity also predict audience affinity?
Our results were surprising.
Measuring audience affinity
Audience affinity can mean many things. We could study watch time, retention, repeat viewing, engagement, or conversion. Each captures a different part of the relationship between a viewer and a piece of content.
For this experiment, we used video view rate (VVR), a standard metric reported by YouTube. It measures the percentage of impressions that resulted in a view, capturing how often an initial exposure developed into more meaningful attention.
We tested 14,520 video–audience placements, using real money in real campaigns with real people. We promoted more than 110 hours of source content against almost 6,000 target channels. In each placement, a source video was promoted against a target channel’s audience, and the resulting VVR was recorded.
We represented both the source content and the content associated with the target audience in four different ways:
- Text: titles, descriptions
- Visual: thumbnails and video representations
- Audio: representations generated from the videos’ audio tracks
- Multimodal (all combined)
For every video–audience pair, we generated embeddings and calculated their cosine similarity, a measure of how closely those two embedded number-lists matched up, from 0 (unrelated) to 1 (nearly identical). If semantic similarity were a strong predictor of audience affinity, we would expect higher similarity scores to correspond with higher view rates.
But that wasn’t what happened.
The Experiment
Using three different measures of correlation, each asking slightly different questions, we found that semantic similarity barely predicted audience response.
Across all 14,520 placements, the results of each of our three measures of correlation were all close to zero. This was the opposite of what we expected.
The weak correlation was surprising. We expected audiences to respond more strongly to content that resembled what they already watched. We were in fact discovering the apparent blind spot in how AI models have approached media. They can identify similarities. They cannot account for taste, which involves combinations of multiple factors.
But our results did not necessarily mean that the embeddings contained no useful information. So we decided to look more deeply at the various things that embedding captures: topic, genre, tone, pacing, and visual style.
After all, an audience might respond to a particular visual style only when it is paired with the right genre, or to a familiar topic only when it is presented in the right tone or format. Some viewers might love a Formula 1 series focusing on the characters of the drivers, but have no interest in technical discussions about engine performance. These relationships would not necessarily appear as a simple undifferentiated increase in view rates. But they are very important to how humans decide what they want to watch.
So we looked at whether a nonlinear model could learn the interactions that we were missing.
Nonlinearity recovered signal
What we discovered was that affinity depended on nonlinear combinations of the characteristics encoded within them, not simply on how close two pieces of content were overall. In other words, a particular characteristic can resonate differently depending on what it is combined with: topic might matter when presented in a certain visual style, a host might appeal to one audience but not another. F1 cars might appeal to one audience when they are filmed from the driver’s POV inside the cockpit, while other audiences might prefer an expert commentator explaining mechanical improvements or race strategies.
Human taste is complicated!
We then explored representations built from text, imagery, audio, and video, as well as different ways of combining them. Across these experiments, AI-generated summaries describing characteristics such as subject matter, themes, tone, and likely audience provided the most reliable signal.
This points to a different way of thinking about the problem. Rather than reducing content to a single measure of similarity, it may be more useful to represent it through richer, descriptive information about its characteristics and learn which combinations of those characteristics matter for a particular audience.
But better content representation may still only be part of the problem. The content associated with an audience does not necessarily describe the audience itself.
Two channels can publish semantically similar content while attracting viewers with different expectations, habits, identities, and levels of responsiveness. One audience may follow a channel because of its host; another may respond to its tone, format, or community. These differences may not be visible in the content published by the channel.
This is where existing content discovery models keep hitting a ceiling. The LLMs are very good at spotting patterns in data. However they are simply not built to discern the role of human experience in dictating how people interact with content and with products. This is a matter of taste, and taste embraces the relationship between content and people in the real world.
Modeling the audience, not just the content
As our testing has shown, each audience has a latent behavioral representation: an audience vector shaped not only by what its members watch, but by how they respond. Affinity emerges from the interaction between two representations:
content representation × latent audience representation → observed response
Chronicle takes both of these factors into account by focusing on the details of content and the lessons of observed human response. The result is a usable model of audience taste.
Content production is exploding in the current era. However almost nobody is focusing on how to tell if audiences want it.
Chronicle is. And we have 15,000 tests that point the way forward.
Because the next advances in content discovery will come not just from better AI understanding of content, but also from AI working out why different people have different responses to it.
Our goal is to bring those two sides together: to create a richer understanding of what a piece of content is, and a learned understanding of how a particular audience responds. Stay tuned as we continue testing that direction!
