Research · Part II
Announcing UBQT: The First Frontier Model for Audience Forecasting and Simulation
The obvious audience is the wrong one.
ChronicleOctober 202610 min read
Reach share of channels that delivered
Affinity how strongly they responded
Figure 1: Reach and affinity, UBQT-alpha vs. Gemini 3.8 Flash, GPT-6 Astra and Claude Fable 5.1. Reach: share of recommended channels where the video actively delivered (150 or more impressions). Affinity: how strongly the reached audiences responded, from retention, watch rate and engagement.
Inputs
Model
Predictions
How far content travels, from broad to niche
How strongly those audiences will engage
Whether audiences take the next step
Rolling out to exclusive partners through the Ubiquity Platform and Ubiquity Labs.
Today we are announcing our first frontier model, UBQT-alpha, rolling out to exclusive partners through our Ubiquity Platform and Ubiquity Labs program. Built on real public and proprietary audience data, UBQT is the first model to accurately predict how audiences will engage with any brand on social, from product concepts to trailers to podcasts and beyond.
Audience simulation solves one of the biggest problems for brands and studios today: who is my audience, how can I reach them, and what do they want next? For brands, this means gauging demand for a product before launching into full production. For studios this means optimizing creative and marketing strategy for specific audience segments instead of going in blind with one-size-fits-all trailers.
There are hundreds of different use cases our partners have identified, all supported by our core audience simulation models.
But the question we asked ourselves as we saw success across many partnerships and applications was: why can’t LLMs do this yet?
We looked at this from two perspectives. First, in our prior post, “Why AI Still Can’t Predict Audience Taste”, we tested the general effectiveness of semantic similarity as it is used by the LLMs. Does matching of content with similar content predict audience affinity? We found it didn’t work at all.
We believed Chronicle’s model was much better. So, we benchmarked the performance of our model on a real (and critical) industry problem: predicting audience affinity for a specific product video. Audience affinity is a standard metric that captures the way audiences interact with content on social, including both click-through rate and view duration. Audiences that are high affinity click, watch and engage with the content. Low affinity audiences might click and immediately swipe away or not even click at all. We compared our models to three leading AI models (Claude, GPT, Gemini) to see which was best at discovering high affinity audiences and predicting which ones would actually become fans.
As an example, for a baking video, the three general AI models chose the same three mainstream cooking channels on YouTube as high affinity audiences to test. Obvious choices, on the surface. We promoted the video on each of them. Watch rates ranged between 1.6% and 3.4%, with not one viewer watching through to the end.
For the same video, Chronicle’s model, by contrast, recommended a mystery/conspiracy channel with no apparent connection to baking. 40% of the audience chose to watch it.
That was just the beginning. We served the content to 480 recommended audiences, generating nearly 50,000 impressions. Chronicle didn’t just find more viewers; it found better ones. 85% of its recommended audiences actually reached people at scale, it generated more than twice as many views and engaged viewers 83% more often than the general AI models. Meanwhile the three general models recommended 18 channels that either did not exist or had fewer than 1,000 subscribers.
The test
Every system got the same brief: here is a piece of content, find the 10 audiences most likely to watch it. The general models answered with the YouTube channels whose audiences they expected to respond. Chronicle AI predicted micro-targeted audience cohorts from its model of observed viewer behavior. Each piece of content was then served to every recommended audience, on equal terms: same budget, same format.
- The content: 12 pieces spanning pop-culture talk, science, horror audio drama, indie film, food, retro kids’ animation, adult animation, tabletop gaming and traditional craftsmanship.
- The measure: affinity, a proprietary metric of how well an audience responds to a piece of content. It takes into account watch rates, retention and engagements.
- Requirements: an audience needed at least 150 impressions for a statistically meaningful result. Predictions that never reached that goal score zero, because a recommendation that reaches no one is worth nothing to a content owner. In cases where models recommend a channel that doesn’t exist, the score is also zero.
The general models were Claude Fable 5.1 and GPT 6 Astra, both at extra-high (xhigh) reasoning effort, and Gemini 3.8 Flash at high reasoning effort. The tests ran September 25 to 28.
UBQT is a multimodal model that encodes each piece of content from a rich set of features, among them thumbnails, frames, audio, and generated text summaries. Each audience gets its own representation, built from observed watch behavior and clustered into cohorts of viewers who respond alike. The model is trained to predict whether a given audience will respond to a given piece of content, supervised by real exposures across thousands of audiences. Several rounds of human-in-the-loop refinement then recalibrate the model.
Watch rate 2.5x the best general model
Engagement rate 1.7x the best general model
Figure I: Average watch rate across all 10 recommended audiences per piece of content; audiences that reached too few viewers to count score zero. Engagement rate is the share of people shown the content who interacted with it.
ChronicleGeneral AI models (Claude, GPT, Gemini)Chronicle vs. general AI
- Retro kids’ animation6.2%36.3%5.9x
- Food: barbecue5.4%29.4%5.4x
- Food: baking3.0%16.1%5.4x
- Adult animated comedy9.8%38.3%3.9x
- Indie film15.2%24.2%1.6x
- Horror audio drama17.5%27.6%1.6x
- Science & nature27.6%37.3%1.4x
Figure II: Watch rate by genre, Chronicle vs. the three general models’ picks combined. Here, watch rate is measured among everyone shown the content: the share who watched 30+ seconds or interacted. Scale 0 to 40%.
Results at a glance
Chronicle’s audiences watched more and engaged more than those of every general model, by a wide margin.
The gap is widest where it matters most
The general models did best on broad, well-documented genres, where the obvious audience and the real one overlap. In niche genres, where most new IP lives, their picks collapsed and Chronicle’s held up.
Case studies: the obvious audience vs. the real one
An adult animated comedy
Asked who would watch an adult animated comedy series, the general models went straight for the obvious: animation channels like South Park Studios and Adult Swim. Claude also recommended a Family Guy channel that doesn’t exist. The obvious audiences barely noticed. On South Park Studios, 1 view in 152 impressions. On Adult Swim, 2 in 170.
Chronicle’s model pointed instead to comedy-podcast and pop-culture audiences, including Stavvy’s World, the podcast of comedian Stavros Halkias, and Astraway. There, the same content landed:
- Stavvy’s World: 62.5% chose to watch, 81% engaged, and 30% watched to the very end.
- Astraway: 58.1% chose to watch, 79% engaged, and 33% watched to the very end.
Nearly 1 in 3 people shown the content on these channels watched every second of it.
A baking video
For a baking video, all three general models made the same choice: the biggest baking and recipe channels on YouTube, including Preppy Kitchen, Natasha’s Kitchen and Food Network. Watch rates on those channels ranged from 1.6% to 3.4%, and not a single viewer watched to the end.
Chronicle’s picks included The Why Files, a mystery and conspiracy channel with no obvious link to baking. There, 40% chose to watch. Across the whole campaign, 12.5% of people shown the video through Chronicle’s audiences watched it to the end, against 1.3% through the general models’: nearly 10 times as many.
A science short
For a science video, GPT recommended two nature and wildlife channels, The Dodo and Brave Wilderness. They were shown the content 342 times and produced zero views.
Chronicle placed the same video on Nightcap, a sports and culture commentary show. Nearly half the people shown it watched past the halfway mark, and 35% watched it all the way through.
That is the pattern our first study predicted. The obvious audience for a piece of content is the one whose content looks most like it. The real audience is the one whose taste responds to it, and the two are often not the same.
- Chronicle~3,100
- Gemini~1,800
- GPT~1,500
- Claude~1,400
Full watches = impressions × share of impressions played to 100%, rounded to the nearest hundred.
Reached enough viewers to countToo few viewers to countChannel didn’t exist or under 1,000 subscribers
- Chronicle85%
- GPT51%
- Gemini50%3 fake or near-empty
- Claude44%12 fake or near-empty
Figure III: Every recommended audience, by outcome. “Too few viewers to count” means under 150 impressions, not enough for a statistically meaningful result.
Finding the core fandom, not just the crowd
A view is a start. What content owners actually want is an audience that cares: people who watch, interact and want more. On every measure we tracked, Chronicle’s audiences were stronger.
They engaged almost twice as often. Chronicle’s audiences engaged with the content 46% of the times they were shown it, against about 25% for the general models’ audiences. Engaged viewers are the seed of a fanbase: the people most likely to comment, share and return for the next release.
More of them watched to the end. Chronicle’s audiences produced about 3,100 full watches, people who stayed with the content to the very last second, against about 1,800 for the best general model. Every model recommended the same number of audiences, so this is a like-for-like comparison, and finishing a video is one of the clearest signals of real interest there is.
They were real, active communities. 85% of Chronicle’s recommended audiences were large and active enough to reach viewers. Fewer than half of the general models’ picks were. The general models also recommended 18 channels that either did not exist or had fewer than 1,000 subscribers. That is worse than a weak pick. Every one of those recommendations is budget spent on an audience that isn’t there, and it’s a sign the model is producing plausible-sounding names rather than drawing on real audiences. A content owner relying on it would need to check every channel by hand before spending a dollar.
They delivered more attention for the money. Chronicle’s audiences generated 8,397 views, more than double the 4,010 from the general models’ picks, at a 13% lower cost per engagement. Budget went to people who wanted the content rather than people who happened to be near it.
Why general AI misses the audience
These results follow directly from what our first study found. General models are built to understand content, and they reason about audiences through that lens. Three things go wrong.
Resemblance is not affinity. Asked who will watch, a general model looks for channels whose content resembles the source. Our 15,000 tests showed that resemblance barely predicts response: the correlation between semantic similarity and view rate was close to zero. Choosing the most similar audience means choosing on the wrong signal.
Taste lives in combinations. Audiences respond to specific mixes of topic, tone, format, pacing and host. The same topic can win over one audience in the right style and lose another that loves the topic but not the format. A single similarity score cannot capture those interactions. In our first study, a linear model could not predict response at all, while a model that learned nonlinear combinations explained more than half of the variation in response.
A channel’s content does not describe its audience. Two channels can publish near-identical videos and attract viewers with different habits, expectations and reasons for watching. That difference only shows up in how people respond. Chronicle models that difference directly, learning a representation of each audience from how it actually behaves and pairing it with a rich description of the content.
A general model has neither the audience representation nor the observed response. It can only offer a plausible guess, which is also why some of its recommended channels turned out not to exist.
What this means for content owners
Most creators and IP owners still release new content blind, and wait for a platform algorithm to decide who sees it. General AI does not fix that. Asked who will watch, it answers with who should watch: the biggest, most on-topic audiences it can recall.
The cost of that guess is real. More than half of the general models’ recommended audiences never meaningfully reached viewers, and even the ones that did held attention far less often. 18 of their picks were either too small to register or didn’t exist at all. In niche genres, where most new IP lives, the gap in watch rates between Chronicle and the LLMs’ recommendations grew to as much as 6x.
The general models all reach for the same names. You’d expect three different models to give three different answers. They didn’t. Of the 10 channels each model picked per piece of content, all three agreed on at least 3 on average. And the picks were the most visible names in each genre: Oprah and Mel Robbins for a celebrity interview, Food Network for a baking video, Adult Swim for an adult animated comedy.
Consensus didn’t help. The channels all three agreed on were watched 19% of the time, no better than picks only one model made (18%). For the celebrity interview, Oprah, Mel Robbins, and Tamron Hall delivered zero views.
General-purpose LLMs aren’t built for this. They pick whatever is most talked about, so every brand that asks gets the same crowded list. Trends fade. Fandoms stay. UBQT-alpha is built to find the fandom.
LLMs generally surfaced the most obvious audiences, because their models are just trained to recognize content. Chronicle found the most responsive audiences, by starting from the other direction. Instead of asking which content to show a viewer, it asks which of thousands of niche audiences will respond to a specific piece of content, and where a larger fanbase may be waiting. The result is an audience map grounded in observed response, which content owners can use to decide where to launch, which niches to test and how to grow.
Content production is exploding. The question is shifting from can AI understand my content? to does AI know who wants it? On the evidence of this test, general models don’t. And the reason they don’t is the reason we built Chronicle’s audience discovery model the way we did. Because audiences are not obvious.
