The Hallucinated Empathy Problem
An AI model can’t be your (synthetic) audience. But it can analyze your audience.
Overview
The rapid rise of synthetic audiences in strategic communications and marketing, e.g. use of AI to understand how audiences might react to a campaign instead of using focus groups, has as many traps as it does opportunities. Strategists need to understand where these traps lie, where the science proves rigor, and how to identify the solutions which create opportunities instead of liabilities.
There are three simple questions a sophisticated buyer can ask to separate science from fiction. But before we get to that, let’s break down the promise, what the science says, and the hallmarks of a rigorous solution.
Today’s Synthetic Audience Solutions
The use of synthetic audiences often looks like this: ask a chatbot to pretend to be your audience and sidestep the rigorous testing of messaging before you go to market. The pitch is becoming irresistible. Open ChatGPT, prompt it to act as “a 45-year-old suburban mother.” Then “a Gen Z environmental activist.” Then “a conservative small-business owner.” Ask what each thinks of your campaign. Test a dozen - or a hundred, thousand - synthetic personas in an afternoon while skipping the cost and time of focus groups.
This is what some companies mean when they say they’re “using synthetic audiences.” The impersonation prompt is followed by a confident-sounding answer and the appearance of rigorous testing. The gap between what this shortcut appears to do and what it actually does is now well-documented enough that continuing to rely on it is no longer a benign oversight to be blamed on the nascent nature of the field. It’s neglect at best, malpractice at worst.
To be clear, synthetic audience research is a real and established discipline. The serious companies doing it build their analyses on real demographics, psychographics, and behavioral data rather than on a model’s impersonations. By some estimates, the synthetic data generation market is projected to clear $3.2 billion by 2032, and the rigorous players in that market produce solid work strategists can act on. The problem isn’t the “synthetic” part, but rather the companies and solutions which aren’t rigorous, and the ability for buyers to effectively evaluate what’s rigor versus what’s smoke-and-mirrors.
LLMs are extraordinary at one thing: analyzing how language encodes belief, bias, and identity. They are bad at embodying any of those things. Asking a language model to pretend to be a person and treating the answer as if a real person had given it is hallucinated empathy.
What we’re actually doing when we prompt a model to “be” an audience
Ask any major LLM to impersonate a member of a demographic and then watch what comes back. The AI response vocabulary shifts, the tone adjusts, and a few cultural markers appear: a reference to a job, a region, a family situation, etc.. The output reads like a highly believable audience response.
But it isn’t. What you’ve retrieved is the model’s statistical archetype of that demographic, compressed from the same training data that produced everything else the model says. The model is not channeling a 45-year-old suburban mother from its data. It is generating the median utterance the model associates with that label, drawn from the cultural and linguistic distributions of its entire training corpus. It’s not only including the data it has from that audience, but all the data from what others believe and have said about that audience, including the stereotypes, biases, and racist tendencies which infest the internet.
The results falsely read like empathy.
This approach works well enough for narrow, mechanical tasks like UI flow testing, agent-to-agent interaction design, or transactional UX research. It fails entirely for strategic messaging, narrative testing, and anything that depends on belief, identity, or cultural reasoning, which is really to say, the work that actually matters in strategic communications.
The empirical case keeps getting harder to argue with

There’s a growing body of research which makes clear that asking an LLM to perform identity cannot do what strategists want it to do.
The first issue is distribution collapse. Shumailov and colleagues, in a 2024 Nature paper, found that when generative models train on their own outputs, the tails of the original semantic distribution vanish first. The model converges toward central tendencies: the most common and culturally dominant patterns. The rare and marginal voices, the cultural edges, the surprising perspectives, all the elements that make audience research valuable, are the first to disappear. With less and less novel training data available - after all, the models are basically trained on the entirety of the internet at this point - it’s unlikely a new, relatively free data source will appear to solve this homogenization at the model level.
The second issue is the expression-cognition gap. Work on modeling human subjectivity in LLMs (Giorgi et al., 2024) found that while models can reproduce surface demographic markers, such as vocabulary, idiom, and references to a region, they fail to capture the underlying interaction between belief, emotion, and context that defines human reasoning. LLMs simulate expression, not cognition. The persona sounds right but reasons wrong. For strategic communications, where the entire point is to predict how an audience will weigh a message against its existing beliefs, this is fatal.
The third issue is stereotype amplification. A 2025 audit of AI-generated personas (Venkit et al.) tested how three major LLMs (GPT-4o, Gemini 1.5 Pro, and DeepSeek 2.5) generated synthetic personas across racial identities. The result was consistent across all models. Each AI model disproportionately foregrounded racial markers, overproduced culturally coded language, and produced personas that were, in the authors’ words, “syntactically elaborate yet narratively reductive.” Worse, the reductive results become masked by a layer of positive affect that makes the stereotype harder to detect. The persona effectively read as a caricature. With later models the analytical capabilities advance, but the underlying biases persist.
A fourth finding deserves a brief mention. When LLMs are deployed as personified agents, such as autonomous pricing agents in a market simulation, they exhibit emergent behaviors no human would. Fish, Gonczarowski, and Shorrer (2024) found that LLM-based pricing agents quickly arrive at supracompetitive prices without being instructed to collude, behaving in ways that match no realistic human merchant. The point isn’t that LLMs are malicious, rather they don’t act like humans even when we want them to. Asking a hundred, a thousand, or ten thousand AI personas together to interact in a simulation (such as mimicking social media communities) compounds the issue rather than relieving it. Treating them as audience proxies inherits the same disconnect.
The ethics problem nobody’s asking about
The impersonation-prompting shortcut isn’t only methodologically broken. It is ethically fraught, and the industry has been remarkably quiet about it.
Asking a model to “be” a Black suburban mother, a Puerto Rican teenager, a Hasidic small-business owner, a recently naturalized Japanese citizen, or a member of any community the strategist doesn’t belong to is a form of automated identity ventriloquism. The AI produces a flat, fabricated version of that person and offers up its results as if the group itself had spoken. No one in the group consented to being represented this way. No one has a way to push back when the representation is wrong. And the representation is, according to the research, almost always wrong, often in the ways which amplify stereotypes and flatters the user’s assumptions.
If a researcher told an Institutional Review Board “I plan to write detailed personas of community X based on my intuitions, then use those personas to test how to message community X,” the protocol would be denied. We’ve built consent frameworks, community review boards, and disciplinary norms over decades to prevent exactly this kind of impersonation. The easy access to AI chatbots has removed the friction which used to surface this ethical problem.
Persona prompting produces its most confident-sounding outputs for groups best represented in training data: Western, Educated, Industrialized, Rich, Democratic (WEIRD). It produces its least reliable outputs for marginalized, underrepresented, and culturally distinct groups, which are frequently the groups whose representation matters most in public communications, in policy work, in public health, or in any communications work with society-wide consequences. The technology fails worst exactly where the ethical stakes are highest, and it hides the failure under a veneer of fluent, confident, sympathetic-sounding language which makes the failure harder to detect.
These findings describe what happens when an LLM is asked to perform. But there are methodologies which use the model differently sidestep these failures almost entirely. While still a substitute, and therefore inferior to doing the hard work of primary research, it avoids the pitfalls of performative approaches.
A better question for the machine
Generative AI models aren’t the problem in synthetic audience research. The question we’re asking the AI is.
The question, “How would this audience respond to this message?” assumes the model can become the audience. It cannot. There is a different question, one the model can answer with rigor, that is far more useful for strategic communications:
Given what we know about this audience, where does this message conflict with their existing cognitive and cultural norms? This is an oversimplified representation, but the approach is sound.
What changes, practically, when you stop asking the model to be the audience and instead analyze the gaps?
You stop trusting answers that feel like empathy and start trusting analyses that show their work. You stop generating an unending series of impersonations that all sound the same because they all came from the same underlying semantic distribution. You start running rigorous diagnostics against audience profiles built from real data, with the LLM serving as the most powerful language analyst ever built, not the most plausible impersonator.
LLMs are exceptional at this kind of analysis. They are linguistically and conceptually fluent enough to map where a message’s framing collides with an audience’s values, where its tone trips moral foundations, or where its messenger lacks the credibility markers the audience requires.
This is semantic analysis.
With semantic analysis, the model reads the message against a rigorously constructed audience profile and semantic landscape, then surfaces the points of convergence and divergence from the norm, thereby identifying specific points of friction. Where will this framing trigger belief conflict? Where will this language fail moral alignment? Where will this messenger lack the credibility this audience expects?
The approach sidesteps the audience impersonation problem. The audience profile is constructed from real data about the actual audience: surveys, behavioral patterns, demographic and psychographic research against a backdrop of relevant scenarios - semantic landscapes which encode reactions, values, beliefs, interests, experiences, etc. The model’s job is to analyze how a message lands against that profile and the backdrop of relevant scenarios.
The output isn’t “the audience will say X.” It’s “this message diverges from this audience’s known cognitive and cultural norms in comparable situations at these specific points.”
The difference matters more than it sounds. When you ask a model to impersonate the audience, you cannot tell whether the answer is the reflection of audience norms or model biases. Any interesting insight isn’t born of surprising human cognition, but temperature and probability variances in the AI model.
When you ask the model to analyze the message against an audience profile you can verify, the answer is anchored to data you control. The model’s job is the analysis, and there’s no conflation between the audience data and the model data. They stay in their respective lanes, and the difference becomes clear in real-world results.
Put another way, AI models are fantastic at analyzing data fed to it. The underlying context and audience data that serves as the corpus for analysis is most critical – in fact, significantly more-so than the actual AI models used to analyze it.
What semantic analysis actually measures
At Lancea, we measure against multiple dimensions which contribute to the acceptance or rejection of a message. These are techniques proven in real-world use, yet are incredibly intuitive.
Belief conflict. Where does the message’s framing contradict the values and worldview the audience expresses? A claim about around efficiency lands differently with an audience who values equity. The analysis identifies the specific message framings that will trigger resistance, with reasoning the strategist can interrogate.
Language alignment. Tone, register, moral foundations, and symbolism can each fail independently of the others. A message can be linguistically polished and yet structurally misaligned. The analysis surfaces the gap between how the messenger speaks and how the audience speaks.
Psychological resistance. Some messages trigger natural reactions, such as the audience’s instinct to reject a message that feels like manipulation, condescension, or threat. The analysis maps where prior narrative patterns suggest a message will hit resistance before the audience ever groks the substance.
Credibility and trust. Messages travel through messengers. A claim that lands as authoritative from one source lands as suspicious from another. The analysis diagnoses where messenger attributes (expertise, authenticity, proximity, and more) fail to match the trust requirements of the specific audience.
Cultural context. Idiom, identity-coded expression and rituals carry meaning beyond their literal content. The analysis flags where cross-cultural resonance is likely to fail, including the blind spots many teams aren’t aware exist.
Importantly, we don’t ask the models to predict each of these based on training data. Instead, we provide near-real-time context, dynamically selected scenarios, language, and audience data aligned with the campaign being tested. Then we model tens of thousands of simulations in a proven, statistically rigorous manner across hundreds of ensembles to measure from a variety of angles. The role of the AI is to evaluate the disparity between the base priors and the new inputs, e.g. a message, campaign, or asset. The result is a system which not only tells you what happened, but tells you whether you can trust the results. Showing the variance and probabilities allows the user to evaluate where the models are confident and where they’re not.

None of these outputs is a role-played reaction. Each is a diagnostic - a measurable gap between the message you’ve written and the audience you’re trying to reach, expressed with the kind of statistical confidence you can transparently evaluate.
Choosing your synthetic audience partner
The next generation of communications intelligence belongs to methodologies, and to the companies, that build on what LLMs are actually good at, not what they (quite convincingly) pretend to be good at. Strategists evaluating synthetic audience tools should ask a simple question of any vendor or methodology: are you asking the model to be the audience, or to analyze the audience?
Vendors who answer the first are selling a shortcut.
Three clarifying questions can help communications professionals test a vendor’s solution quickly.
First: what is the source of the audience profile the analysis runs against? If the profile itself is generated by the model, you’re sampling model biases, not audience biases.
Second: what does the output look like? If it reads like a transcript of what a synthetic audience supposedly said, the model is performing. If it reads like a diagnostic map of where the message and the audience diverge, with reasoning you can interrogate, the model is analyzing.
Third: what statistics exist showing the confidence of the models on your specific simulation? If the model can’t tell you if it lacks confidence, you’re buying convenient confirmation bias, not rigor.
The use of synthetic audiences stands to become a monumental leap forward for communications strategists who need to rapidly test and deploy messaging, but like all things with AI, skipping rigor for convenience can be catastrophic. Those who can identify the difference between performative impersonation and rigorous analysis will reap the advantages when they execute their campaigns. The rest will realize the error when it’s too late.
-Greg Young


