Every vendor in the research space is talking about synthetic data. Not all of them mean the same thing. Not even close.
The word "synthetic" is currently being applied to four meaningfully different things, each with different validity profiles, different failure modes, and wildly different levels of commercial readiness. This article gives you actionable ways to distinguish between the terms and technologies and choose what is right for the business need at hand.
What "synthetic" actually covers
The clearest way to think about synthetic research is as a spectrum, from the most granular to the most aggregate.
Synthetic panels and digital twins are the most ambitious version of the idea: one AI model per real consumer, queryable like a survey or a focus group. The promise is a panel you can run against instantly, with no recruitment, no incentives, no logistics. Academically, this is the most valid form, if you could actually build it, it would be the most representative. In practice, the blocker is the data. To model a single person's decision-making accurately, you need rich qualitative data captured in-context, tied to behavior, and continuously updated. That dataset does not exist at the scale required. Current models are also not trained to model human variance, the messiness of real choice-making, which means even well-resourced attempts produce demos that look compelling and production outputs that teams cannot confidently bet on.
Synthetic personas are a step down in ambition: a single aggregate chatbot representing a target audience, queryable through conversation. Simpler to build, simpler to interface with. But the aggregation process is fuzzy, how data from multiple people gets combined into a single persona isn't clearly defined, and the outputs lack statistical significance. You cannot model distributions of consumer preferences from a single persona, which limits what questions it can actually answer.
Synthetic communities sit in between: an aggregate model of a target audience that can produce quantitative outputs and model preference distributions. More rigorous than a persona, less precise than a digital twin. The open questions are about how reliably the underlying data is combined and what "accuracy" actually means when there's no clean ground truth to validate against.
Simulated testing is the most distinct category, and the one most likely to deliver reliable value in the near term. Rather than modeling individual consumers from scratch, simulated testing takes existing research and insights and applies them to new use cases. It combines, quantifies, and extrapolates from what is already known. The key characteristic is that it is auditable: every output traces back to the source material that supports it, so teams can see exactly where the confidence is high and where the evidence runs thin.
Why the distinction matters
The difference between these four approaches is not just technical. It's epistemological.
Digital twins and synthetic panels ask you to trust a model to represent consumers you have never directly spoken to. Simulated testing asks you to trust a structured extrapolation from consumers you have spoken to. Those are fundamentally different asks, with fundamentally different answers to the question "how do I know this is right?"
With simulated testing built on real research data, you can trace an output back to a verbatim. You can see the gap when the evidence does not support the inference. With a synthetic panel or digital twin, the output arrives without that trail, and as anyone who has sat in a research debrief knows, "the AI said so" is not a sufficient answer when a business decision is on the table.
The research industry has known for decades that even traditional methods have this problem. A landmark meta-analysis found that a medium-to-large shift in stated intention produces only a small-to-medium shift in actual behavior (Webb & Sheeran, 2006). Surveys have always asked people to predict behavior in a frictionless context that does not resemble the shelf. Synthetic panels inherit this limitation and add new ones: the model does not know what mood the consumer is in, who they are shopping with, or what the competing brand just did on promotion.
AI demos are compelling. Production is harder. That gap, between what synthetic research looks like in a pitch and what it delivers when teams need to stake decisions on it, is the thing that most vendor conversations skip past.
Where simulated research is useful
When the task is directional and the timeline is short, simulated testing grounded in real consumer data is a strong fit.
Early-stage concept screening is the clearest example. If you have ten product concepts and need to narrow to three before committing a fieldwork budget, you need relative ranking more than precise measurement. Simulated testing can tell you that concept A outperforms B and C on relevance and differentiation with your core segment. It cannot tell you what concept A's market share will be in eighteen months, and it should not try.
Messaging and stimulus iteration is another strong use case. Running copy variants through a simulated framework before live fieldwork is pre-flight testing, you are looking for obvious misses, not final calibration. It compresses a cycle that would otherwise require recruiting, fielding, and waiting.
Hard-to-reach segments also benefit. Recruiting a niche B2B persona is slow and expensive. Simulated testing built on existing research from similar segments can produce directional signals while live recruitment runs in parallel. The key is being transparent about what the evidence base actually is.
Where it falls apart
The failure modes are specific and worth naming.
Synthetic panels and digital twins, in their current state, are not ready for general commercial use on high-stakes research questions. The ground truth problem is real: there is no clean benchmark to validate against, because every traditional method, surveys, focus groups, IDIs, has its own known flaws. That does not mean the approach is wrong. It means the infrastructure to build it reliably does not yet exist.
Simulated testing, meanwhile, is only as good as the research it extrapolates from. If the underlying consumer conversations were narrow, outdated, or poorly executed, the simulated outputs will inherit those problems. Garbage in, garbage out has always been true in research; it is just less visible when AI is doing the extrapolation.
Neither approach works for "why" questions where the answer requires fresh consumer context. Most hard business questions across innovation, brand, messaging, and business intelligence ultimately come down to why: why is this category declining, why did this launch underperform, why is a competitor taking share. The answers to those questions come from consumers who are context-aware decision-makers making real choices. Simulations can help you frame the question and surface hypotheses. They cannot replace the consumer.
Emotionally complex and socially sensitive topics are a clear exclusion for any form of simulated research. So are any final go/no-go decisions where the stakes are high enough that someone in the room will eventually ask how you know. "We simulated it" is not the answer for those moments. "We spoke to five hundred real consumers and here are their voices" is.
What Keplar actually builds
Keplar's approach starts with a premise that most synthetic vendors skip: real consumer voices are not optional. They are the foundation.
The platform runs AI-moderated voice interviews with real consumers at scale, conversations that produce more than ten times the data of a survey open-end, automatically evaluated for quality. That research data then powers simulated testing: evidence-backed concept and message tests that are fully auditable, tied to source verbatim, and clearly marked where the evidence is strong and where it runs thin.
This is not a digital twin product or a synthetic panel. It is a structured way to extend the shelf life of real research, to apply what you already know from consumers to new questions, quickly and transparently, without pretending the simulation is the same thing as talking to people.
The output can be traced. The gaps can be seen. That is the bar that makes simulated testing trustworthy enough to act on in production, not just in a demo.


Want to see Keplar's simulated testing in action? Book a demo with us.
FAQ
What is synthetic research data? Synthetic research data is a broad term covering several distinct approaches: digital twins and synthetic panels (individual AI models meant to represent real consumers), synthetic personas (aggregate chatbots representing a target segment), synthetic communities (aggregate models with quantitative outputs), and simulated testing (extrapolation from existing research). These have meaningfully different validity profiles and readiness levels.
Are synthetic panels and digital twins reliable? The concept is academically sound but practically not yet ready for general commercial use. The main blockers are the absence of the right training data, the difficulty of measuring accuracy without a clean ground truth, and current models' inability to represent human behavioral variance.
What is simulated testing and how is it different? Simulated testing applies existing market research and consumer insights to new questions, quantifying, combining, and extrapolating from what is already known. Unlike synthetic panels, it is fully auditable: every output traces back to the source evidence. It is best used as an early-stage sandbox for concept and message screening, not as a replacement for live research.
When should I not use any form of simulated research? For final launch decisions, precise volume forecasting, "why" questions that require fresh consumer context, emotionally sensitive topics, and any situation where stakeholders need to see the evidence traced back to real people.
How does Keplar's simulated testing work? Keplar grounds simulated testing in real AI-moderated consumer conversations. The simulations extrapolate from actual research data, so every output can be traced to a source verbatim. Gaps in the evidence base are surfaced rather than hidden.

