Synthetic respondents and synthetic data: where they support research and where they can be misleading

Monika

Before commissioning a project based on data generated by a language model, it is worth knowing at what point such a simulation stops reflecting reality and starts producing convincing-sounding fiction. Synthetic respondents in research can speed up questionnaire pretesting and generate initial hypotheses, but in some applications they substitute a statistically probable answer for a genuine one. This article shows where the boundary lies and how to recognize it in research design practice.

What are synthetic respondents, and when is it worth using them at all?

Synthetic respondents are profiles that answer research questions not based on their own experiences, but on patterns learned by a language model from vast text datasets. In research practice, they are known by several names: AI personas in research, simulation agents, or simply synthetic data generated in place of responses from real people. They have one thing in common: the answer does not come from a person who actually bought a product, felt frustrated at a self-checkout, or canceled a subscription.

This distinction determines everything. The model does not report an experience; it predicts what a typical statement from a person with specified characteristics would sound like. As long as this is what is needed – something probable, averaged, and consistent with a pattern – synthetic respondents in research can be useful. The problem begins where the aim of the research is to capture what deviates from the pattern.

The most justified applications are supportive rather than substitutive. They are worth listing because they define the real scope of meaningful use:

  • Questionnaire pretesting – checking whether questions are understandable, whether scales work logically, and whether there are any double-barreled or leading questions.
  • Generating hypotheses before the actual fieldwork – quickly mapping possible motivations that can later be verified with a real sample.
  • Filling data gaps in quantitative datasets using statistical methods (imputation), where the synthesis is based on actual distributions from the study.
  • Training and calibrating analytical tools, where real data are sensitive or unavailable and only material with a realistic structure is needed.

In each of these cases, synthetic data serve as scaffolding rather than the foundation for conclusions. This distinction recurs in every subsequent section because it is at the core of the entire issue.

How can response simulation be used in research design without distorting the results?

The key research design decision is whether the study is intended to describe a known distribution or discover an unknown one. Response simulation may work in the former case and is risky in the latter. If the aim is to test whether respondents understand the wording of a question about purchase frequency, the model will generate a realistic range of responses and reveal where the question “breaks down.” However, if the aim is to find out why loyal customers suddenly leave, synthetic respondents in research will provide the reasons most commonly described in the texts on which the model was trained – rather than the one that no one has yet named.

This is where the deepest methodological problem lies. As Hume’s Institute experts point out, a model can write the answer one would like to hear – but in research, the answer that is not expected is needed. The value of qualitative research comes from anomalies, inconsistencies, and statements that challenge the project team’s assumptions. A language model typically gravitates toward more probable patterns and may smooth out outliers. The resulting material is polished, internally consistent, and often devoid of precisely what qualitative research is conducted to uncover.

This is why response simulation should be incorporated into the research process in specific, limited roles. In Hume’s Institute projects, synthetic data provide the greatest benefits at the instrument development stage and the fewest benefits when drawing conclusions about motivations. The practical sequence of steps is as follows:

  1. Define whether the research question concerns a known structure or the discovery of the unknown – this determines whether simulation is permissible at all.
  2. Use synthetic data for pretesting and refining the questionnaire before launching fieldwork with a real sample.
  3. Treat generated hypotheses as material to be verified, never as findings.
  4. Compare every synthetic result with a portion of real data – even if smaller, it must be genuine.
  5. Document which elements of the analysis come from synthesis and which come from real respondents – otherwise, after a month, no one will be able to tell them apart.

The final point is often overlooked, yet it determines the credibility of the entire report. Without clear source labeling, synthetic data seep into the set of real responses and contaminate the interpretation in a way that cannot later be reversed.

When do synthetic data mislead? The most common pitfalls

The most dangerous errors do not arise because the model is obviously wrong. They arise because it is convincingly wrong. A synthetic response sounds credible, has correct grammar, a coherent narrative, and a logical structure – yet it does not describe any real person. The following are the pitfalls that occur most often in projects and should be recognized before they become part of a report:

  • Loss of the distribution tail – the model may miss extreme, niche, and unusual opinions, even though these often drive real purchase decisions or cancellations.
  • Apparent agreement – AI personas in research can confirm almost any claim if the question is even slightly leading, because the model seeks consistency with the context.
  • Factual hallucinations – a synthetic respondent may “remember” products that never existed, prices that no one set, and experiences they could not have had.
  • Reinforcement of biases from the training dataset – if one cultural or demographic perspective dominated the input data, the synthesis may replicate it and present it as representative.
  • False representativeness – it is easy to generate a thousand responses and mistake their number for evidential strength, even though a thousand variations of the same pattern cannot replace several dozen real conversations.

The comparison with the traditional approach is instructive. A real respondent can be inconvenient: they interrupt, contradict themselves, and say something the researcher did not want to hear. This “inconvenience” is not noise – it is a signal. A synthetic respondent is often polite, on topic, and consistent, which is a drawback rather than an advantage in exploratory research. Therefore, synthetic data should not be positioned against real fieldwork as a cheaper alternative for the same purpose – they serve different purposes.

Another pitfall concerns the validation of synthetic data. Verification solely within the model itself is insufficient – the model may confirm its own output because it operates on the same pattern. Meaningful validation of synthetic data always requires a point of reference outside the model: comparison with a real sample, historical research findings, or hard behavioral data. If there is nothing to compare it with, there is no way to determine whether the synthesis describes the market or merely itself.

What should be checked before using synthetic data in a project? A checklist

Before synthetic data are introduced at any stage of a project, it is worth going through a set of control questions. An answer of “I don’t know” to any of these points is a signal to hold off on synthesis and return to real material:

  • Does the research question concern confirming a known structure or discovering unknown motivations?
  • Is there a real reference dataset against which synthetic data can be validated?
  • Does the report clearly indicate which sections come from response simulation?
  • Is the decision based on these data reversible, or does an error carry high risk?
  • Does the team understand that the number of synthetic responses does not translate into evidential strength?
  • Has the risk been ruled out that the prompt given to the model is leading and forces apparent agreement?

This checklist is not intended to discourage the use of the method. It is intended to help use it where it genuinely saves time and reduces costs without compromising rigor – primarily at the preparatory stage rather than when drawing final conclusions about behaviors and attitudes.

Frequently asked questions

What are synthetic respondents?

They are profiles that answer research questions based on patterns learned by a language model rather than their own experiences. They do not report the experiences of a specific person, but predict what a probable statement from someone with specified characteristics would sound like. For this reason, they work well in supporting tasks rather than as a source of final conclusions about real attitudes.

When do synthetic data help, and when do they distort results?

They help when the objective is to reproduce a known structure – questionnaire pretesting, generating initial hypotheses, or filling gaps based on actual distributions. They may distort results when the task of the study is to discover what is unusual and non-obvious, because the model may smooth distribution tails and gravitate toward averaged responses. The boundary lies between confirming what is known and discovering what is unknown.

How can the credibility of AI-generated responses be verified?

Validation of synthetic data requires a point of reference outside the model itself – a real sample, previous research findings, or hard behavioral data. Verification within the model is not sufficient because it may confirm its own output based on the same pattern. It is also worth testing resilience to leading questions and checking whether the synthesis replicates biases from the training dataset.

Ask where synthetic data make sense in your research project – Hume’s Institute specialists will help define the boundary between what can be safely simulated and what requires a real respondent. Contact us to discuss the methodological assumptions of your study.