How to design a survey questionnaire that measures what you actually want to measure

Monika

A survey that collects thousands of responses but is based on poorly designed questions generates data that appears reliable but is actually misleading. The question of how to design a survey questionnaire is not simply about choosing the number of items or the form’s aesthetics – it is about whether the tool actually measures the construct you want to study or merely captures artifacts of wording, order, and scales.

How do you start designing a questionnaire without distorting the results?

The first step in answering the question of how to design a survey questionnaire is to define the research objective precisely and operationalize the constructs. Operationalization means translating an abstract concept (e.g., “customer loyalty,” “satisfaction with a service,” or “purchase intent”) into a set of observable indicators that can be measured with questions. Without this step, a research questionnaire becomes a collection of questions “about something” rather than a measurement tool.

In practice, it is worth starting with three control questions: what exactly is to be measured, at what level of detail, and what decision will be made based on the result. If a survey result does not change any decision, the question is unnecessary. This principle eliminates most “just in case” questions, which lengthen the questionnaire and reduce response quality through respondent fatigue.

The next step is to select the population and the sampling frame. Even the best-designed questionnaire will not compensate for sampling error – if a survey about purchase preferences reaches only the brand’s current customers, it measures the loyalty of the existing customer base rather than market potential. This logic also determines the design of screening questions, which direct respondents to the appropriate questionnaire path.

The questionnaire structure should lead respondents from general to specific questions, from neutral to potentially sensitive ones, with demographic questions usually placed at the end. This order reduces the risk of survey abandonment in the first few minutes and limits the priming effect, in which earlier questions influence responses to subsequent ones.

How should survey questions be worded and measurement scales selected?

The wording of a question determines what you actually measure. The basic principles are well known but consistently violated in practice: one question, one construct; the respondent’s language rather than industry jargon; emotional neutrality; and avoidance of double negatives. The question, “Don’t you think that our new app is more intuitive than the previous version?” contains three errors at once: it suggests an answer, imposes a comparison, and assumes familiarity with both versions.

A particular case involves double-barreled questions, which combine two constructs in one sentence. “Are you satisfied with the price and quality of the product?” forces the respondent to average their evaluation – if the price is acceptable but the quality is disappointing, any answer will be inaccurate. The rule is simple: one question should ask about one thing.

The choice of measurement scale should follow from the level of measurement required for the analysis. Nominal scales (e.g., device type) are suitable for segmentation, ordinal scales (ranking preferences) for establishing hierarchies, interval scales (e.g., rating scales treated in analysis as an approximation of interval scales) for measuring attitudes, and ratio scales (age measured in years, number of transactions) for statistical analyses requiring reference to an absolute zero. The most common choice in satisfaction and attitude research is the Likert scale, while NPS questions (0-10) are often used in loyalty research.

The number of scale points is not neutral. Five-point scales are intuitive and quick, but differentiate respondents less effectively. Seven- and ten-point scales offer greater resolution but increase completion time. Even-numbered scales, without a midpoint, force respondents to take a position, which may be desirable in purchase intent research but risky in attitude measurement, where a “neither yes nor no” response carries information. As Hume’s Institute experts point out, a poorly designed question will not produce a good answer even if you collect a thousand completed surveys – a scale selected without considering the construct turns a precise tool into a source of noise.

In Hume’s Institute projects, pretesting the questionnaire on a small sample, usually from a dozen or so to several dozen respondents from the target group, combined with the think-aloud technique, in which respondents comment aloud on their understanding of questions, has proven to add the greatest value. Pretesting identifies:

  • questions interpreted differently than the researcher intended,
  • jargon terms that are not understood outside the industry,
  • scales on which respondents are unable to place themselves,
  • skip logic that works contrary to its intended purpose,
  • fatigue points at which response quality drops sharply.

Only after pretesting can a research questionnaire be considered ready for fieldwork. Skipping this stage is a common reason why data from large surveys proves unusable at the analysis stage.

Which survey errors most often distort results?

Survey errors fall into three categories: question design errors, order and context errors, and respondent response errors. Each category requires a different risk-mitigation strategy.

Question design errors include leading questions, loaded questions, vague response categories (e.g., “often” or “rarely” without time anchors), and the absence of “don’t know” or “not applicable” options, which forces respondents to provide fictional answers. If a respondent does not use the service covered by a question, the lack of an exit option generates entirely fabricated data.

Order errors include the priming effect, in which an earlier question shapes the interpretation of the next one; the halo effect, in which an overall evaluation carries over into detailed ratings; and the contrast effect, in which adjacent questions are compared with one another rather than evaluated independently. Rotating items within question blocks and randomizing the order of response options in closed-ended questions help limit these effects.

Respondent response errors primarily include social desirability bias, the tendency to give socially acceptable answers; acquiescence bias, the tendency to agree; central tendency bias, the avoidance of extreme responses; and straightlining, or selecting the same answer across matrix blocks. Limiting these distortions requires, among other measures, the careful use of reverse-coded statements in scales, short matrices, and consistency check questions.

An alternative or complement to surveys, when a construct is difficult to measure directly through self-reported responses, such as values, unconscious motivations, or emotional reactions, is qualitative research – in-depth interviews (IDIs), focus group interviews (FGIs), or projective techniques. How does a survey differ from an interview in this context? A survey measures well what respondents are able and willing to report consciously; an interview reaches layers that a questionnaire cannot capture. In mixed-methods projects, the two approaches are combined: qualitative exploration generates hypotheses and language, while a quantitative survey verifies them using an appropriately selected sample.

What should you check before launching a survey? Checklist

Before sending a questionnaire into fieldwork, it is worth reviewing a checklist that organizes the most important design decisions:

  1. Does each question address a specific research objective and influence a decision?
  2. Have the constructs been operationalized, and does each have more than one indicator where justified?
  3. Are the survey questions clear, neutral, and free of double-barreled questions?
  4. Are the measurement scales appropriate for the level of analysis and consistent within blocks?
  5. Does the question order minimize priming and the halo effect?
  6. Have “don’t know” / “not applicable” options been included where substantively justified?
  7. Has the skip logic been tested across all paths?
  8. Is the completion time estimated during pretesting within an acceptable range for the channel (CAWI, CATI, CAPI/PAPI)?
  9. Do the demographic questions enable the planned segmentation analyses?
  10. Has the questionnaire undergone qualitative pretesting with respondents from the target group?

A checklist will not replace methodological expertise, but it reduces the risk of the most serious survey errors, which only become apparent at the analysis stage – when corrections are no longer possible.

Frequently asked questions

Which scales should be used in a survey?

The choice of scale depends on the construct being measured and the planned analysis. The standard for measuring attitudes and opinions is the Likert scale, either 5- or 7-point; for loyalty, an NPS question (0-10); for attribute importance, an ordinal scale or the MaxDiff method; and for behavioral frequency, scales with time anchors rather than labels such as “often/rarely.” The higher the level of measurement, the broader the possibilities for statistical analysis.

How can you avoid leading questions?

All evaluative wording, judgmental adjectives, and assumptions that the respondent has not confirmed should be removed from the question. Instead of asking, “How much do you like the new, improved version of the product?” it is better to ask, “How do you rate the current version of the product?” and add a balanced response scale. It is also helpful to give respondents negative and neutral options on equal terms with positive ones.

How many questions should a survey include?

There is no universal number – completion time and cognitive burden, rather than the number of items, are what matter. In CAWI research, the optimal completion time is usually around ten to fifteen minutes, while in CATI it is shorter. Every question should pass a purpose test: if the answer does not affect any decision or hypothesis, the question should be removed rather than “kept because it might be useful.”

Ask about a survey design tailored to your research objective

If you are planning quantitative research and want to ensure that the questionnaire measures what truly matters for your decision, contact Hume’s Institute – the team of methodologists will prepare a survey design tailored to your research objective, target group, and data collection channel.