{"id":3708,"date":"2026-09-13T00:00:00","date_gmt":"2026-09-12T22:00:00","guid":{"rendered":"https:\/\/humes.pl\/slownik\/missing-data-and-imputation\/"},"modified":"2026-09-24T09:33:21","modified_gmt":"2026-09-24T07:33:21","slug":"missing-data-and-imputation","status":"publish","type":"slownik","link":"https:\/\/humes.pl\/en\/glossary\/missing-data-and-imputation\/","title":{"rendered":"Missing data and imputation"},"content":{"rendered":"<p>Missing data and imputation refer to the identification, assessment and treatment of absent values in a dataset. In market research, the way missing responses are handled can materially affect reported preferences, segment sizes, driver analyses and business decisions based on survey findings.<\/p>\n<p>Data imputation is not simply a technical correction. It is a controlled analytical procedure that should reflect why information is missing, how much is missing and whether replacing absent values could introduce bias.<\/p>\n<h2>What is missing data and imputation?<\/h2>\n<p>Missing data occur when a variable has no recorded value for one or more cases in a research dataset. In survey research, this may mean that a respondent skips a question, abandons the questionnaire before completion, selects \u201cprefer not to say\u201d, or cannot provide an accurate answer. Missing values can also arise through routing errors, data integration problems, interviewer omissions or unavailable administrative records.<\/p>\n<p>Missing data and imputation describe two linked stages of data preparation. The first stage is diagnosing the missingness: determining which variables are affected, which respondents are affected and whether the absence of data follows a recognisable pattern. The second stage, data imputation, involves replacing selected missing values with plausible estimates derived from available information.<\/p>\n<p>The central methodological issue is not whether a dataset contains missing values, but whether those values are likely to distort the analysis. For example, income questions often have higher non-response than questions about product awareness. If respondents with higher incomes are more likely to omit the income question, calculating averages only from completed answers may understate the income level of the target population.<\/p>\n<p>Researchers commonly distinguish between three mechanisms of missingness:<\/p>\n<ul>\n<li><strong>Missing completely at random:<\/strong> the probability of a value being missing is unrelated to observed or unobserved respondent characteristics. This is uncommon in practice but causes the least concern for statistical bias.<\/li>\n<li><strong>Missing at random:<\/strong> conditional on observed information in the dataset, missingness is not related to the unobserved value itself. For instance, younger respondents may be less likely to answer an employment-status question, while their age is known.<\/li>\n<li><strong>Missing not at random:<\/strong> missingness is related to the unobserved value itself. A respondent may decline to disclose income specifically because it is unusually high or low. This is the most difficult situation to address reliably.<\/li>\n<\/ul>\n<p><\/br> <\/p>\n<p>These categories guide decisions about whether data imputation is appropriate and which method is defensible. They do not prove the true reason for every missing answer, but they provide a structured framework for evaluating risk.<\/p>\n<h2>Application of missing data and imputation in practice<\/h2>\n<p>Missing data and imputation are used primarily in quantitative market research, especially when a dataset supports segmentation, modelling, forecasting, tracking or comparisons between customer groups. They are relevant to online surveys, telephone interviews, panel studies, customer-experience measurement, employee surveys and data products that combine survey results with CRM, transaction or behavioural data.<\/p>\n<p>Knowing how to handle missing data in survey research is particularly important when a key variable is missing for a non-random group of respondents. Deleting all incomplete questionnaires may reduce the usable sample and change its composition. Conversely, retaining incomplete cases without a clear rule may produce inconsistent bases for different survey questions.<\/p>\n<p>In practice, the appropriate approach depends on the analytical purpose. Typical applications include:<\/p>\n<ul>\n<li><strong>Customer segmentation:<\/strong> imputing selected profile variables can preserve cases needed to assign respondents to segments, provided the assumptions are documented and validated.<\/li>\n<li><strong>Brand tracking:<\/strong> consistent rules for missing data and imputation help maintain comparability between waves and prevent operational differences from being interpreted as market change.<\/li>\n<li><strong>Driver analysis:<\/strong> models of satisfaction, loyalty or purchase intention can be sensitive to missing predictor values. Imputation may allow the retention of useful cases while reducing avoidable loss of information.<\/li>\n<li><strong>B2B research:<\/strong> respondents may not know precise company-level metrics, such as procurement budgets or headcount. A missing-value strategy helps distinguish genuine lack of knowledge from refusal or survey-design problems.<\/li>\n<li><strong>Data integration:<\/strong> when survey data are linked with customer databases or external datasets, missing identifiers and unmatched records require explicit treatment before modelling or reporting.<\/li>\n<\/ul>\n<p><\/br> <\/p>\n<p>Hume&#8217;s Institute may apply missing data and imputation procedures after reviewing questionnaire logic, fieldwork quality and the structure of non-response. The method should be selected before final analysis, documented in the technical reporting and considered when interpreting findings for business stakeholders.<\/p>\n<h2>How to handle missing data in survey research<\/h2>\n<p>A sound missing-data procedure starts with prevention rather than imputation. Clear questionnaire wording, appropriate answer options, logical routing, mobile-friendly survey design and interviewer training can reduce avoidable item non-response. Questions perceived as sensitive should include carefully designed response alternatives where appropriate, rather than forcing an answer that may be inaccurate.<\/p>\n<p>When missing values remain, the analytical process should usually include the following steps:<\/p>\n<ol>\n<li><strong>Measure the extent and location of missingness.<\/strong> Missing values should be reviewed by question, respondent group, survey wave and data source.<\/li>\n<li><strong>Identify the likely mechanism.<\/strong> Analysts assess whether missingness is associated with known characteristics such as age, customer status, channel, device or questionnaire length.<\/li>\n<li><strong>Decide whether to retain, exclude or impute cases.<\/strong> Not every missing value should be replaced. A question that is central to the research objective may require a different treatment from an optional descriptive variable.<\/li>\n<li><strong>Select a suitable data imputation method.<\/strong> The method should match the type of variable, the amount of available auxiliary information and the intended analysis.<\/li>\n<li><strong>Validate and report the decision.<\/strong> Results should be checked for implausible distributions, altered relationships between variables and sensitivity to alternative assumptions.<\/li>\n<\/ol>\n<p><\/br> <\/p>\n<p>Common imputation methods include replacement with a mean, median or mode; hot-deck imputation using values from similar respondents; regression-based estimation; and multiple imputation, which creates several plausible versions of missing values and incorporates uncertainty into subsequent analysis. Simple replacement methods are easy to implement but can reduce natural variation in the data. More advanced methods may better preserve relationships between variables, but they also require stronger modelling decisions and careful quality control.<\/p>\n<p>Data imputation should not be used to conceal poor fieldwork or to create information that respondents were unable or unwilling to provide. In qualitative research, missing information is usually handled differently. An unaddressed interview topic, incomplete observation or unavailable participant is interpreted in context, with attention to recruitment, interview conduct and potential gaps in the evidence rather than statistically imputed.<\/p>\n<h2>Missing data and imputation in relation to other methods<\/h2>\n<p>Missing data and imputation are closely related to data cleaning, weighting, quality control and statistical modelling, but they serve different purposes. Data cleaning corrects errors such as invalid formats, duplicates or impossible values. Imputation addresses absent values. Weighting adjusts the influence of completed cases to improve alignment with a target population, whereas imputation estimates values within individual records. Both may be used in the same project, but one does not replace the other.<\/p>\n<p>The distinction is especially important in survey research. Weighting can reduce bias when some population groups are underrepresented among respondents. Data imputation can retain records with missing answers needed for a multivariate analysis. If a particular group is underrepresented and also has high item non-response, both issues may need to be assessed together.<\/p>\n<p>Missing data and imputation also differ from exclusion methods. Listwise deletion removes every respondent with a missing value in any variable used in an analysis. Pairwise deletion uses all available answers for each specific calculation. These approaches can be acceptable in limited circumstances, particularly when missingness is low and plausibly random, but they can produce changing sample bases and biased results when non-response is systematic.<\/p>\n<p>In mixed-methods research, missing-data diagnostics can inform the qualitative stage. For example, if survey participants frequently omit questions about switching suppliers, follow-up interviews may explore whether the topic is unclear, sensitive or not applicable to particular customer groups. This links quantitative data quality assessment with richer interpretation of respondent behaviour.<\/p>\n<p>Ultimately, the purpose of missing data and imputation is to preserve the analytical value of research data without overstating certainty. A transparent approach makes it possible to distinguish observed evidence from estimated values and to evaluate whether key conclusions remain robust under reasonable alternative assumptions.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Missing data and imputation cover identifying, assessing and treating absent values in a dataset. How gaps are handled can materially affect survey findings and business decisions.<\/p>\n","protected":false},"template":"","slowa_kluczowe":[],"class_list":["post-3708","slownik","type-slownik","status-publish","hentry"],"acf":[],"_wp_attached_file":null,"_wp_attachment_metadata":null,"wpml_media_processed":null,"_wpml_media_usage_in_posts":null,"_wp_attachment_context":null,"_oembed_35c905c64c03156f243b94f18c4eb80f":null,"_wp_attachment_image_alt":null,"rank_math_description":"Concept definition: Missing data and imputation. Application in market research and methodology practice. Check the Hume's Institute glossary.","rank_math_focus_keyword":"Missing data and imputation","rank_math_contentai_score":null,"_wpml_post_translation_editor_native":null,"_menu_item_type":null,"_menu_item_menu_item_parent":null,"_menu_item_object_id":null,"_menu_item_object":null,"_menu_item_target":null,"_menu_item_classes":null,"_menu_item_xfn":null,"_menu_item_url":null,"_wp_page_template":null,"rank_math_og_content_image":null,"_wp_trash_meta_status":null,"_wp_trash_meta_time":null,"_wp_desired_post_slug":null,"rank_math_primary_category":null,"_acf_changed":null,"wp_pattern_sync_status":null,"_form":null,"_mail":null,"_mail_2":null,"_messages":null,"_additional_settings":null,"_locale":null,"_hash":null,"_config_validation":null,"_wp_old_slug":null,"rank_math_internal_links_processed":"1","_top_nav_excluded":null,"_cms_nav_minihome":null,"_thumbnail_id":null,"_last_translation_edit_mode":null,"_wpml_word_count":"1692","_dp_original":null,"_edit_last":null,"_edit_lock":null,"rank_math_seo_score":null,"_wpml_location_migration_done":null,"_wpml_media_duplicate":null,"_wpml_media_featured":null,"_wp_old_date":"2026-09-24","copied_media_ids":[],"referenced_media_ids":[],"rank_math_title":"Missing data and imputation - definition | Hume's Institute","job_department":null,"_job_department":null,"job_location":null,"_job_location":null,"job_offer_external_link":null,"_job_offer_external_link":null,"footnotes":null,"inline_featured_image":null,"blog_podtytul":null,"_blog_podtytul":null,"blog_czas_czytania":null,"_blog_czas_czytania":null,"blog_dalsza_lektura":null,"_blog_dalsza_lektura":null,"slownik_krotka_definicja":null,"_slownik_krotka_definicja":null,"slownik_cytat":null,"_slownik_cytat":null,"slownik_na_stronie_glownej":"1","_slownik_na_stronie_glownej":null,"slownik_slowa_kluczowe":null,"_slownik_slowa_kluczowe":null,"slownik_w_praktyce":null,"_slownik_w_praktyce":null,"slownik_powiazane":null,"_slownik_powiazane":null,"slownik_kluczowe_punkty":null,"_slownik_kluczowe_punkty":null,"_links":{"self":[{"href":"https:\/\/humes.pl\/en\/wp-json\/wp\/v2\/slownik\/3708","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/humes.pl\/en\/wp-json\/wp\/v2\/slownik"}],"about":[{"href":"https:\/\/humes.pl\/en\/wp-json\/wp\/v2\/types\/slownik"}],"version-history":[{"count":1,"href":"https:\/\/humes.pl\/en\/wp-json\/wp\/v2\/slownik\/3708\/revisions"}],"predecessor-version":[{"id":3709,"href":"https:\/\/humes.pl\/en\/wp-json\/wp\/v2\/slownik\/3708\/revisions\/3709"}],"wp:attachment":[{"href":"https:\/\/humes.pl\/en\/wp-json\/wp\/v2\/media?parent=3708"}],"wp:term":[{"taxonomy":"slowa_kluczowe","embeddable":true,"href":"https:\/\/humes.pl\/en\/wp-json\/wp\/v2\/slowa_kluczowe?post=3708"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}