{"id":3768,"date":"2026-09-06T00:00:00","date_gmt":"2026-09-05T22:00:00","guid":{"rendered":"https:\/\/humes.pl\/churn-analysis-how-to-predict-customer-churn-using-data\/"},"modified":"2026-09-24T15:24:57","modified_gmt":"2026-09-24T13:24:57","slug":"churn-analysis-how-to-predict-customer-churn-using-data","status":"publish","type":"post","link":"https:\/\/humes.pl\/en\/churn-analysis-how-to-predict-customer-churn-using-data\/","title":{"rendered":"Churn analysis: how to predict customer churn using data"},"content":{"rendered":"<p>The report shows that several percent of the customer base churned last quarter, but it does not answer the question of which customers will leave in the next quarter and why. A <strong>churn rate<\/strong> calculated as a single figure has reporting value rather than operational value &#8211; only breaking it down by cohorts, segments, and behavioral signals makes it possible to act before the contract is terminated. Below is a practical guide to measuring churn, building a predictive model, and combining transactional data with research into the reasons for customer attrition.<\/p>\n<h2>What is churn rate and how should it be measured correctly?<\/h2>\n<p>Churn rate is the rate of customer attrition over a defined period &#8211; most often calculated as the number of customers lost during a given period divided by the number of customers at the beginning of that period. The formula appears simple, but the real difficulty lies in the definitions: who counts as a customer, what constitutes &#8220;loss,&#8221; and which period is used as the unit of measurement.<\/p>\n<p>In subscription models, the matter is relatively clear &#8211; churn means contract termination, non-renewal, or subscription expiry on a specific date. In transactional models (e-commerce, retail, services purchased irregularly), there is no explicit point of cancellation. The customer simply stops buying, and the analyst must arbitrarily establish an inactivity window after which the customer is considered lost. This window should be based on the actual distribution of intervals between purchases in a given category, rather than on a round number chosen for convenience. A pet food store and a furniture showroom have entirely different natural repurchase cycles.<\/p>\n<p>For the proper interpretation of <strong>churn rate<\/strong>, it is also important to distinguish between several variants, as each answers a different question:<\/p>\n<ul>\n<li><strong>Customer churn<\/strong> &#8211; the percentage of lost customer relationships, regardless of their value.<\/li>\n<li><strong>Revenue churn<\/strong> &#8211; the percentage of lost revenue; it shows whether small or key customers are leaving.<\/li>\n<li><strong>Net churn<\/strong> &#8211; accounts for the expansion of remaining accounts (upsells, upgrades), which means it may be lower than gross churn and, in some business models, even negative.<\/li>\n<li><strong>Voluntary and involuntary churn<\/strong> &#8211; cancellation driven by the customer&#8217;s decision versus attrition caused by a failed payment, card expiration, or a change in formal circumstances. These are driven by different causes and require different responses, so combining them obscures the picture.<\/li>\n<\/ul>\n<p><\/br><\/p>\n<p>A second requirement for meaningful measurement is moving away from an average for the entire customer base. An aggregated metric can mask a situation in which churn falls in a low-value segment while rising in a strategic segment. This is why <a href=\"https:\/\/humes.pl\/en\/glossary\/churn-analysis\/\">churn analysis<\/a> begins with decomposition by acquisition channel, plan, starting cohort, region, device, and type of first purchase. The more granular the breakdown, the clearer it becomes that &#8220;churn&#8221; is not a single phenomenon, but several different processes occurring simultaneously.<\/p>\n<p>A third element is the relationship lifetime perspective rather than the calendar quarter. Attrition is unevenly distributed &#8211; the highest risk is usually concentrated in the first weeks or months after activation, then gradually declines. A monthly rate calculated for the entire customer base averages customers at vastly different stages of the relationship and therefore says nothing about whether onboarding is actually working.<\/p>\n<h2>How do you build a churn predictive model and combine it with research into the causes?<\/h2>\n<p>A practical approach to churn involves several stages that should be carried out in sequence &#8211; skipping the earlier ones means the model at later stages will produce results that cannot be implemented.<\/p>\n<p><strong>Stage 1. Cohort analysis.<\/strong> <a href=\"https:\/\/humes.pl\/en\/glossary\/cohort-analysis\/\">Cohort analysis<\/a> groups customers by the point at which they entered the customer base (month of first purchase, week of registration) and tracks what proportion of each group remains active after one, three, or twelve months. A retention table shows the shape of the attrition curve and whether newer cohorts are retained better or worse than older ones. It is one of the simplest diagnostic tools available to an analytics team, while also serving as a benchmark for evaluating any changes in the product, pricing, or communication. Cohorts can also be built by acquisition channel &#8211; this often reveals that some traffic sources deliver customers with structurally shorter lifecycles.<\/p>\n<p><strong>Stage 2. Defining the event and prediction window.<\/strong> The model learns to predict a specific event within a specific horizon: &#8220;cancellation within the next 30 days&#8221; or &#8220;no purchase within 90 days.&#8221; Without a clear definition of the window, the model cannot be validated. It is also necessary to establish a data cutoff point &#8211; features must come exclusively from the period before the prediction window, otherwise information leakage occurs and produces artificially high performance at the testing stage.<\/p>\n<p><strong>Stage 3. Feature engineering.<\/strong> This stage has the greatest impact on model quality &#8211; more so than the choice of algorithm. Variables worth considering fall into several groups:<\/p>\n<ul>\n<li><strong>Behavioral<\/strong> &#8211; frequency and recency of interactions, number of logins, use of product features, and changes in usage intensity relative to the customer&#8217;s own previous baseline.<\/li>\n<li><strong>Transactional<\/strong> &#8211; purchase value and frequency, basket trends, plan downgrades, cancellation of add-ons, and longer intervals between orders.<\/li>\n<li><strong>Service-related<\/strong> &#8211; number and type of contacts with customer service, case resolution time, complaints, escalations, and repeated inquiries on the same issue.<\/li>\n<li><strong>Payment-related<\/strong> &#8211; failed charges, delays, and changes to the payment method.<\/li>\n<li><strong>Contextual<\/strong> &#8211; relationship tenure, acquisition channel, the presence of a promotional offer at sign-up, and an approaching end of the commitment period.<\/li>\n<li><strong>Declarative<\/strong> &#8211; ratings from satisfaction surveys and relationship metrics, provided their use complies with the appropriate legal basis, information obligations, and data protection requirements.<\/li>\n<\/ul>\n<p><\/br><\/p>\n<p>A strong predictor is often not the absolute value but a change in trajectory: a decline in activity relative to the customer&#8217;s individual pattern, rather than relative to the average across the customer base.<\/p>\n<p><strong>Stage 4. Modeling and validation.<\/strong> Logistic regression is most commonly used to classify churn risk (transparent, easy to interpret, and useful as a benchmark), along with tree-based models and gradient boosting (often more accurate but more difficult to interpret, partly offset by methods for explaining variable contributions). Alternatively, survival analysis can be used. Rather than providing a binary &#8220;will leave\/will stay&#8221; answer, it estimates the probability of survival over time &#8211; which can be more convenient when the question concerns not only whether a customer will leave, but also when. Validation must be time-based: the model should be trained on an earlier period and tested on a later one. Evaluation only on a randomly split dataset may inflate results because it ignores seasonal changes and product changes. The <strong>churn rate<\/strong> itself remains a control metric here &#8211; the model should improve the accuracy of targeting actions, not replace measurement of the phenomenon.<\/p>\n<p><strong>Stage 5. Researching the causes.<\/strong> This is where what can be read from the database ends. Predictive analytics identifies probabilities and variables correlated with risk, but correlation is not a mechanism. An increase in customer service contacts before cancellation does not reveal whether the problem was an outage, unclear pricing, or competitor behavior.<\/p>\n<p>A predictive model indicates who is likely to leave, but research among customers who have already left can explain why &#8211; and without this second component, an organization risks optimizing its response to a symptom rather than to the cause.<\/p>\n<p>The research toolkit at this stage relies on several techniques, selected depending on the question:<\/p>\n<ul>\n<li><strong>Exit survey<\/strong> &#8211; a short survey conducted at the point of cancellation, with a list of reasons developed through earlier qualitative exploration and an open-ended question to guard against overlooking unanticipated motives.<\/li>\n<li><strong>In-depth interviews with former customers (IDI)<\/strong> &#8211; reconstructing the decision-making journey: when the first thought of leaving emerged, what triggered it, what customers considered among competitors, and what could have retained them.<\/li>\n<li><strong>Quantitative research on a sample of lost customers<\/strong> &#8211; makes it possible to estimate the share of individual causes and compare it across segments identified by the model.<\/li>\n<li><strong>Control research among remaining customers<\/strong> &#8211; without a reference group, it is impossible to determine whether an identified cause differentiates customers who leave or affects the entire customer base.<\/li>\n<li><strong>Content analysis of customer service contacts<\/strong> &#8211; coding transcripts and service requests as a source of hypotheses for further testing.<\/li>\n<\/ul>\n<p><\/br><\/p>\n<p>Combining both layers creates a structure in which the model is responsible for prioritization (who, when, and with what probability), while research determines the content of the response (why, and what should be changed in the offer, process, or communication). In addition, the findings from research into causes may point back to new variables worth recreating in behavioral data.<\/p>\n<h2>Which errors most often distort churn analysis?<\/h2>\n<p>Attrition analysis is sensitive to several recurring methodological pitfalls. Their consequence is not a lack of results, but misleading results that appear credible.<\/p>\n<ul>\n<li><strong>Inconsistent definition of churn.<\/strong> Changing the inactivity window or the method used to calculate the customer base during monitoring makes the time series no longer comparable, and an apparent improvement in the churn rate turns out to be an artifact of the definition.<\/li>\n<li><strong>Averaging across the entire customer base.<\/strong> A single metric for all segments conceals opposing trends. Customer retention should be reported by cohorts and value segments.<\/li>\n<li><strong>Future data leakage.<\/strong> Including in the feature set variables that arise at the point of cancellation (contact with the retention team, a cancellation discussion, or a reason code) produces a model with excellent metrics and no usefulness &#8211; it predicts an event that has already happened.<\/li>\n<li><strong>Ignoring involuntary churn.<\/strong> Attrition resulting from failed payments requires operational rather than promotional action. Putting it into a single pool distorts the structure of causes.<\/li>\n<li><strong>Evaluating a model using accuracy alone.<\/strong> When the proportion of churned customers in the customer base is low, a model that predicts &#8220;no one will leave&#8221; achieves high accuracy and zero value. Appropriate metrics include precision, recall, ROC AUC or PR AUC, and lift in the highest risk deciles.<\/li>\n<li><strong>No control group for retention actions.<\/strong> Without a randomly assigned control group, it is impossible to distinguish the effect of an action from the natural behavior of high-risk customers, some of whom would have stayed anyway.<\/li>\n<li><strong>Asking about reasons too late.<\/strong> Research conducted many months after churn is affected by rationalization and forgetting, which can alter the stated reasons.<\/li>\n<li><strong>Basing research into causes solely on those who responded to the exit survey.<\/strong> Customers who leave without saying anything often have a different profile from those who provide a reason. It is necessary to assess the response structure and, where possible, compare respondents with the entire population of customers who churned.<\/li>\n<li><strong>Treating churn rate as a metric that replaces diagnosis.<\/strong> The metric level itself indicates neither the cause nor the point of intervention &#8211; it serves as a signal that initiates analysis, not as its outcome.<\/li>\n<\/ul>\n<p><\/br><\/p>\n<p>A limitation of predictive analytics itself is also its dependence on the past. A model reproduces patterns from the training period; after a pricing change, the entry of a new market player, or a product redesign, its effectiveness may decline, requiring monitoring of variable stability and regular retraining. An alternative or complement may be an approach based on early warning indicators &#8211; simple behavioral rules that are clear to operational teams, statistically less accurate but easier to implement and explain.<\/p>\n<h2>When should you launch churn analysis, and when should you start with qualitative research?<\/h2>\n<p>The choice of starting point depends on what data are available and which question is genuinely open at a given time. The following criteria help determine the sequence of work.<\/p>\n<ol>\n<li><strong>The customer base contains transaction and event histories covering at least several full purchase cycles, as well as a sufficient number of churn observations<\/strong> &#8211; you can start with cohort analysis and a predictive model, then design research into causes for the segments identified by the model.<\/li>\n<li><strong>Data are dispersed across systems or customer identification is missing in some channels<\/strong> &#8211; the first step is to organize the definitions and integrate the sources; a model built on an inconsistent dataset will reproduce data errors.<\/li>\n<li><strong>The churn rate has increased suddenly and sharply<\/strong> &#8211; the priority is rapid qualitative diagnosis (interviews with former customers, analysis of customer service requests), because a model trained on historical data may not account for the new cause.<\/li>\n<li><strong>Attrition is stable but high in a specific starting cohort<\/strong> &#8211; this indicates an issue with onboarding or the sales promise; research should focus on the first weeks of the relationship.<\/li>\n<li><strong>The organization already has risk lists, but retention actions are not producing results<\/strong> &#8211; the problem may lie not in prediction but in a lack of knowledge about the causes or in the absence of measurement of intervention effectiveness using a control group.<\/li>\n<\/ol>\n<p><\/br><\/p>\n<p>In practice, both layers operate in a loop: measurement and the model narrow the area of focus, research explains the mechanism, tests with a control group verify the effectiveness of the response, and the results may indicate new variables and risk segments.<\/p>\n<h2>Frequently asked questions<\/h2>\n<h3>How do you calculate churn rate?<\/h3>\n<p>The basic formula is the number of customers lost during a period divided by the number of customers at the beginning of the period, expressed as a percentage. Customers acquired during the period are usually excluded from the denominator or reported separately so that acquisition growth does not reduce the metric. It is also worth calculating revenue churn in parallel, because losing one large account may matter more than losing many small ones.<\/p>\n<h3>What data can predict customer churn?<\/h3>\n<p>Strong signals come from behavioral and transactional data describing changes in behavior over time: a decline in frequency of use relative to the customer&#8217;s own baseline, longer intervals between purchases, cancellation of additional services, and a move to a cheaper plan. They are supplemented by service-related data (repeated requests, complaints, case resolution time), payment data, and contextual data such as relationship tenure and acquisition channel. Declarative data from satisfaction surveys can improve model accuracy if their use complies with the appropriate legal basis and data protection requirements.<\/p>\n<h3>How can you research the reasons why customers cancel?<\/h3>\n<p>The standard setup is an exit survey launched at the point of cancellation, supplemented by in-depth individual interviews with former customers that reconstruct the decision path from the first sign of dissatisfaction to the choice of an alternative. The findings should be compared with research among customers who remain &#8211; otherwise, it is impossible to determine whether the identified reason genuinely differentiates those who leave. Research should be conducted as close as possible to the point of churn, because the risk of rationalization and memory errors increases over time.<\/p>\n<h2>Ask about churn analysis and research into the causes of customer attrition<\/h2>\n<p>Hume&#8217;s Institute designs churn measurement, risk models, and research into cancellation reasons as a single analytics and research process. <a href=\"https:\/\/humes.pl\/en\/contact\/\">Ask about churn analysis and research into the causes of customer attrition<\/a> to determine a scope tailored to the data available.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Churn rate shows how many customers leave, but not why. We explain how to build a predictive churn model and combine it with research into the reasons customers cancel.<\/p>\n","protected":false},"author":4,"featured_media":0,"comment_status":"closed","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"inline_featured_image":false,"footnotes":""},"categories":[912],"tags":[],"slowa_kluczowe":[],"class_list":["post-3768","post","type-post","status-publish","format-standard","hentry","category-badania-i-analizy"],"acf":[],"_wp_attached_file":null,"_wp_attachment_metadata":null,"wpml_media_processed":null,"_wpml_media_usage_in_posts":null,"_wp_attachment_context":null,"_oembed_35c905c64c03156f243b94f18c4eb80f":null,"_wp_attachment_image_alt":null,"rank_math_description":"Churn rate is only the start. Learn how to measure customer attrition, build a predictive churn model and link it to research into why customers leave.","rank_math_focus_keyword":"churn rate","rank_math_contentai_score":null,"_wpml_post_translation_editor_native":null,"_menu_item_type":null,"_menu_item_menu_item_parent":null,"_menu_item_object_id":null,"_menu_item_object":null,"_menu_item_target":null,"_menu_item_classes":null,"_menu_item_xfn":null,"_menu_item_url":null,"_wp_page_template":null,"rank_math_og_content_image":null,"_wp_trash_meta_status":null,"_wp_trash_meta_time":null,"_wp_desired_post_slug":null,"rank_math_primary_category":null,"_acf_changed":null,"wp_pattern_sync_status":null,"_form":null,"_mail":null,"_mail_2":null,"_messages":null,"_additional_settings":null,"_locale":null,"_hash":null,"_config_validation":null,"_wp_old_slug":null,"rank_math_internal_links_processed":"1","_top_nav_excluded":null,"_cms_nav_minihome":null,"_thumbnail_id":null,"_last_translation_edit_mode":"native-editor","_wpml_word_count":"2753","_dp_original":null,"_edit_last":null,"_edit_lock":null,"rank_math_seo_score":null,"_wpml_location_migration_done":null,"_wpml_media_duplicate":null,"_wpml_media_featured":null,"_wp_old_date":"2026-09-24","copied_media_ids":[],"referenced_media_ids":[],"rank_math_title":"Churn rate: how to measure and predict it | Hume's Institute","job_department":null,"_job_department":null,"job_location":null,"_job_location":null,"job_offer_external_link":null,"_job_offer_external_link":null,"footnotes":null,"inline_featured_image":null,"blog_podtytul":null,"_blog_podtytul":null,"blog_czas_czytania":null,"_blog_czas_czytania":null,"blog_dalsza_lektura":null,"_blog_dalsza_lektura":null,"slownik_krotka_definicja":null,"_slownik_krotka_definicja":null,"slownik_cytat":null,"_slownik_cytat":null,"slownik_na_stronie_glownej":"1","_slownik_na_stronie_glownej":null,"slownik_slowa_kluczowe":null,"_slownik_slowa_kluczowe":null,"slownik_w_praktyce":null,"_slownik_w_praktyce":null,"slownik_powiazane":null,"_slownik_powiazane":null,"slownik_kluczowe_punkty":null,"_slownik_kluczowe_punkty":null,"lang":"en","translations":{"en":3768},"pll_sync_post":{},"_links":{"self":[{"href":"https:\/\/humes.pl\/en\/wp-json\/wp\/v2\/posts\/3768","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/humes.pl\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/humes.pl\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/humes.pl\/en\/wp-json\/wp\/v2\/users\/4"}],"replies":[{"embeddable":true,"href":"https:\/\/humes.pl\/en\/wp-json\/wp\/v2\/comments?post=3768"}],"version-history":[{"count":1,"href":"https:\/\/humes.pl\/en\/wp-json\/wp\/v2\/posts\/3768\/revisions"}],"predecessor-version":[{"id":3769,"href":"https:\/\/humes.pl\/en\/wp-json\/wp\/v2\/posts\/3768\/revisions\/3769"}],"wp:attachment":[{"href":"https:\/\/humes.pl\/en\/wp-json\/wp\/v2\/media?parent=3768"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/humes.pl\/en\/wp-json\/wp\/v2\/categories?post=3768"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/humes.pl\/en\/wp-json\/wp\/v2\/tags?post=3768"},{"taxonomy":"slowa_kluczowe","embeddable":true,"href":"https:\/\/humes.pl\/en\/wp-json\/wp\/v2\/slowa_kluczowe?post=3768"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}