Wróć do słownika

Logistic regression

Logistic regression is a statistical method used to estimate the probability that an observation belongs to a defined category, such as purchase versus non-purchase, churn versus retention, or preference for one offer over another. In market research, a logistic regression model helps translate survey, behavioral, CRM, or experimental data into interpretable drivers of choice and risk.

What is logistic regression?

Logistic regression is a predictive and explanatory modeling technique used when the dependent variable is categorical, most often binary. Instead of predicting a continuous value, as in linear regression, logistic regression estimates the probability of an event occurring. The result is usually expressed as a value between zero and one, which can be interpreted as the likelihood of a respondent, customer, company, or household falling into a specific outcome category.

The method is called “logistic” because it uses the logistic function to transform a linear combination of predictors into a probability. Predictors may include demographic variables, attitudes, brand perceptions, price sensitivity, past behavior, exposure to communication, product attributes, or contextual factors. A logistic regression model estimates how each predictor changes the log-odds of the outcome, or the odds when coefficients are exponentiated, while holding other variables constant.

In market research, logistic regression is especially useful because many business questions are categorical by nature. Examples include whether a consumer will buy a product, whether a customer will recommend a brand, whether a lead will convert, whether an employee will leave, or whether a respondent will choose one concept over another. The model does not merely classify cases. It also identifies which factors are associated with higher or lower probability of the target outcome.

Logistic regression is widely used in quantitative research, but it can also support mixed-methods projects. Qualitative research may first identify decision criteria, barriers, motivations, and category language. These insights can then inform the selection of variables for a logistic regression model built on survey or behavioral data. In this way, qualitative understanding improves model design, while quantitative modeling tests which factors matter at scale.

Application of logistic regression in practice

Logistic regression is applied by market researchers, data analysts, CRM teams, customer experience teams, product managers, and marketing analysts when the objective is to understand or predict a discrete outcome. It is useful when the research question concerns probability, classification, or the relative influence of competing drivers.

Typical applications of logistic regression in market and customer research include:

  • Purchase prediction: estimating the probability that a consumer will buy a product based on needs, price perception, brand awareness, category usage, and communication exposure.
  • Churn modeling: identifying customers with a higher likelihood of cancellation, non-renewal, or inactivity by combining satisfaction, usage, complaint, and transactional data.
  • Lead scoring: ranking B2B prospects according to their probability of conversion using firmographic variables, engagement history, stated needs, and sales funnel behavior.
  • Concept and product testing: assessing which features, claims, benefits, or price points increase the likelihood of choosing a tested offer.
  • Brand and communication analysis: examining whether brand associations, advertising recall, trust, or perceived differentiation are linked to consideration or preference.
  • Customer experience research: modeling whether service interactions, wait time, issue resolution, or perceived effort influence retention, recommendation, or complaint behavior.


The phrase “how logistic regression predicts consumer choice” refers to the model’s ability to estimate the probability of selecting an option from observed attributes and respondent characteriztics. For example, in a survey about banking products, a logistic regression model may show whether perceived security, fees, digital convenience, branch access, or trust is most strongly associated with choosing a given bank. The model can then support segmentation, prioritization of product improvements, or targeting of marketing messages.

In practical research work, logistic regression often starts with a clearly defined target variable. This may be a behavioral measure from customer databases or a survey-based variable such as “intends to buy”, “prefers brand A”, or “would switch provider”. The quality of the model depends on the clarity of this outcome, the relevance of predictors, the structure of the sample, and the correct interpretation of coefficients, odds ratios, and predicted probabilities.

Logistic regression and related methods

Logistic regression belongs to a broader ecosystem of statistical modeling methods used in market research and business analytics. It is related to linear regression, discriminant analysis, decision trees, conjoint analysis, machine learning classifiers, and segmentation models, but it has a distinct role because it combines probability estimation with relatively transparent interpretation.

The main distinctions are important for correct methodological choice:

  • Linear regression: predicts continuous outcomes, such as spending, satisfaction score, or usage frequency. Logistic regression is used when the outcome is categorical, especially binary.
  • Multinomial logistic regression: extends logistic regression to outcomes with more than two unordered categories, such as choosing between several brands, channels, or product variants.
  • Ordinal logistic regression: applies when the dependent variable has ordered categories, for example low, medium, and high likelihood to recommend, and when model assumptions such as proportional odds are appropriate.
  • Decision trees and random forests: can capture non-linear patterns and interactions more flexibly; single trees may be relatively easy to explain, while ensembles such as random forests are often less straightforward to interpret than a logistic regression model.
  • Conjoint and choice models: focus specifically on trade-offs between product attributes. Logistic regression may support or complement these methods when the aim is to model binary choice or acceptance.
  • Cluster analysis: groups respondents by similarity. Logistic regression can then explain which segments are more likely to buy, churn, adopt, or prefer a brand.


In mixed-methods research, logistic regression may be combined with qualitative interviews, focus groups, social listening, desk research, or expert workshops. Qualitative stages help define meaningful predictors and interpret unexpected statistical relationships. Quantitative modeling then estimates the strength and direction of associations in a structured dataset. This combination is particularly valuable when managerial decisions require both measurable evidence and contextual understanding.

Hume’s Institute applies logistic regression in projects where the objective is to identify drivers of market behavior, forecast customer outcomes, or quantify the probability of choice under defined conditions. The method is most valuable when decision makers need a model that is both actionable and explainable, not only technically accurate.

Limitations and good practice in logistic regression

Logistic regression is robust and interpretable, but it is not automatically suitable for every research problem. A logistic regression model should be used when the dependent variable is correctly defined as categorical, the predictors are theoretically meaningful, and the dataset contains sufficient variation in the outcome.

Key limitations and good practices include:

  • Association is not causation: logistic regression can identify variables associated with higher probability of an outcome, but causal interpretation requires appropriate design, such as experiments, longitudinal data, or strong controls.
  • Model specification matters: omitted variables, poorly coded predictors, and irrelevant measures may distort interpretation.
  • Multicollinearity should be checked: highly correlated predictors can make coefficient estimates unstable and difficult to interpret.
  • Probabilities should be communicated clearly: business users usually need predicted probabilities and practical scenarios, not only statistical coefficients.
  • Validation is necessary: a model should be assessed on data not used for estimation when prediction quality is important.
  • Segment differences may matter: a single model can hide differences between consumer groups, markets, categories, or customer types.


For market research users, the strongest value of logistic regression is the connection between statistical rigor and managerial interpretation. It answers not only whether an outcome is likely, but also which measurable factors increase or decrease that likelihood. Used carefully, logistic regression provides a clear bridge between data analysis and decisions about targeting, product development, pricing, communication, and customer retention.