The database contains several hundred thousand transaction records, and the question from management is straightforward: what is an individual customer worth, and which customers are worth targeting with separate communications at all? The answer requires two tools calculated from the same data: customer lifetime value, meaning the value a customer brings over time, and RFM analysis, which organizes the database according to actual purchasing behavior. Below is an overview of the methodology: what data are needed, what the calculation looks like, and where the results most often go wrong.
What is customer lifetime value and when is it worth calculating?
Customer lifetime value (CLV) is the cumulative value a customer brings to a company throughout the relationship, most often calculated as margin rather than revenue. In analytical practice, two distinct measurement objects are separated:
- Historical CLV – the total realized margin from the customer’s transactions to date. This is a fact recorded in the data, not a forecast.
- Predictive CLV – the expected value of future purchases, based on a retention model, purchase frequency, and basket value.
The distinction is operationally important. Historical CLV answers the question “how much have we earned so far,” while predictive CLV answers “how much are we likely to earn in the future.” Two different questions, two different methods, and two different levels of uncertainty.
The basic predictive formula for a subscription model or a model with stable purchase frequency is as follows:
CLV = ARPU × margin × (1 / churn rate)
where ARPU is average revenue per customer over a given period, margin is the share of gross profit in revenue, and the inverse of the churn rate gives the expected number of relationship periods, assuming constant churn. If monthly churn is five percent, the expected customer lifetime is twenty months. The value of future cash flows should be discounted, particularly when a substantial share of the expected margin falls in distant periods; in that case, a formula that accounts for both churn and the discount rate is used.
In non-subscription models – retail, e-commerce, and services purchased irregularly – customers do not “terminate a contract,” so the churn rate must be defined arbitrarily or replaced with a model of the probability of a subsequent purchase. This is why approaches such as BG/NBD and Gamma-Gamma are popular: they estimate separately the probability that a customer remains active and the expected value of their transactions.
The point at which calculating customer lifetime value stops being an academic exercise and becomes necessary can usually be recognized by several signals: acquisition costs are rising faster than revenue, the database is already large enough that individual transactions do not distort averages, and the marketing and customer service teams disagree over priorities for the same accounts. CLV does not settle these disputes for anyone, but it provides a shared unit of measurement.
How do you calculate CLV and conduct RFM segmentation using transaction data?
Both tools require the same minimum data: a customer identifier, transaction date, and transaction value. Valuable additions include margin at the order-line level as well as service and return costs, without which CLV measures turnover rather than value.
The sequence of work in an analytical project usually looks like this:
- Cleaning and deduplicating the database. One customer should have one identifier – otherwise, every subsequent figure is biased. Accounts created for the same person under different email addresses reduce apparent purchase frequency and inflate the size of the customer base.
- Defining the observation window. The window must be longer than the typical purchase cycle in the category. For a category purchased once every two years, a twelve-month window will generate apparent churn.
- Calculating baseline metrics. Average order value, purchase frequency, unit margin, and retention rate across successive periods.
- Building cohorts. Customers are grouped by the month or quarter of their first purchase, or alternatively by acquisition channel. This is the step most often skipped – and the one organizations most often pay for.
- Calculating CLV by cohort and segment with an uncertainty interval rather than as a single number.
- RFM analysis as a behavioral layer overlaid on the financial result.
As Hume’s Institute experts point out, customer lifetime value calculated as an average for the entire customer base may not reflect the value of a typical customer – the CLV distribution is often strongly right-skewed, and a small group of customers with high purchase frequency raises the average above the median. A reliable result requires a breakdown by acquisition cohorts or behavioral segments and reporting the median alongside the average.
How to build RFM segmentation step by step
RFM analysis describes each customer using three variables calculated solely from transaction data:
- Recency (R) – the number of days since the last transaction. The lower the number, the better.
- Frequency (F) – the number of transactions within the observation window.
- Monetary (M) – total spending or margin within the observation window, and less often average basket value.
The standard procedure involves dividing the customer base into quintiles for each of the three dimensions. Each customer receives three scores from 1 to 5, producing a code such as 555 (purchased recently, purchases frequently, spends a lot) or 111 (has not purchased for a long time, purchased infrequently, spent little). Formally, this creates 125 combinations, which is not operationally usable – so the codes are combined into several to a dozen groups with a consistent interpretation: highest-value customers, loyal customers of average value, promising new customers, dormant customers, customers at risk of churn, and lost customers.
Business thresholds are also used instead of quintiles – for example, Recency based on the category’s actual purchase cycle rather than the distribution in the customer base. The quintile method offers comparability between periods only if the thresholds are fixed; when recalculated from scratch each month, they will always place twenty percent of customers in each group, regardless of whether the customer base is actually improving.
How to combine RFM with CLV
RFM and customer lifetime value answer complementary questions. RFM organizes the customer base according to the current state of the relationship – it is fast, transparent, and easy for operational teams to understand. CLV values the relationship in monetary terms and over time. The two are usually combined by calculating CLV within RFM segments, making it possible to determine whether the “loyal” segment actually generates margin or simply generates many low-value orders with high service costs. In practice, the segment with the highest purchase frequency may have lower CLV than a segment that purchases less frequently but buys in higher-margin categories and has a lower return rate.
What mistakes most often undermine the results of CLV and RFM analysis?
Most unsuccessful implementations result not from an error in the formula, but from decisions made before the calculation.
- Calculating CLV based on revenue rather than margin. A customer generating high turnover with low margin, a high return rate, and frequent contact with customer service may have a value close to zero. CLV calculated based on revenue systematically rewards such customers.
- Omitting acquisition costs. CLV without reference to CAC does not indicate whether a relationship is profitable, only how much margin it generates before acquisition costs. Comparison requires consistent definitions, periods, and levels of aggregation; it is particularly useful within the same cohorts or acquisition channels.
- No discount rate over a long horizon. In multi-year relationships, omitting discounting inflates the result in a way that increases with the forecast horizon.
- No defined forecast horizon. A formula based on the inverse of churn assumes constant churn and no predetermined end to the relationship, although its expected duration is finite. For planning purposes, it is worth defining a horizon aligned with the business objective and validating the model using data that were not used to build it.
- Treating churn as a constant. Churn may differ substantially across relationship periods and cohorts. Applying average churn to all customers may overstate or understate their CLV. The same issue applies to the interpretation of behavioral segmentation based on a single time window.
- An observation window that is too short in RFM. In categories with long purchase cycles, customers are assigned to the “lost” segment even though they are in a normal interval between purchases.
- RFM segmentation of the entire customer base without separating product categories. If a company sells both fast-moving products and products purchased once every few years, shared Frequency quintiles mix two incomparable populations.
How RFM differs from other customer segmentation approaches
RFM is customer segmentation based solely on observed transaction behavior. Its strength is that it does not require survey research, self-reported data, or marketing consent – purchase history is sufficient. Its weakness is that it does not explain why a customer behaves in a particular way. A “dormant” segment does not reveal whether the customer has switched to a competitor, changed their needs, or simply postponed a purchase.
For this reason, in mixed-methods projects, RFM most often serves as a sampling framework: it enables controlled sampling of respondents for qualitative or quantitative research so that the sample includes customers with different behavioral profiles. The answers to “why” questions come from the self-reported layer – interviews, needs research, and analysis of satisfaction drivers. Methods based on clustering algorithms go further by incorporating additional variables beyond the three RFM dimensions, as described in the article on data-driven segmentation.
It is also worth remembering that RFM is a descriptive measure, while CLV can be a historical or predictive measure; neither measure is causal in itself. They indicate which customers are valuable, but do not explain what made them valuable. This question requires testing, cohort analysis of acquisition channels, or primary research.
What to check before launching the analysis: a data checklist
Before the first calculation, it is worth verifying several conditions whose absence may substantially undermine the credibility of the result regardless of the model’s quality.
- Is the customer identifier unique and stable over time, including after a change of email address or phone number?
- Does the data include returned and canceled transactions, and are they marked in a way that allows them to be deducted?
- Is margin available at the order-line level, or is only revenue available?
- Does the observation window cover at least two complete purchase cycles typical for the category?
- Are there any one-off events in the window, such as a major promotion, price change, or system migration, that distort the Frequency and Monetary distributions?
- Are B2B and B2C customers and sales channels separated if their purchasing models differ?
- Has it been defined how customers with a single transaction will be treated – as a separate segment or as part of the distribution?
- Will the result be reported with the median and distribution, rather than only the average?
An answer of “I don’t know” to any of these points is a signal to organize the data first and calculate customer lifetime value only afterward. CLV analysis performed on an inconsistent transaction database may lead to incorrect conclusions, and its errors are difficult to detect because the final number always looks credible.
Frequently asked questions
How do you calculate CLV?
In the simplest version, CLV is the average revenue per customer over a period multiplied by margin and by the expected number of relationship periods, meaning the inverse of the churn rate assuming constant churn. Future cash flows should be discounted, particularly when a significant share of the value falls in distant periods. For non-subscription data, models of the probability of a subsequent purchase are used instead of constant churn, for example BG/NBD combined with Gamma-Gamma. The result should always be reported by cohort and with the median provided alongside the average.
What is RFM analysis?
RFM analysis is a method for organizing a customer base according to three transaction variables: time since the last purchase (Recency), number of purchases (Frequency), and spending value (Monetary). Each customer receives a score for each dimension, most often on a quintile scale, and the resulting codes are grouped into several segments with a consistent operational interpretation. The method does not require survey research – a properly prepared transaction history is sufficient.
How does RFM segmentation differ from needs-based segmentation?
RFM segmentation is based on behavior observed in the data and answers the question of how a customer buys. Needs-based segmentation is based on self-reported research data and answers the question of what a customer expects and why they choose a particular offer. The two approaches rarely produce consistent segment boundaries, which is why in mixed-methods projects RFM often serves as a framework for sampling in needs research rather than a substitute for it.
Ask about customer value analysis using your data
If you have transaction data but lack a consistent methodology for calculating CLV and dividing your customer base by purchasing behavior, Hume’s Institute can conduct an analysis using your dataset – from data quality verification to ready-to-use segmentation. Get in touch to discuss the scope and required input data.