You have several thousand open-ended survey responses, interview transcripts, or social media comments in your database – and you know that manual coding will take weeks, while a summary such as “customers want better service” will add nothing. Text mining in market research makes it possible to move from a chaotic collection of statements to an organized structure of themes, sentiment, and relationships that can be meaningfully interpreted. It is not algorithmic magic, but a discipline that combines natural language processing with research methodology.
When does text mining in market research make sense, and when is traditional coding enough?
Open-ended survey questions, product reviews, customer service inquiries, social media posts, and focus group transcripts – all these sources generate unstructured data. As long as the dataset contains a few hundred records, two researchers can manage manual coding in a spreadsheet. The problem begins with several thousand statements, data from multiple channels at once, or projects that require regular monitoring, such as monthly analysis of customer feedback from an NPS survey.
Text mining in market research addresses three specific needs. First, volume reduction – from 10,000 comments to several dozen coherent themes. Second, repeatability – the same algorithm applied to data from the next quarter can deliver more comparable results than a rotating team of coders, provided that a consistent analytical pipeline and quality control are maintained. Third, identifying patterns that are not visible to the naked eye – co-occurring concepts, changes in sentiment over time, and differences between customer segments.
Not every project requires this. If the goal is to gain an in-depth understanding of the motivations of 20 respondents after in-depth interviews, algorithm-assisted text analysis would be excessive – traditional thematic analysis will work better. Text mining starts to pay off where scale makes manual work impossible or where standardization over time is needed.
What does the text mining process look like in a research project?
Text analysis in market research is a sequence of steps in which technology serves as a tool rather than replacing the researcher’s decisions. Each stage requires methodological choices that affect the final conclusions.
A typical project includes the following stages:
- Data collection and consolidation – exporting responses from a CAWI platform, collecting reviews or API data in accordance with the source’s terms of use, and obtaining transcripts. At this stage, formats are standardized, metadata are coded (date, channel, respondent segment), and duplicates are removed.
- Language preprocessing – tokenization, lemmatization (reducing words to their base form, which is crucial in Polish due to its inflectional structure), removal of stopwords, and typo correction. The quality of this step determines the quality of the entire analysis.
- Topic modeling – algorithms such as LDA, BERTopic, or methods based on embeddings from transformer models group statements into thematic clusters. The number of topics is a researcher’s decision, not an algorithmic output.
- Sentiment analysis – classifying statements as positive, negative, or neutral, sometimes with an additional layer of emotions, such as frustration, enthusiasm, or disappointment, if the model and data allow for it. In Polish, this requires models trained on native-language corpora, because direct translations of English-language tools often fail when dealing with irony and colloquialisms.
- Entity and relationship extraction – identifying brands, products, attributes, and locations mentioned in the text, as well as the relationships between them.
- Interpretation and reporting – translating model outputs into the language of business decisions and validating them against a sample of manually coded cases.
NLP in research now makes it possible to do things that were beyond the reach of many market research projects just five years ago – language models handle context, ambiguity, and colloquial language much better than earlier approaches based primarily on dictionaries. As Hume’s Institute experts point out, a thousand customer reviews are a mine of knowledge, but only if you have the right tools and a researcher who can interpret that knowledge – an algorithm alone will identify clusters, but it is a person who determines whether an “app issue” means a technical error, an unclear interface, or frustration caused by something entirely different.
An example from practice: in a bank customer feedback analysis project covering well over ten thousand comments from relationship surveys, topic modeling identifies more than a dozen main areas – from service at branches and the mobile app to fees. Sentiment calculated at the overall level will show the general mood, but only when themes are analyzed alongside sentiment and customer segment does it become clear that frustration is concentrated around a specific issue within a particular group. This is the appropriate level of insight from text mining.
What is most often missing from text mining projects, and what mistakes are most common?
Natural language processing in market research has accumulated myths that lead to disappointment after the first project. It is worth naming them.
The first mistake is treating text mining as a “switch it on and it is ready” tool. Off-the-shelf libraries will produce an output for any dataset, but without validation, there is no way to know whether that output makes sense. The standard should be to manually code a representative sample, usually at least several dozen to several hundred statements depending on the project’s objective and scale, and compare it with the model output – both during calibration and as a quality check throughout the project.
The second mistake is ignoring the specifics of the Polish language. Most available models and tutorials were created for English. Polish inflection, flexible word order, double negations, diminutives, and irony require language models trained on Polish corpora, such as HerBERT, PolBERT, or Polish applications of multilingual models. Using machine translation as preprocessing may introduce distortions that cannot later be reversed.
The third mistake is placing too much trust in sentiment as an indicator. Sentiment is a simplification – a statement such as “they finally changed that app, it was terrible” may be incorrectly classified as negative, even though it expresses satisfaction with the change. In Hume’s Institute projects, simply counting positive and negative statements is rarely enough – value only emerges when sentiment is analyzed in the context of a specific theme and respondent segment.
The fourth mistake is a lack of iteration. The first run of a topic model almost never produces results ready for presentation. Topics merge, split apart, and some clusters are preprocessing artifacts. A good project includes several rounds of calibration involving a researcher familiar with the client’s industry.
The fifth mistake is confusing text mining with qualitative analysis. Topic modeling will show what people talk about and how frequently. It will not explain why they talk about it or what needs lie behind it. In-depth interviews, ethnography, or thematic analysis in its traditional form are still needed for that. The best results come from a mixed-methods approach in which text mining identifies areas for further exploration and qualitative research provides interpretation.
What can realistically be gained from analyzing large text datasets, and in which projects does it work well?
Text mining in market research works best in several recurring applications:
- Analysis of open-ended questions in quantitative research – when a survey includes questions such as “why did you choose this brand?” and the sample consists of several thousand respondents, manual coding becomes a project bottleneck.
- Monitoring customer feedback over time – recurring analysis of reviews, social media comments, and customer service inquiries by theme and sentiment.
- Voice of Customer in transactional data – exploring comments from post-purchase surveys, complaint forms, and chats with customer service representatives.
- Analysis of in-depth interviews in larger-scale projects – when a project includes several dozen IDI transcripts, text mining supports the researcher in identifying themes without replacing interpretation.
- Analysis of discussions on social media and industry forums – mapping topics, identifying key themes, and tracking changes in narratives over time.
In each of these applications, text analysis delivers value provided that it is designed together with the rest of the study – from question design to the reporting approach. Adding text mining as a “bonus” after data collection usually results in weak outcomes because the dataset was not prepared for this type of analysis.
Frequently asked questions
What is text mining?
Text mining is a set of techniques for extracting structured information from unstructured textual data – reviews, comments, transcripts, and documents. It combines statistical methods, machine learning, and natural language processing to identify themes, sentiment, entities, and relationships between them in text. In market research, it is used where the scale of data makes manual analysis impossible or where repeatable results over time are needed.
How does NLP detect sentiment in customer feedback?
Modern NLP models classify sentiment based on patterns learned from large corpora of manually labeled texts. Transformer models, such as BERT and its Polish variants, analyze not individual words but the context of entire sentences, which allows them to handle negations and nuances to some extent. Even so, sentiment classification remains a simplification – irony, sarcasm, and mixed statements require manual validation on a control sample.
When does automated text analysis not replace a researcher?
Whenever a project requires interpretation of “why,” rather than only “what” and “in what proportions.” An algorithm will identify thematic clusters and the sentiment distribution, but it will not explain the motivations behind statements or connect findings to the market context or the specifics of a customer segment. Text mining works well as an exploratory and monitoring layer, while in-depth interpretation, the development of methodological recommendations, and results validation remain the researcher’s responsibility.
Discuss a project involving the analysis of large text datasets
If you are dealing with thousands of open-ended responses, reviews, or transcripts and are looking for a way to draw structured conclusions from them, the Hume’s Institute team can help design an analysis tailored to the specifics of your data and research objective. Get in touch to discuss the project scope and select the right tools.