Open text analysis turns survey open-ends, support tickets, and reviews into structured findings. Most tools stop at sorting responses into existing categories or sentiment scores, so a new complaint gets forced into the nearest bucket or missed as noise. Skimle instead lets themes emerge from the data first, then lets you dig into the segments and quotes behind each one.
Run a search for "open text analysis" and the results are dominated by native survey platforms describing what their AI does to your open-ended question. SurveyMonkey will "automatically categorize your qualitative responses into meaningful themes." QuestionPro offers a word cloud, a text tagging feature, and sentiment scoring. Alchemer's help page is admirably direct about it: with Open Text Analysis, "you can read through responses to each open text question in your survey and bucket them into categories."
All three describe the same underlying operation with different amounts of automation: take the text, sort it into buckets, count what landed where. That is a real strength when the categories are already right: it is fast, it lives inside the same platform where the survey was built, and for tracking a known metric wave over wave it is often all you need. It answers how the responses you received break down against categories you already had in mind. It answers almost nothing about the question most teams actually care about, which is "what did I miss" or "what new is emerging"?
This article is about the discovery step that classification-first tools, and most manual workflows, skip, and it is aimed squarely at market researchers and customer insights teams sitting on open text data from surveys, tickets, or reviews that has only ever been sorted, never actually explored. If you already know your response volume is the bottleneck and you want the practical, step-by-step workflow for coding hundreds or thousands of responses, how to analyse open text responses at scale covers that ground in detail. And if you're earlier in the process, still deciding how to write the questions that will generate this data, our guide to open-ended questions covers question design and the basics of coding responses. This post assumes you have the data and focuses specifically on the difference between sorting it and learning something new from it.
Why does classifying open text miss the most valuable findings?
Every classification approach, whether it is a survey platform's built-in AI, a spreadsheet with manually defined buckets, or a generic sentiment API, shares the same structural limit: it can only sort responses into categories that already exist. The categories come from one of three places. Someone typed them in beforehand (a deductive codebook). The tool trained its sentiment model on generic text (positive, negative, neutral). Or an LLM generated a plausible-sounding set of buckets by skimming a handful of responses (which is what "AI categorisation" usually means under the hood).
In every case, the moment a response describes something outside that set, one of two things happens. It gets forced into the nearest existing bucket, where it dilutes a category with content that does not really belong. Or it lands in a catch-all like "other" or "miscellaneous," where it is invisible in every report anyone actually reads.
Consider a concrete example. A SaaS company runs a quarterly NPS survey with a codebook built around price, support, and feature gaps because that is what past complaints have been about. This quarter, a new integration partner changes their API without warning, and 40 respondents mention broken workflows tied to that specific integration. None of those responses match an existing code cleanly. Some get coded as "feature gaps." Some get coded as "other." The result: a new, urgent, and fixable problem is smeared across two existing categories and a junk drawer, invisible until someone happens to read the raw comments and notices the pattern by hand.
This is not a hypothetical edge case. The same pattern shows up on the inbound side: a retailer's support team tags tickets into a fixed set of ticket types (shipping, returns, billing, product defect), and when a warehouse switch causes a wave of wrong-item deliveries, those tickets scatter across "shipping" and "product defect" rather than surfacing as the single, specific, and highly fixable problem they actually are. It is what predefined categories are structurally built to do: confirm what you already suspected, and struggle with what you didn't. Sentiment scoring has the same blind spot from a different angle: it can tell you a response is negative, but "negative because of a new integration bug" and "negative because pricing went up" score identically. Sentiment analysis is a useful, fast signal layered on top of themes. It is not a substitute for knowing what the themes are, and treating it as sufficient on its own is where most teams' analysis quietly stops being useful.
What is inductive theme discovery, and how is it different from classification?
Inductive theme discovery starts from the opposite direction. Instead of defining categories first and sorting responses into them, you let the categories emerge from what the data actually contains. The method reads across the full set of responses, identifies recurring patterns in what people are actually saying, and proposes themes based on that, rather than on a list someone wrote down beforehand.
The practical difference shows up immediately in the integration-bug example above. An inductive pass over the same 40 responses does not need "broken integration workflows" to already exist as a category. It notices that 40 responses cluster around the same specific complaint, distinct from the existing price and support themes, and surfaces it as its own theme, with the actual count and actual quotes attached. Nobody had to anticipate the problem in advance for it to become visible.
This does not mean deductive coding (working from a fixed codebook you already trust) is wrong. Sometimes you know exactly what you are measuring: a compliance review checking for five specific disclosures, or a longitudinal study tracking the same five categories every quarter to compare trends over time. Deductive coding is the right tool there. The mistake is treating classification as the only mode available, when most open text analysis benefits from starting inductive and adding structure afterwards rather than the reverse. Skimle's predefined categories mode is built around that reality: inductive and deductive coding live in the same mode, so you can start with a codebook, start with nothing, or mix both, and subcategories still emerge from the data underneath whichever framework you chose. The automatic thematic analysis engine underneath does the same whole-corpus read regardless of which starting point you pick.
The table below summarises the practical difference.
| Classification (predefined categories, sentiment) | Inductive discovery | |
|---|---|---|
| Starting point | A codebook or sentiment model defined in advance | The data itself |
| Handles anticipated topics | Yes, reliably | Yes, but rediscovers them rather than assuming them |
| Handles unanticipated topics | Poorly (forced fit or "other" bucket) | Well (surfaces as a new theme) |
| Output | Counts against fixed categories | Themes with variable structure, sized by prevalence |
| Best for | Tracking known metrics over time, compliance checks | Open-ended feedback, first-pass analysis, anything with volume you haven't read yet |
| Risk if used alone | New problems stay invisible | Less comparable across waves unless themes are pinned down over time |
Neither mode replaces the other. The point is knowing which one you are actually running, since most native survey-tool AI runs the left column and calls it analysis.
3 things digging deeper into a theme actually means
Finding a theme is the easy half. Discovery only pays off once you dig into what you found, and in practice that means three concrete moves.
Follow the theme back to its actual quotes. A theme label like "onboarding confusion, 34 responses" tells you almost nothing on its own. The value is in opening it and reading what those 34 people actually said, in their own words, not a paraphrase. Is the confusion about where to start, what a specific setting does, or a broken step in the flow? A theme without traceability back to source quotes is a claim you cannot verify, which matters practically the moment a stakeholder asks how you know 34 people said that, rather than 4 or 340. "The AI found it" is not an answer that survives a client review or a board meeting; a quote count you can click through to source responses is.
Check how the theme varies by segment. Once a theme exists, the next question is always whether it is universal or concentrated. Does "onboarding confusion" hit new signups specifically, or enterprise accounts migrating from a competitor, or one particular acquisition channel? A theme that looks like a broad, systemic problem in the aggregate view can turn out to be a narrow, specific, and much more fixable issue once you cut it by segment. This is the metadata cross-tabulation step, covered in more detail below.
Split an overly broad theme as more data arrives. Early in a dataset, "pricing" might reasonably be one theme. Once you have 500 responses instead of 50, "pricing" as a single bucket usually stops being useful, because it is quietly holding three different complaints: the price increase itself, confusing tier structure, and a perceived mismatch between price and feature set. A discovery-first workflow expects this and lets you split the broad theme into sharper subthemes as volume grows, rather than locking in whatever granularity felt right at 50 responses. This is the opposite failure mode from classification: instead of a new topic being invisible, an old topic becomes too coarse to be actionable, and the fix is the same muscle either way, keep checking what the data supports rather than trusting the first pass.
Doing this by hand across hundreds of responses is slow. Re-reading a 300-response dataset closely enough to check one theme against five segments, by filtering and re-reading in a spreadsheet, is realistically a half-day of work per theme, which is exactly why it gets skipped and analysis stalls at the "here are five themes" slide. Skimle keeps every theme linked back to the respondent and quote behind it, so a category that looks tidy in a summary can be opened and checked in seconds, and split or merged as new data lands.
Does this work the same for survey open-ends and inbound feedback?
Survey open-ends and inbound feedback (support tickets, app reviews, complaint emails, free-text NPS comments) are usually owned by different teams, live in different systems, and get analysed with different tools. Structurally, though, they are the same kind of data: unstructured text produced by a customer, with the same discovery problem underneath. A new complaint pattern is exactly as invisible in a stack of support tickets sorted into ticket-type categories as it is in a survey categorised against a fixed codebook.
Analysing them together, rather than siloed by channel, is where the discovery approach pays off twice over. First, because combining feedback channels into one category structure means a theme that appears faintly in the survey and strongly in support tickets shows up as one connected finding instead of two separate, smaller, and easier-to-dismiss signals. Second, because survey data alone is a narrow window: McKinsey research on customer experience programmes found that the typical CX survey samples only around 7% of a company's customers, meaning inbound channels like tickets and reviews often carry the majority of the unstructured signal a company actually has, simply because nearly everyone who is dissatisfied enough to write in does so unprompted, while very few of them are selected for a survey.
NPS is the clearest single-channel example of the classification trap, because the 0-10 score and the verbatim comment underneath it are almost always analysed at different levels of rigour. The score gets a dashboard; the comment gets skimmed. NPS verbatim analysis at scale covers the specific workflow for that channel: separating promoter and detractor themes (they rarely overlap), and treating the free-text comment as the actual explanation for the score rather than decoration on top of it.
Sentiment scoring sits comfortably underneath all of this as a fast filter (which comments are worth reading first, which segment sounds most frustrated this quarter) but it answers "how do people feel" rather than "what is actually happening," and customer sentiment analysis that stops at the positive/negative/neutral score is the single most common way teams convince themselves they have analysed their feedback when they have only summarised its mood.
If you are looking specifically at tooling for this rather than a native survey feature, our comparisons of focus group analysis software, customer research analysis tools, and customer insights research platforms cover where discovery-first analysis fits against the wider landscape.
How do you compare themes across segments once they exist?
Discovery gives you themes. The next question almost every team asks is whether those themes land evenly or concentrate somewhere specific: a customer tier, a region, a survey wave, a product version, a support channel. This is a cross-tabulation between your themes and your metadata (the structured fields attached to each response, such as plan tier, tenure, region, or channel).
The mechanics matter here because doing this by hand does not scale past a handful of segments. With a spreadsheet, comparing one theme against five customer segments means five separate filters and five separate reads of the filtered rows, repeated for every theme you want to check. Skimle's metadata analysis treats this as a single operation instead: attach segment, tenure, region, or any structured variable to each response as metadata, then see automatically where those variables actually explain a difference in what people are saying, rather than manually re-slicing the same data for every combination you can think to check.
This is also where a brief aside on employee data is worth making, even though this post is written for customer and market researchers rather than HR teams: the same segment-comparison logic (does this theme hit one department, one tenure band, one manager, more than others) is exactly how employee engagement open-text analysis turns a generic "communication" theme into an actionable, department-specific finding. The discovery-then-segment pattern is the same regardless of whether the respondent is a customer or an employee.
Frequently asked questions
What is the difference between open text analysis and thematic analysis?
Open text analysis is a broad term for any process that extracts structured findings from unstructured written responses, including simple word counts, sentiment scoring, and manual bucketing. Thematic analysis is a specific, more rigorous method within that broader category: it identifies patterns of meaning (themes) across a dataset, whether by working from a predefined codebook (deductive) or letting themes emerge from the data (inductive). Most native survey-tool "AI text analysis" features are closer to classification than to thematic analysis in this stricter sense.
Can sentiment analysis replace theme discovery?
No. Sentiment analysis tells you whether a response is positive, negative, or neutral. It cannot tell you why, which means two responses about completely different problems can carry the same sentiment score and look identical in a sentiment-only report. Sentiment is a useful, fast complementary signal (worth checking which segment sounds most negative this quarter) but it needs themes underneath it to be actionable.
How much open text data is normal to have unanalysed?
There is no fixed benchmark, but the underlying imbalance is well documented: McKinsey's research on customer experience measurement found that a typical CX survey reaches only around 7% of a company's customers, which means the inbound channels most teams analyse least rigorously (support tickets, reviews, unprompted complaints) often carry more raw feedback volume than the survey channel that gets the most analytical attention.
Do I need a fixed codebook before I start analysing open text?
No, and starting with one before you have looked at the data is often counterproductive. A better sequence is to run inductive discovery first to see what themes the data actually contains, then formalise a codebook for whichever themes you want to track consistently across future survey waves or reporting periods. This gets you the benefits of both: nothing new is missed on the first pass, and recurring themes still get consistent, comparable tracking over time.
Is inductive discovery slower than using a fixed codebook?
Not in practice with AI-assisted tools, since both modes read the same volume of text. The difference is in what happens to text that does not fit the codebook: a fixed-codebook pass either force-fits it or drops it, while an inductive pass surfaces it as a new theme automatically. Manually, inductive coding is typically slower up front because you are building the codebook as you go rather than applying one that already exists, though it saves the rework of discovering later that your original categories missed something important.
Ready to find what your current categories are missing?
Try Skimle for free and run inductive discovery across your survey open-ends, support tickets, and reviews in one project, with every theme traceable back to the quotes and respondents behind it.
Want to go deeper on adjacent topics? See our practical workflow for analysing open text responses at scale, our guide to writing and analysing open-ended questions, and how to combine insights across feedback channels.
About the authors
Henri Schildt is a Professor of Strategy at Aalto University School of Business and co-founder of Skimle. He has published over a dozen peer-reviewed articles using qualitative methods, including work in Academy of Management Journal, Organisation Science, and Strategic Management Journal. His research focuses on organisational strategy, innovation, and qualitative methodology. Google Scholar profile
Olli Salo is a former Partner at McKinsey & Company where he spent 18 years helping clients understand the markets and themselves, develop winning strategies and improve their operating models. He has done over 1000 client interviews and published over 10 articles on McKinsey.com and beyond. LinkedIn profile
Sources
- SurveyMonkey: AI Text Analysis & Sentiment Tools
- QuestionPro: Analyzing open-ended text data
- Alchemer: Open Text Analysis
- McKinsey & Company: Prediction: The future of customer experience
- ESOMAR Global Market Research 2025, via Research World: Inside the $153bn insights industry
- Miller, A. L. & Lambert, A. D. (2014). Open-Ended Survey Questions: Item Nonresponse Nightmare or Qualitative Data Dream? Survey Practice, 7(5)



