Think Build Implement Repeat
London, UK +44 7367 067226
WhatsApp FOLLOW f in X
  1. Home
  2. Blog
  3. Getting Real Themes Out of Review and Survey Text
AI Integration

Getting Real Themes Out of Review and Survey Text

Thousands of free-text comments and nobody reads them. How topic extraction works, why unsupervised themes disappoint, and the hybrid that earns its keep.

Updated 2 min readBy SpiderHunts Technologies

Free estimateNo obligation

Get a free estimate

Tell us what you need. A senior engineer reads every enquiry.

Takes under a minute. We never share your details.

  • Free consultation
  • No commitment
  • NDA on request

Prefer to talk? Book a free 30-minute call →

Quick answer — TL;DR

Fully automatic theme discovery produces clusters that are hard to act on. Defining the themes you care about and classifying against them works far better, with an unsupervised pass used only to find what you missed.

The pile nobody reads

Most businesses collect far more free-text feedback than anyone reads: review sites, post-purchase surveys, NPS comments, support conversations, cancellation reasons. It is the most direct information available about why customers behave as they do, and it usually sits unused.

The instinct is to run topic modelling across the lot and see what emerges. That produces something, but rarely something anyone can act on.

Why unsupervised themes disappoint

Unsupervised methods group text by statistical similarity, with no knowledge of what matters to your business. The clusters that come back tend to split on vocabulary rather than meaning, and mix issues you would never put together.

A typical output is a cluster containing delivery complaints, packaging complaints and a few comments about the courier's manner - related by words, but owned by three different teams with three different fixes.

They are also unstable. Re-run next quarter with new data and the clusters shift, so you cannot track whether a theme is growing.

Define the themes, then classify

The approach that works is less exciting and considerably more useful: decide what you need to know, then classify each comment against those categories.

  1. Read a couple of hundred comments by hand. This is not optional and it is always worth the afternoon.
  2. Write a theme list tied to who would act on each one - delivery speed, packaging condition, product fit, pricing, website usability.
  3. Label a sample against that list, checking that two people agree on the labels.
  4. Train a classifier, or use a language model with the definitions supplied, and measure it against the held-out labels.
  5. Track theme volumes over time, which is where the value actually is.

Sentiment alone tells you nothing

A sentiment score is not a finding. Knowing 22% of comments are negative does not tell anyone what to change, and the number moves for reasons that have nothing to do with your product.

Theme plus sentiment is actionable: negative comments about delivery timing rose this month while negative comments about product quality did not. That points at a specific team and a specific question.

OutputCan someone act on it?
Overall sentiment scoreNo
Sentiment by themeYes
Theme volume over timeYes
Theme by product or regionYes, and usually the most useful cut

Keep the unsupervised pass for discovery

Clustering still has a role: finding themes your list does not contain. Run it periodically over comments the classifier found hard to place, and read what comes out.

That is the right division of labour - structured classification for tracking, clustering for discovery. Used that way the awkward cluster is a feature, because it is where a new problem shows up first.

Comments nobody reads are not feedback. They are storage.

FAQ

Frequently asked questions

The questions readers ask us after this guide.

Still have a question?

Ask us directly — a senior engineer will get back to you.

Ask about your project

How many comments do I need?

For tracking themes, enough per theme per period that a change is distinguishable from noise. For building the classifier, a few hundred well-labelled examples per theme is a reasonable start.

Can a language model do this without training data?

Often yes, given clear theme definitions in the prompt. You still need a labelled sample to measure whether it is right, and to catch drift.

Should I analyse reviews from third-party sites?

If you can access them within their terms, yes - they often contain franker feedback than your own surveys. Keep the sources separate when tracking, since audiences differ.

What about comments in other languages?

Translate then classify, or use a multilingual model, but validate on each language separately. Accuracy usually varies more than expected.

Keep reading

More on AI Integration

Start here

Want machine learning project details from us?

Tell us what you are trying to predict and roughly what data you hold. We will come back with an honest view on whether machine learning is the right tool, what the work would involve and a realistic cost range. If a spreadsheet would do the job, we will say so.

  1. You tell us what you needTwo minutes on the form, or a message on WhatsApp.
  2. A senior engineer reviews itAnd comes back with questions, a realistic range and an honest view on fit.
  3. Free 30-minute scoping callWe talk through scope, options and a realistic estimate — with no obligation.
Free estimateNo obligation

Talk to someone who builds this

Send a short brief and we will come back with an honest view and a realistic range.

Takes under a minute. We never share your details.

  • Free consultation
  • No commitment
  • NDA on request

Prefer to talk? Book a free 30-minute call →