Review analysis

How to Read Customer Reviews Properly

Using 819,253 WordPress plugin reviews, this article shows which reviews are worth reading and how to turn them into useful customer and product insights.

Start reading
Long-form research 22 min read WordPress plugin reviews
819,253 reviews analysed
24,871 rated plugins
82.5% average 4–5 stars
10 reusable techniques

Why Most Reviews Are Noise

Review analysis sounds like something you need a very large dataset to do. Ironically, the main point I will be making with this one is that you don't really need one of your own. To put it bluntly (as I will do throughout this), a lot of customer reviews are noise and review datasets are inherently biased. In this corpus, 34.8% add no useful detail beyond the rating.

However, with the proper lens, you will be amazed at the insights you can find: customer pain points, what customers actually want and the language they use to describe it. Everything I'll show you works for someone testing the waters for a start-up idea, right up to a large enterprise looking for extra insight to help build its product roadmap.

It doesn't matter whether you offer a service, software or physical products. They don't even have to be your own reviews: looking across several competitors can give you fantastic landscape insights.

The Review Analysis Dataset

Using public WordPress.org plugin review pages, I built a database containing 819,253 reviews for this review analysis. These approaches are agnostic, so I will pull in other public sources along the way.

Why Reviews Follow a J-Shaped Distribution

Review distribution by star rating
01Review distribution by star rating

For those of you not familiar with looking at reviews, welcome to the phenomenon known as the "J-shaped distribution". If you are looking at the chart thinking, "Hey, this looks like my great review profile," I'm afraid you are in the majority. The act of submitting a review is inherently biased.

First, only people who tried the product can review it, and that group was already willing to give it a chance. Then, among those who did try it, people with moderate opinions are less likely to bother. You come back to brag or to moan, not to say, "It was fine." Hu, Pavlou and Zhang describe these as acquisition bias and underreporting bias.

The result is a distribution dominated by five-star praise, a smaller spike of one-star anger and very little in between. Not because every product is amazing or terrible, but because the people leaving reviews are not a random sample. I'd better back that up quickly.

Trustpilot's transparency report shows the same broad shape: 73% of reviews were five-star, 14% one-star, and only 13% fell across ratings two to four. That makes 81% four- or five-star.

So do we see it in our dataset? Unsurprisingly, yes. 711,015 of 819,253 reviews are five-star (86.8%).

At plugin level, 20,507 of 24,871 listings with at least one rating average between 4.0 and 5.0 stars (82.5%). Different denominator, same heavily positive shape.

Average plugin rating distribution
02Average plugin rating distribution

I believe reviews have fundamentally split into two brackets. Some tools, predominantly in ecommerce, focus on descriptive information to help the decision-making process.

The much larger bracket uses reviews as social proof. Where do returns diminish here? Do 10,000 reviews at 4.5 stars have the same impact as 100,000? Sadly, I don't have the dataset to test it, but my instinct is there is a point where gathering more 4.5-star reviews no longer adds much as a signal.

Where reviews sit between social proof and product insight
03Where reviews sit between social proof and product insight

Obviously social proof and trust are fundamental. But the very achievable quest for a good score can cause us to overlook a rich vein of insight. So here is my take on how to squeeze every drop of value out of your reviews.

What Each Star Rating Can Tell You

1 & 2 stars - Obvious, but this is where the catastrophes live. One-star reviewers can be so blinded by rage that they focus on anger rather than critique, so look for the event underneath it. Two-star reviews are richer: they are the longest in the corpus, and 24.6% run past 500 characters. They are useful for reconstructing what failed, what the user tried and what they chose instead.

3 stars - This is where the "howevers" start to appear: credit first, then the missing feature or awkward bit. In the WordPress dataset, 29.0% are feature requests. This is where you find a more nuanced assessment of a product's current state.

4 & 5 stars - Four-star reviews are particularly strong for feature requests: 36.8%. Five-star reviews, despite their highly desired status, contain a lot of noise: 38.5% say nothing more specific than generic praise. However, the longer five-star reviews can be perfect for snagging copy. I'll cover that later.

That gives us a useful starting point, but star rating alone is a blunt filter. The next job is finding the reviews within each bracket that are actually worth reading.

How to Find Customer Pain Points in Reviews

"Losses loom larger than gains." - Kahneman & Tversky, 1979.

Prospect theory tells us that losses tend to loom larger than gains. That does not prove loss aversion is driving our reviews, but we may be seeing a similar imbalance in effort.

In this dataset, one-star reviewers use 313 characters on average to describe their suffering, against 144 for five-star praise - about 2.2 times longer. Five-star language is also more repetitive: its top 50 words make up 27.1% of all text, compared with 20.9% in one-star reviews.

Review length and vocabulary concentration by star rating
04Review length and vocabulary concentration by star rating

Spotting Recurring Review Themes

From there, a word cloud can give you a quick top-level view of repeated language. It is not proof of a theme, but in this dataset one concern jumps out: support is paramount (shock horror).

Common words by star rating
05Common words by star rating

Do you see words like "however", "issue" and "problem"? That should prick up the ears of anyone working in product (or trigger anxiety).

A basic word match tagged 88,076 reviews - 10.8% of the dataset - as candidates to inspect. The words do not prove there is a product problem; they simply narrow the pile.

Reviews containing “however”, “issue” or “problem”
06Reviews containing “however”, “issue” or “problem”

Signal Words That Reveal Useful Reviews

Signal Word Example 1 Example 2
however [3 star] ..."This plugin works perfectly… However, after using it for years I am now leaving it. All new features are pay-features and it will constantly add issues …to force you to pay"... [4 star] ..."It’s a greatly helpful and easy-to-use plugin. However, there are things that can improve the experience: some readability features are not detected in some languages"...
issue [1 star] ..."A serious security issue was discovered two weeks ago, leaving it vulnerable. No updated version has been released yet.…"... [1 star] ..."I just spent 2 hours with a team member getting to the bottom of an issue where this plugin was removing backslashes from edited files on save"....
problem [1 star] ..."Been using this really good plugin for about 2 years. As of v4 it has caused more problems than it’s worth. Widgets disappearing has been my primary concern…."... [2 star] ..."The problem here appears to be the plugin using an old version of ajax-admin. Two requests take an average of 6-10 seconds. Disabling the plugin gave me a performance grade of A-95 and load time of 465ms"...

This is hardly scientific, but it has already produced six leads on bugs, features or customer pain points. Any of them could come down to opinion or user error, so at this stage I would give each one a descriptive tag, then comb through the wider set and tally what recurs.

This is where your own knowledge matters. Look for other signal words and keywords that are important in your market.

AI can accelerate this (I'll cover that later), but I recommend learning to flex this muscle first: thrust without control over direction rarely ends well.

To stay true to the goal of this project, I am going to jump out of the plugin space and pick a random product in my room to put this to the test.

Seeing This in a Real Review

I landed on Fussy, a product whose exceptionally clever and well-designed advertising inspired me to buy it and see what all the fuss(y) was about.

Fussy’s rating distribution on Trustpilot
07
Fussy’s rating distribution on TrustpilotSource: Fussy on Trustpilot ↗

Unsurprisingly, it shows the standard J-shaped distribution of product reviews. I do not need anything sophisticated or technical here. I can literally search for "however", "problem" and "issue".

Fussy reviews containing “however”
08Fussy reviews containing “however”

From just two examples, I already had a couple of comments:

Travel-sized packaging

Courier options

Looking at a small handful of results for each signal word gave me a longer list of leads. This is an illustration, not a representative sample:

Signal Word Findings
however [5*] Travel sized packaging needed [3*] Delivery not trackable, customer experience team gives no real answers [3*] Refills arriving dented or crushed for third time in a row [5*] All 3 scents very faint, disliked them all [4*] Last order didn’t arrive, previous order incomplete [3*] No financial incentive to buy direct vs supermarkets [4*] Delivery too slow, please change supplier [3*] Will never be 5 stars whilst EVRi is the courier [2*] Holder awkward to use, doesn’t stand upright, scent doesn’t last a working day [5*] Couldn’t add limited edition case to existing subscription
problem [1*] Nearly 2 weeks no delivery, automated response doesn’t address the problem [1*] Rancid grease smell at end of day (shea butter / coconut oil) [1*] Metallic odour left in clothes, impossible to remove, taints other stored clothes [1*] Clothing staining ruined 3 t-shirts and bedding, cancellation form lists it as a known reason [1*] Developed sores under arms [2*] Deodorant causes stains on shirts that don’t wash out [2*] Case mechanism breaks turning action stops working after adding refills [2*] Clumps come off stick with unshaved underarms problem you don’t get with roll-ons [3*] Doesn’t work in hotter weather fine in winter, unreliable in summer [3*] Case not accessible for people with arthritis nothing to grip [3*] Subscription replaces chosen scents with random alternatives when out of stock [3*] Applicator mechanism doesn’t move the stick up [5*] Case design unclear which end is pull and which is twist
issue [5*] Delivery issue with refill resolved by Sadie immediately [5*] Had an issue, resolved straight away still smelt fresh after tough bike ride [5*] Delivery issue sorted quickly [4*] One scent didn’t go with skin money credited within a day [2*] After 3 days not impressed, marketing very misleading, different body types not mentioned [5*] Delivery issue resolved by Vivienne professionally [4*] Delivery took a week via Royal Mail only reason deducted a star [5*] Delivery issue, team proposed alternative quickly [5*] No body odour for two days with single application [5*] Order issue dealt with fast and responsive

I would then turn those leads into themes and begin tallying. As you scale up, clusters start to form. In this tiny sample, they looked like this:

Themes Count
Case design issues 6
Delivery delays / missing orders 5
Delivery issues resolved by support 5
Scent doesn’t last 4
Clothing staining / odour transfer 4
Subscription issues 2
Skin reactions 1
Doesn’t work in hot weather 1
Refills damaged in transit 1
No incentive to buy direct 1
Travel sized packaging needed 1
Marketing misleading 1

And yes, the same trick works with services.

A MoneySuperMarket review containing “however”
09A MoneySuperMarket review containing “however”

How to Do Sentiment Analysis on Customer Reviews

Put simply, review sentiment analysis reads text and labels it positive, negative or neutral. The three common routes are:

An off-the-shelf library such as VADER or TextBlob

An LLM asked to classify the text

A model trained on your own data

To be perfectly honest, I can only see much use for basic sentiment analysis when feedback does not already come with a star rating. If the review has one star or five, you already know the broad direction. The more useful question is what the customer liked or disliked inside the review. That is where aspect-based sentiment analysis (ABSA) earns its keep.

What Is Aspect-Based Sentiment Analysis?

Simply put, it chunks one review into the separate things the customer mentions.

Take a customer review. ABSA separates the features, parts or topics it names, then gives each one its own positive, negative or neutral label.

Instead of one label for the whole review, you get a list of labelled pieces.

The review: "Easy to install and the support team is brilliant, but it slowed my site to a crawl after the latest update."

What ABSA returns:

Setup: positive

Support: positive

Performance: negative

Update: negative

Each piece is an aspect, hence the name. In one review we have gone from one star rating to four pieces of product insight.

It sounds great, and it can be. But none of these advanced techniques will simply fit your data out of the box. They need human adjustment. Pattern matching is weak on context, and a model is not automatically a subject-matter expert.

ABSA With Trained Models

This is the gold-standard approach, but it is not entry level. It needs machine-learning talent, clear labelling guidelines and enough budget to create a substantial training set. A well-trained BERT-based model can combine aspect extraction with sentiment classification at scale: humans label thousands, then the model labels millions.

I could not reproduce that approach here (and definitely could not afford it), but Wayfair's ABSA case study is a useful real-world example. Its team used a BERT-based system to identify product aspects and sentiment, then used the results to build long-tail shopping pages.

ABSA With AI

Fortunately, we can build a rougher version with an LLM. For this experiment I used the WordPress review data and the OpenAI API. The scripts defaulted to gpt-4o-mini when I ran them; treat that model and the cost as historical because API availability and pricing change.

I started with a popular WordPress plugin that bundles dozens of features into one install. It is polarising: across 2,387 reviews, 24% were one-star, 60% five-star, and the average was 3.75.

ABSA Pre-Labelling

The quick version lets the model invent aspect names as it goes. The script sends one review at a time to gpt-4o-mini, asks for aspect-sentiment pairs, and writes one CSV row for each pair. That makes it easy to start, but messy to aggregate: "support", "customer support" and "help desk" can become separate labels.

It is not entirely useless. You can still see how users respond to updates or customer support, and start benchmarking mentions of errors and bugs. Just do not pretend that every free-form label is directly comparable.

Free-form aspect mentions split by positive, neutral and negative sentiment
10Free-form aspect mentions split by positive, neutral and negative sentiment

ABSA Post-Labelling

This is the more involved version. Instead of letting the model invent labels for each review, we build a category list first and tag every review against that fixed list. Three scripts do three jobs: discover the categories, classify the reviews, then drill into one category to find its specific sub-issues.

A Practical Example from the WordPress Dataset

To put it through its paces, I picked another popular WordPress plugin. It has a different shape from the multi-feature one above: single-purpose, narrow in scope, but with a long track record and 2,156 reviews.

The discovery script samples reviews in batches of 100, asks the model what categories it sees, and stops when three consecutive batches add nothing new. A final cleanup pass collapses near-duplicates. It returned 14 candidate categories.

Three needed a manual touch before classification:

Dropped 'Functionality' because it was too broad and overlapped Site Performance and Plugin Stability.

Merged 'Integration' into 'Plugin Compatibility' to create one broader category for third-party interactions.

Renamed 'User Expectations' to 'Setup & Onboarding' because the scope was sharper and clearer.

That left 12 categories. Six are shown below; the full breakdown comes later. This step matters. In this run, the model got me roughly 90% of the way. The last 10% was judgement about overlap and naming, and skipping it would have made the downstream analysis noisier.

Category What it covers
Form Customization Comments related to building and customizing forms: styling, field configurations, layout, and the form-creation experience.
Setup & Onboarding Comments about installing, configuring, and getting the plugin working for the first time, including the learning curve and initial expectations.
Captcha Integration Comments about the captcha feature, including reCAPTCHA versions, configuration, and effectiveness against spam.
Email Delivery Comments about email delivery from the contact forms: whether messages send, arrive, get marked as spam, or fail silently.
User Interface Comments about the design and usability of the plugin's admin interface, including form-builder navigation and overall aesthetic.
Documentation Comments about the clarity, completeness, and usefulness of the plugin's documentation, including troubleshooting guides.

The classification script then tagged every review against the taxonomy. 1,741 of 2,156 reviews (80.8%) touched at least one category. The other 415 (19.2%) were mostly generic praise such as "love it" or "easy to use", with no specific aspect attached. That is a finding in itself: nearly one in five reviewers did not explain why.

Another 127 reviews (7.3% of the classified set) contained both positive and negative categories. That is the per-aspect view doing real work, splitting apart detail that a star rating cannot.

Plotting mention count against percentage negative gives the chart below. Far right means heavily discussed; higher means more negative. The top-right quadrant is a candidate priority area because complaints are both frequent and negative. It is a place to investigate, not an automatic product roadmap. The bottom-right shows the high-volume areas where users volunteer praise.

Aspect volume and negative sentiment priority map
11Aspect volume and negative sentiment priority map

Captcha Integration, Updates and Email Delivery sit in that top-right quadrant, each with more than 60% negative mentions. These are frequent, negative clusters worth investigating first.

Form Customization, User Interface and Setup & Onboarding sit in the bottom-right strengths zone: high volume and mostly positive. This is what the product appears to be winning on.

Full Breakdown

Category Mentions % positive % neutral % negative
Form Customization 668 82% 6% 11%
Plugin Stability 619 49% 1% 48%
User Interface 364 75% 4% 19%
Setup & Onboarding 324 69% 7% 22%
Captcha Integration 256 20% 5% 73%
Plugin Compatibility 237 65% 10% 24%
Updates 223 16% 4% 79%
Support Quality 213 40% 6% 53%
Email Delivery 145 15% 4% 79%
Spam Experience 100 10% 5% 85%
Documentation 97 44% 8% 47%
Site Performance 70 42% 4% 52%

The quadrant tells you where to spend time. To understand what is actually going wrong inside a category, you drill in. I picked Captcha Integration because it sits firmly in the priority quadrant and is focused enough to break down cleanly.

The drill-down repeats the same approach inside one category. It extracts a free-form sub-aspect from each Captcha Integration review, consolidates those phrases into a smaller taxonomy, then classifies the reviews against it. Eight sub-aspects emerged.

Captcha Integration sub-aspects
12Captcha Integration sub-aspects

One dominates: reCAPTCHA Versions, with 195 mentions and 83% negative sentiment. That is the core of the Captcha Integration pain. Three snippets from those reviews:

… Go all the way down, and opt for an earlier version otherwise it's the cat and do not update this plugin as long as the author does not reconsider his decision of reCaptcha V3 that just does not work (empty space). …
… This used to be my go to contact form. However, after the most recent update it no longer works with Google Recaptcha, and forms fail to be submitted. …
… Three stars because this plugin has worked for so well for so long. However, the decision to for users to transition to reCaptca V3 was a bad one. It's way too heavy and causes my site to become unresponsive. Have downgraded to the last release version. …

Beyond the headline, the other negative sub-aspects - Integration Issues, Spam Protection, Support and Documentation - tell a similar story. The fix list is no longer abstract. There are concrete, named issues to work through.

The more involved workflow has four stages: discover a candidate category list, clean it manually, classify every review against it, then drill into the high-volume negative categories. A final script combines the outputs into a report.

How to Find Swipeable Copy in Customer Reviews

Good copy and taglines will never fix a bad product or service. But they can have a significant effect on a good one. What you are looking for is a line that hooks someone - and I do not mean that in showbiz marketing slang.

I mean a line that resonates because it captures a reader's desire or anxiety in that situation. You want a small extra moment of consideration: a nudge further down the funnel, or a line memorable enough to return when the time to buy arrives.

Test them in ads, meta titles and landing-page copy. You can also place the original review next to the claim, so the customer's words do the proof.

There are two angles to look for in your reviews:

Good reviews are direct swipes. The customer has already written your headline. Your job is to spot it and steal it.

Bad reviews are inverted swipes. The pain a customer describes becomes the positive promise your product makes. "I wasted four days fighting bugs" becomes "Stop wasting days fighting bugs".

Four Ways to Find Useful Customer Language

This is a manual hunt. You could layer in a script, but a human still has to vet the result. Here are some clues for where to look:

In Five-Star Reviews, Start with the Longer Ones

This limits the amount of noise you have to get through:

★★★★★

"Believe me! [Brand Name] is the only service you need to secure your site completely... It never affects your site speed... It completely works offsite. Reliable backups!!! ... I might have performed restoring backups more than 1000 times with [Brand Name], with a 100% success rate."

In Five-Star Reviews, Look for Capitalisation

★★★★★

"a task that was so overwhelming and terrifying has come together using simple, repeatable steps that give BEAUTIFUL results"

Look for Before-and-After Arcs

Use the signal-word tricks again and look for phrases such as "I was". My top tip is to write a few small theoretical reviews yourself and note the words you naturally reach for: "tried", "I was", "searching", "looking".

Signal phrases reveal before-and-after arcs
13Signal phrases reveal before-and-after arcs

Look for Quantification: Hours, Days, Weeks

"Every time a add, edit or cancel anthing on a cell, it takes 45 seconds or more to respond. It has taken me over 1 day of work to do a simple 303 table"

Inverse swipe:

Edit hundreds of rows in seconds, not days.

These are the four placements I would test: a Facebook ad, a Google search ad, a product listing card and a landing-page testimonial.

Using Customer Feedback Analysis to Understand a Market

The useful question is what to do with customer feedback once it covers multiple vendors. At that point, customer feedback analysis offers a chance to understand a market rather than one product. But comparisons only make sense where the products share a job. Every item here is a WordPress plugin, but a caching plugin should not be judged against one that sends calendar invites. For more granular insight, group products by the function they perform.

How to Do N-Gram Analysis on Reviews

N-grams are a basic way to see what customers mention frequently. They lose most of the context, so they are better for orientation than diagnosis. They can show the vocabulary worth investigating and the terms you may want to include or exclude from later processes.

This frequency pass used 817,999 non-empty reviews. It did not apply a corpus-level language filter, so obvious names and foreign-language fragments still need manual cleaning.

What Frequency Can Tell Us

If users keep talking about something, good or bad, it is probably notable to them. In this pass, 26 single words appeared in the top 50 of both one- and two-star reviews and four- and five-star reviews.

Shared vocabulary in positive and negative reviews
14Shared vocabulary in positive and negative reviews

Support is king. It ranked first in both brackets: 16,609 mentions in one- and two-star reviews and 126,185 in four- and five-star reviews.

Free versus freemium. "Free version" ranks near the top of both brackets, while "paid version" and "premium version" rank far higher in negative reviews. People love a free plugin, but dislike being pulled into a freemium product.

Time is precious. "Time" ranks fourth in negative reviews and eighth in positive ones: people describe either time wasted or time saved.

"It had better work." The bigram "not work" appears 7,975 times in one- and two-star reviews, against 5,113 in four- and five-star reviews.

Strip out the words that rank high on both lists and what remains is the diagnostic vocabulary of unhappy reviews. Three themes dominate: broken updates, technical errors and the feeling of having wasted money (as we all know).

What’s uniquely bad
15What’s uniquely bad

Ease of use, speed and usefulness dominate the positive vocabulary.

What’s uniquely good
16What’s uniquely good

This is easy to do, and there is value in understanding the dictionary of reviews. It can help build the backbone of later processes: which terms to include, exclude or investigate. But it remains a limited view because frequency cannot tell you the full context.

Combining ABSA and LLMs for Market Analysis

We have already seen how ABSA plus an LLM can turn one product's reviews into named aspects. For a market view, the first rule is stricter: compare products that do the same job.

For the worked example, I selected 15 WordPress speed, cache and performance-optimisation plugins, covering 24,024 reviews. The category-discovery sample contained 900 reviews.

Discovery did not converge automatically. I manually reviewed the taxonomy, merged overlaps, removed sentiment-loaded names and finished with nine neutral aspects: performance impact, feature capabilities, configuration experience, ease of use, compatibility and stability, support quality, documentation quality, advertising and branding, and commercial terms.

Only after that did I run a 900-review proof of concept, inspect samples by aspect and plugin, and move to the full cohort in 5,000-review increments with QA between each block.

The final run classified 24,024 of 24,024 reviews with no API errors. It produced 26,198 aspect mentions. Performance Impact led with 7,934 mentions, followed by Feature Capabilities with 5,269 and Support Quality with 3,760.

That is not the same as 100% classification accuracy. After a recovery pass, 9,385 reviews (39.1%) still had no aspect. Most were very short praise such as "great plugin", so the method probably understates thin positive feedback. The output passed QA for exploratory analysis and charting, not publication-grade measurement.

The point of the landscape is not to create one magic winner score. It lets you compare competitors on the things customers actually mention: speed, setup, stability, support, commercial terms and the rest.

That turns one solid pile of good or bad into a map of where each product is strong, where complaints cluster and where a gap might be worth investigating.

The useful output is a set of aspect-level comparisons, not a league table pretending every difference is equally important.

Used carefully, this gives you a landscape of customer concerns rather than a ranking built from star averages alone.

Landscape analysis: recurring customer concerns
17Landscape analysis: recurring customer concerns

Conclusion: Turning Review Analysis Into Product Insight

The point of good review analysis is not to replace reading reviews with a model. It is to stop spending equal attention on every review when so many say almost nothing.

Start manually. Learn the star brackets, use signal words, spot the before-and-after arcs and build your own judgement. When the volume becomes too large, n-grams and ABSA can help you decide where to look.

Reviews are biased, noisy and overwhelmingly positive. They are also full of bug reports, feature requests, anxieties and copy written in the customer's own language. The trick is not to trust the average. It is to squeeze every useful drop from the words underneath it.

References

Kahneman, D. & Tversky, A. (1979). Prospect Theory: An Analysis of Decision under Risk. Econometrica, 47(2), 263-291.

Trustpilot (2024). Transparency Report 2024, p. 15.

Hu, N., Pavlou, P. A. & Zhang, J. (2017). On Self-Selection Biases in Online Product Reviews. MIS Quarterly, 41(2), 449-471.

Saad, A. & Mosse, G. (2022). Wayfair's New Approach to Aspect-Based Sentiment Analysis Helps Customers Easily Find Long-Tail Products.

OpenAI. Developer Quickstart.

Insights live · tools in development

Ready to
play
FORTE?

Ask about the research, request early access to a tool, or tell me what you are working on.

Your details are used only to respond to this enquiry. Privacy notice.