Every product page on ProductDome carries a scorecard, every category page opens in Recommended order, and sometimes both disagree with the raw Amazon stars sitting right next to them. That disagreement is the point. One shared ranking engine computes every score, order and badge on this site; here’s exactly what it does.
The problem with raw star averages
A product with three reviews averaging 5.0 looks better than one with five thousand averaging 4.6. It usually isn’t. Small samples are noisy and easy to seed with friendly reviews; large samples have survived years of real buyers, defective units, shipping damage, and expectation mismatches. Raw averages treat both as equally trustworthy. We don’t.
Bayesian smoothing, with the actual numbers
Our satisfaction score starts every product at a prior - the category’s average Amazon rating when the category has at least 8 rated products, otherwise a global 4.2 - and lets its own ratings pull it away as evidence accumulates:
smoothed rating = (rating × rating count + prior × 25) / (rating count + 25)
The prior weight of 25 acts like 25 phantom ratings at the category average (25 is roughly the lower-quartile rating count in our catalog). Below ~25 ratings the average dominates; by a few hundred ratings the product’s own rating does. A 5.0 with three ratings lands near the pack; a 4.6 with five thousand keeps its 4.6.
The ProductDome score: six components
The overall score (0-100, shown as 0-5 on product pages) blends six components with fixed, published weights:
- Owner satisfaction, 50% - the smoothed Amazon rating above. The closest signal to real owner outcomes, so it carries the most weight.
- Evidence strength, 15% - rating volume on a log scale, capped at 10,000 (so mega-listings can’t drown everything else).
- Buyer demand, 10% - Amazon’s “bought last month” figure, log-scaled and capped. Useful, but deliberately capped so hype can’t outrank quality.
- Value, 10% - price position versus the category, blended with satisfaction so “cheap” alone never wins. Only scored when we can display compliant, fresh pricing - currently we can’t, so value is excluded and the other weights renormalize.
- Data quality, 10% - does the product have a resolved brand, imagery, a description, and meaningful specs? Incomplete listings score lower.
- Freshness & availability, 5% - how recently we refreshed the product’s data, and whether the listing looks unavailable.
Missing data is handled honestly. A missing “bought last month” figure is treated as unknown, not as zero demand - the component is simply excluded and the remaining weights renormalize. Missing availability is never assumed to mean “in stock.”
Confidence is separate from quality
Every ranked product also gets a confidence level, shown wherever we make recommendations:
- High - 300+ ratings, fresh data, meaningful specs, resolved brand.
- Medium - useful but incomplete evidence.
- Limited evidence - fewer than 25 ratings, stale data, or incomplete product information. These products can still score well provisionally, but they’re labelled, and they can never take a top badge.
Badges are computed, never decorative
Badges are assigned per category by the same engine, one per product, and a category with no qualifying product gets none:
- Best Overall - the top-scoring product, high confidence required. Limited-evidence products are ineligible.
- Popular Pick - the strongest real demand (100+ bought last month minimum).
- Best Value / Premium Pick - require compliant, fresh pricing, so neither is currently awarded.
- Human-approved badges - a badge claiming human editorial selection can only ever appear with a genuine approval record behind it. We have none today, so you won’t see one.
Percentiles, not absolutes
A 4.3-star rating means something different in a category averaging 4.6 than in one averaging 3.9. The percentile chips on product pages (“better rated than Y%”, “more reviewed than Z%”) are computed against every model we track in that category. We only show a scorecard when the category has at least eight tracked products, so the percentiles mean something.
What the scorecard is not
It’s not a lab test - we say this on every page. We don’t own the products; we track their data at scale: specs, prices, ratings, rating velocity, brand footprint. The ProductDome score is our computed synthesis of that data and is always shown as distinct from the Amazon rating itself. And the one thing we never do is put Amazon’s star rating into search-engine schema as if it were our own verdict - visible on the page, labelled as Amazon’s, kept out of our structured claims.
Why this makes rankings feel “wrong” sometimes
If you sort our category pages by “Best rated”, you’re seeing the smoothed rating - so a shiny 5.0 with a dozen ratings sits below a workhorse 4.5 now and then. That’s the smoothing doing its job. The 5.0 might be genuinely excellent - and if it is, its score will rise as its rating base grows. Until then, we’d rather under-rank a promising newcomer than over-rank a padded one.