We Scored Ourselves Against 117 Rivals. Then We Changed One Column and Lost.

We ran the same comparison four times this week. Same 54 products on our side, same 117 rival entries on theirs, same afternoon. The first run had us winning 37 and losing 10. The second had us winning 36 and losing none at all. The third had us losing.
Nothing about any product changed between those four runs. We changed which column we read.
This is the second time we have opened up our own scoring in public. Earlier this month we published the fact that a single product of ours carries three different scores in three different fields, and that the three disagree. That article was about our side of the sheet. This one is about the other side, where it turns out we have been keeping three numbers per rival as well, plus a fourth copy in a different table that does not match.
Our comparison charts read one column out of six available, and that choice decides the outcome. Read our screening_score against the rival score and we win 37–4–10. Read curator against curator and we win 36 and lose nothing, because our curator scale only has three values in it. Read curator against the rival UX column and we lose, 7 wins to 11. The rival scores we published from are also a stale copy: 18 of 101 matched records disagree with the source table, and 17 of those 18 are higher in our copy than in the record. We are not fixing this by picking the flattering column. We are printing all four.
Three columns, three winners
Every product page carries a metafield called custom.amazon_benchmark, a small JSON array of the rivals we lined it up against. There are 117 entries across the catalogue. Twenty-five products have three rivals, sixteen have two, ten have exactly one, and three products have an empty array with no comparison at all.

Each entry carries three numbers, not one. There is score, which runs to a decimal place and averages 8.73. There is ux_score, whole numbers only, averaging 8.28. There is curator_score, also whole numbers, averaging 7.00. They are not copies of each other. Compare them pairwise and score differs from ux_score in 110 of the 112 entries where both exist, and from curator_score in 114 of 115.
Our side has two. screening_score averages 9.30 across the 54 products. curator_score averages 8.43. Those two agree with each other on exactly two products out of 54.
So there are six ways to answer one question. Here are four of them, each run over every product where the comparison is possible, our number against the best rival number in that product's set.
| Which columns | Products compared | Win / tie / loss | Average gap |
|---|---|---|---|
Our screening vs rival score
|
51 | 37 / 4 / 10 | +0.30 |
Our curator vs rival curator_score
|
51 | 36 / 15 / 0 | +0.92 |
Our curator vs rival ux_score
|
50 | 7 / 32 / 11 | −0.08 |
| Our screening vs the source table | 41 | 40 / 0 / 1 | +0.32 |
The version of this article we took down ran line one of that table and stopped there.
We published from the wrong copy
The rival scores live in two places. There is a metaobject table called competitor_product with 327 records in it, one score per record, no blanks anywhere. And there is the copy stapled to each product page inside amazon_benchmark. The second is supposed to be the first.

Match the two by product name and 101 of the 115 scored entries line up to a record. Of those 101, 83 carry the same number. Eighteen do not. Seventeen of the eighteen are higher in the copy on our product page than in the source table.
Cliganic Organic Jojoba Oil sits at 8.6 in the table and 9.2 on our page. CeraVe's Mineral Sunscreen Stick is 9.0 in the table and 9.5 on our page. Mighty Patch and Rael are both 9.3 in the table and both 9.5 on the page, which is the shape the errors keep taking: a row of rivals under one product all typed at the same round number. Our spot patch has three rivals and two of them are 9.5. Our wrapping mask has three and two of them are 9.5.
Every one of those seventeen errors runs against us. A rival scored half a point too high is a rival we are closer to losing to. Re-run the comparison against the source table instead, restricted to the 41 products where every rival matches a record, and the result is 40 wins, no ties, one loss. Our screening average holds at 9.30 and the rivals sit at 8.98. The single loss is our HIDIFF cleansing kit, 9.0 against 9.1.
That is the uncomfortable part. The bad copy was running against us in 17 cases out of 18, so correcting it makes us look better, and neither of those facts makes the number trustworthy. We also cannot tell you which of the two copies was written first. Neither table carries a timestamp, and we did not keep one.
Reproducibility is worse than the drift. Illiyoon's Ceramide Ato Concentrate Cream appears four separate times in the 327-record table, scored 9.5, 9.7, 9.1 and 8.8. Three of those four are the same shop. Dermalogica's Daily Microfoliant appears twice, both from Amazon, at 8.4 and 9.1. Aestura's Atobarrier365 Cream, twice, 9.7 and 9.2. When we score the same product on two different days we move it by up to nine tenths of a point.
Which is worth holding next to the headline we thought we had. Across all 327 records the Amazon average is 8.677 and the Olive Young average is 8.711. The gap between two entire retail platforms is 0.03. The gap between us and ourselves on one moisturiser is 0.9.
A three-point scale cannot lose
Look again at line two of that table, the one where we win 36 and lose nothing.

It is the most flattering of the four lines and it is worthless.
Our curator_score takes exactly three values across all 54 products. Twenty-seven nines, twenty-three eights, four sevens. That is the whole distribution. The rival curator_score spans six values from 4 to 9 and averages 7.00 against our 8.43. Two scales with different floors and different spreads, put side by side and declared a contest. Nobody loses a contest they are not on the same axis for.
Line three fails the other way. The rival ux_score only ever takes 6, 7, 8 or 9, and our curator column only takes 7, 8 or 9, so 32 of the 50 comparisons come out as exact ties. We win 7 and lose 11 in what is left, and the average lands at −0.08. That is not a finding about product quality. It is a finding about two coarse rulers.
This is not a problem we invented. Tamblyn and colleagues took 3,156 grant applications to Canada's federal health funder and had reviewers both score and rank the same proposals. Inter-rater agreement moved with the format: an ICC of 0.54 for rating against 0.59 for ranking in the first phase, and 0.25 against 0.38 in the second. The interpretation bands move too. Hallgren's tutorial reports Cicchetti's cutoffs, where "good" reliability starts at 0.60. McHugh, writing about kappa, quotes the conventional table in which 0.41–0.60 reads as "moderate" and then rejects it outright, arguing that anything below 0.60 means "little confidence should be placed in the study results." Same coefficient. Two verdicts.
We read 33 ingredients of ours and 6 of theirs
Here is the asymmetry underneath all four lines of that table.
When we score one of our own products we read the entire ingredient list. The median is 33 entries, the mean 34.2. The shortest is Pure Somme's jojoba oil at one ingredient and the longest is our spicule set at 87.
When we score a rival we read six. Not roughly six. Exactly six, in 88 of the 117 entries, which is 75.2 percent of them. Two entries have none at all. The handful that run longer reach 47 and 44, so the field can hold more than six and mostly does not. That is a cap, not a count.
A six-ingredient reading is a reading of the top of the label, and the top of the label is where the water and the humectants are. In the United States the order is set by 21 CFR 701.3, which requires ingredients "in descending order of predominance" but lets everything at one percent or less be listed "without respect to order of predominance." Most actives in a serum sit under that line. We have written up what a position on an ingredient list can and cannot tell you about concentration, and the answer was: less than you would like. Reading six of them tells you less still.
The EU's common criteria for cosmetic claims put it more sharply than we would have. Claims "shall not attribute to the product concerned specific (i.e. unique) characteristics if similar products possess the same characteristics," and where studies are used as evidence they must follow methods that are "valid, reliable and reproducible." A rubric that reads 33 ingredients on one side and 6 on the other is not reproducible in the direction that matters.
Three of our 54 products have no rival on file, so they are simply absent from every line of the table. Ten more are being compared against a single product. We have not gone back to check whether that one product was the obvious rival or just the first one somebody found.
What our screening found
The honest summary of our own benchmark is that we win under three of the four framings and lose under one, and that the framing we lose under is the one where both rulers are too coarse to say anything. The framing we win biggest under is worse than useless, because our curator column has three values in it. What survives is the first line and the fourth: reading full ingredient lists against a truncated six-item reading of the competition, we come out ahead by roughly three tenths of a point. We think that gap is real and we think it is smaller than it looks, because we read our own products more carefully than we read anyone else's. Under United States rules, an officer or manager writing a review of their own company's product has to disclose the relationship. We are not writing reviews, we are publishing a rubric, and the same logic applies: the person holding the ruler is selling you the thing being measured. Here is the ruler.
Gyeol Haus Skin Barrier Jojoba Oil
Four of its eleven ingredients are flagged mid risk and all four are fragrance material: jasmine oil, ylang ylang, linalool and benzyl benzoate. It is also the only product in our catalogue carrying a Preservative-free claim, which we checked and which is true. We put it here because we lose with it: 8.7 against Cliganic's 9.2. Except Cliganic is 8.6 in the source table, which would make it a win. Both of those sentences are our own data.
The practical takeaway
When a brand shows you a comparison chart, the number on it is the last thing that happened. Before it there was a choice of which field to read, a choice of how many ingredients to read on each side, and a choice of which rivals to put on the shelf next to it. All three of those choices were made by the party that benefits.

Ask the small questions. How many ingredients did they read on the other product. How many rivals, and who picked them. Is the scale the same scale on both sides, or is a five-point house rating being weighed against a decimal one. Thompson and colleagues went through 100 best-selling facial moisturisers and found that the marketing category alone tracked both price and rating: anti-aging averaged $14.99 an ounce and expert-approved $5.91, while fragrance-free products averaged 4.35 stars out of five and natural ones 3.49. The label you assign a product moves the numbers attached to it before anyone opens the bottle.
We are keeping the 327-record table as the record from here and treating what sits on the product page as a copy. We are going to widen the rival readings past six ingredients, which will probably narrow our lead. And the four-line table above stays, because a benchmark that only shows you its best framing is an advertisement wearing a lab coat.
You can see the whole scoring method, and what it does and does not check, on our awards and screening page. The earlier pieces on the three scores we keep on a single product and on what a position on an ingredient list can tell you are the two halves of this one. If you would rather watch a number get taken apart in a smaller room, the price-per-millilitre piece does it with arithmetic you can check on a shelf.
- R. Tamblyn, N. Girard, J. Hanley, B. Habib, A. Mota, K. M. Khan, C. L. Ardern, "Ranking versus rating in peer review of research grant applications," PLOS ONE 18(10):e0292306 (2023). link
- K. A. Hallgren, "Computing inter-rater reliability for observational data: an overview and tutorial," Tutorials in Quantitative Methods for Psychology 8(1):23–34 (2012). link
- M. L. McHugh, "Interrater reliability: the kappa statistic," Biochemia Medica 22(3):276–282 (2012). link
- A. M. Thompson, B. Kromenacker, T. Y. Loh, C. M. Ludwig, R. Segal, V. Y. Shi, "Allergenic potential, marketing claims, and pricing of facial moisturizers," Dermatology Online Journal 26(7) (2020). link
- US FDA, 21 CFR 701.3, "Designation of ingredients." link
- US FTC, 16 CFR Part 465, "Rule on the Use of Consumer Reviews and Testimonials," §465.5 (insider consumer reviews and testimonials). link
- Commission Regulation (EU) No 655/2013, common criteria for the justification of claims used in relation to cosmetic products, Annex, criteria 3 and 4. link

