Whether on Amazon, Booking.com or Google Reviews, star ratings are now among the most important factors helping consumers make decisions when shopping online. Platforms summarise the reviews of thousands of customers into a single metric, usually the arithmetic mean – in other words, the classic average. But does this calculation actually reflect the way in which consumers perceive and process rating distributions? A new study by an interdisziplinär research team from Paderborn University and Ludwig Maximilian University of Munich now provides experimental answers to this question for the first time. The findings have been published in the international journal ‘PLOS ONE’.
Previous research in this field has primarily focused on whether the arithmetic mean is a useful indicator for purchasing decisions. However, the aggregation principles that consumers actually apply when looking at rating distributions – that is, how and in what way they take these into account – have remained largely unexplored. Do they really use the simple average, or do they give greater weight to certain rating categories than others, for example by focusing on particularly negative or positive ratings?
To answer this question, the economists and Computer Science experts conducted a controlled laboratory experiment with 107 participants. The participants were presented exclusively with the rating distributions for three products at a time and had to rank them according to their personal preferences without knowing the product names or any other characteristics. An incentive-compatible study design ensured that the participants revealed their true preferences: the higher they rated a product, the more likely they were to receive it. The ranking decisions obtained in this way were subsequently analysed using various statistical models and compared with different theoretically derived aggregation functions.
The results clearly show that it is not only the average rating of a product that is relevant to potential customers’ purchasing decisions. Whilst the majority do indeed aggregate rating distributions in line with the arithmetic mean, more than 40 per cent of participants systematically deviate from this. A significant proportion follow a binary strategy, distinguishing only between positive (4–5 stars) and negative (1–2 stars) ratings, whilst the middle category (3 stars) is largely ignored. Other groups focus predominantly on negative ratings, particularly 1-star ratings. It is also noteworthy that none of the participants used the median as an aggregation principle, even though this is considered particularly robust in the theoretical literature. These patterns proved to be stable. They persisted regardless of whether, in addition to the graphical rating bars, the participants were also shown the percentage shares of the individual star categories and the average value, and were not significantly influenced by individual characteristics such as age, gender or online shopping experience.
“The results suggest that the common practice of summarising product ratings exclusively using the arithmetic mean does not, in fact, provide the optimal basis for decision-making for a significant proportion of consumers. Platform operators could benefit from offering customisable aggregation options,” explains Dr Benhud Mir Djawadi, former head of the ‘Business and Economic Research Laboratory’ at Paderborn University. This would allow consumers to choose whether they prefer a traditional average rating, a metric focused on negative experiences, or a simplified positive-negative breakdown. “As all the necessary rating data is already available, such adjustments would be technically feasible and could lead to more satisfied customers and better purchasing decisions,” says Mir Djawadi.
The study “Aggregation Processes in Customer Rating Systems — Insights from an Economic Decision Experiment” by Dr Dirk van Straaten, Dr Behnud Mir Djawadi, Vitalik Melnikov, Prof. Dr Eyke Hüllermeier and Prof. Dr René Fahr is available online. The publication was funded by Paderborn University’s Open Access Publication Fund and by the German Research Foundation (DFG) as part of the Collaborative Research Centre (CRC) 901 ‘On-the-Fly Computing ’.