Ranked all 571M Amazon reviews from 2023 by category profanity rate. Video games is 6× the cleanest category.
Our take
The recent analysis of Amazon's vast repository of reviews, as highlighted in the article discussing the McAuley Lab's 2023 dataset, unveils intriguing insights into consumer behavior across various product categories. With over 571 million reviews analyzed based on factors such as profanity, capitalization, and punctuation, the findings reveal not only the emotional intensity behind products but also the distinct characteristics that different categories elicit from consumers. For instance, video games emerge as the most expressive category, with a striking 6.54% of reviews containing strong profanity, reflecting a passionate engagement often found in gaming communities. In contrast, categories like gift cards and handmade items exhibit significantly lower profanity rates, underscoring the difference between cultural engagement and transactional utility. This disparity raises important questions about how product nature influences consumer expression, a theme that resonates with ongoing discussions about user feedback in platforms like Conditional formatting for specific character count or Does anyone have issue of stock prices stopped updating?.
One of the most striking revelations from the dataset is the phenomenon of the "angriest category"—subscription boxes. With nearly 16% of reviews rating these products at one star, it's clear that the expectation of curated surprises often leads to disappointment. This sentiment raises significant implications for businesses operating in this space, highlighting the necessity for transparency and accuracy in marketing these products. As consumers increasingly seek novel experiences in their purchases, brands must navigate the fine line between excitement and regret, ensuring that they meet or exceed consumer expectations. This tension mirrors broader conversations in the digital landscape about user experience, as seen in discussions about AI's impact on workflows in articles like Your AI Use Is Breaking My Brain: Why 10 Minutes of Prompting Fries Us[D.
Methodologically, the approach taken to analyze this immense dataset is commendable for its clarity and reproducibility. By employing a straightforward rule-based system rather than relying on complex models, the findings are transparent and accessible. This choice reflects a broader trend in data analysis where simplicity often leads to greater understanding and usability. However, it also invites discussion about the limitations of such an approach, particularly in its English-only scope and the potential for misinterpretation through quoted titles. As the landscape of data continues to evolve, the challenge remains to balance thoroughness with accessibility, ensuring that insights gleaned from datasets can inform and empower users without overwhelming them with complexity.
Looking ahead, the implications of these findings extend beyond mere consumer sentiment. They prompt a deeper exploration into how brands can better engage with their audiences by acknowledging the emotional undercurrents tied to their products. This analysis serves as a call to action for businesses to not only refine their marketing strategies but also enhance their product offerings based on consumer feedback. As we continue to dissect the nuances of online reviews, one cannot help but wonder: how will the evolving landscape of consumer expectations shape the future of product development and brand loyalty in an increasingly digital marketplace?
I read the McAuley Lab's full 2023 Amazon Reviews dataset, 571,544,386 reviews and 275 GB on the HuggingFace CDN, and ranked every single review on four simple signals: how many strong-profanity word hits it has, how much of it is in ALL CAPS, the longest single run of consecutive exclamation marks, and how long it is. The question I started with was "how do people actually behave in Amazon reviews, and does the category they're reviewing change that?"
Live site, per-category breakdown, and the Wall of the loudest reviews: https://burla-cloud.github.io/amazon-review-distiller/
What surfaced:
- Video Games is the rowdiest category by a huge margin. 6.54% of video game reviews hit the strong-profanity list. Compare that to Gift Cards at 1.19% and Handmade at 1.08%. Movies & TV, CDs & Vinyl, Subscription Boxes, and Kindle Store fill out the top five. Cultural products attract feelings, consumer goods attract utility.
- Subscription Boxes is the angriest category. 15.89% of subscription box reviews are one-star. Almost 1 in 6. Charging people monthly for a curated surprise generates a lot of regret.
- The longest exclamation-mark run is 10,594 in a row. The review itself is two words ("love these") on a baby product. One person held one key down for a long time.
- The longest all-caps review is 1,169 words. Posted on a Mozart CD by a self-described disabled Vietnam veteran and Mozart scholar. He opens by apologizing for the caps (macular degeneration) and then keeps going for 1,169 more words.
- Forty reviewers gave a product five stars and wrote zero or one word. One five-star review of a cherry cough drop was just "Taste." That's the whole text.
- Books, music, and games write essays. Gift card buyers write nothing. Average review length: CDs & Vinyl 428 chars, Books 423, Kindle Store 367, Digital Music 340, Video Games 308. Gift Cards is at the bottom by a wide margin. Culture gets words, utility gets silence.
Methodology, plain version:
- The dataset is 34 separate
.jsonl.gzfiles on HuggingFace, one per Amazon category, totaling 275 GB. The usual workflow is to download all 275 GB to a laptop, then iterate. I didn't want to do that. - The HuggingFace CDN supports HTTP Range requests. A worker can ask for "give me bytes 1,000,000,000 to 1,500,000,000 of this file" and get just that slice without downloading the whole file. I split the 34 files into 545 chunks of about 500 MB each, on byte-range boundaries.
- Each chunk runs on its own worker. The worker streams its byte range row by row, scores every review on the four signals, and writes the top scoring reviews to a shared folder.
- A separate reducer container merges the per-chunk top-K shards into the final ranked lists per finding.
Map step: 3.21 minutes. Reduce step: 9.2 seconds. End to end under four minutes for 571 million reviews.
The pipeline runs on Burla using remote_parallel_map(worker, jobs, func_cpu=1, func_ram=4, max_parallelism=1000, grow=True). In English: "ask for up to 1000 parallel workers, each with 1 CPU and 4 GB of RAM, and let the cluster grow to meet that demand." In practice the cluster peaked around 500 concurrent workers and held there for the run. Workers run on a stock python:3.12 Docker image, and Burla auto-installs my local Python packages onto each one. The shared output folder is a Google Cloud Storage path that every worker writes to like a network drive.
(Disclosure: I work on Burla. The script and the live site are open source on GitHub. The dataset is the McAuley Lab's 2023 corpus on HuggingFace.)
Caveats worth being upfront about:
- Scoring is rule-based, not model-based. Word lists for strong, medium, and mild profanity, plus caps ratio, plus longest exclamation run. No sentiment model. That's deliberate: every score is reproducible and you can see exactly why a review got it.
- English-only. Reviews not in English get scored only by length, caps, and punctuation, because the word list is English. A multilingual sentiment model would do better here.
- Quoted titles leak in. A review of "Dick Tracy" can match the strong word list. There's a rescorer that penalizes capitalized-noun matches but it's imperfect.
- 2023 snapshot. The dataset is the McAuley Lab 2023 release, so it doesn't include reviews posted after mid-2023.
Repo with the full pipeline: https://github.com/Burla-Cloud/amazon-review-distiller
If anyone has a cleaner pattern for streaming huge HuggingFace datasets without materializing them locally, I'd love to hear it. I went with requests.get(..., stream=True) plus manual line splitting to keep the worker dependency surface tiny, but the datasets library probably has a cleaner Range-based path.
[link] [comments]
Read on the original site
Open the publisher's page for the full experience