Data Extraction

How to Scrape Amazon Reviews in 2026: Python Script, NLP & Fake Review Detection

A comprehensive 2,200+ word engineering guide on extracting, filtering, and analyzing Amazon reviews. Learn how to detect fake reviews, bypass pagination limits, and visualize NLP data.

Alex Chen, Lead Data Engineer9 min read

TL;DR (Bottom Line Up Front): To scrape Amazon product reviews without triggering CAPTCHAs, you must automate pagination handling (looping across the /product-reviews/ASIN sub-pages), rotate residential proxies to prevent rate-limit blocks, filter out incentivized/unverified reviews, and extract structured review metadata (Rating, Review ID, Helpful Votes, Date, Body text) for NLP sentiment analysis.

Amazon product reviews are arguably the most valuable source of consumer sentiment data on the internet. Every day, millions of customers leave highly detailed feedback on everything from the durability of a $10 kitchen gadget to the battery performance of a $2,000 laptop.

For brand managers, product developers, hedge funds, and market research firms, this unstructured text corpus is an absolute goldmine. Scraping and analyzing these reviews allows you to identify your competitors' fatal flaws, discover emerging market trends, engineer superior products, and monitor brand reputation in real time.

However, Amazon fiercely protects its customer review data. Extracting 10,000 reviews is not as simple as running a basic Python loop. In this definitive 2,200+ word technical guide, we will explore the engineering challenges of review scraping, provide complete Python extraction scripts, examine how to detect and filter out incentivized "fake" reviews, and build an end-to-end NLP (Natural Language Processing) sentiment analysis pipeline.


1. The Technical Architecture of Amazon Reviews

To extract reviews at volume, you must understand where Amazon stores review data and how the pagination pipeline functions.

Unlike product detail pages (/dp/ASIN), which only display the top 8 to 10 "helpful" customer reviews, Amazon maintains a dedicated customer review hub at the following URL path:

https://www.amazon.com/product-reviews/{ASIN}/?pageNumber={PAGE}&sortBy=recent&filterByStar={STAR_RATING}
[Target ASIN]
      │
      ▼
[/product-reviews/ASIN/] ──(Page 1: 10 Reviews) ──> [Extract Review IDs & Text]
      │
      ▼ (Asynchronous Concurrent Pagination)
[/product-reviews/ASIN/?pageNumber=2] ────────────> [Extract Review IDs & Text]
      │
      ▼
[/product-reviews/ASIN/?pageNumber=N] ────────────> [Extract Review IDs & Text]

Key Pagination Challenges

  1. The 500-Page Ceiling: Amazon displays 10 reviews per page, up to a maximum of 500 pages (5,000 public reviews). Even if a product has 80,000 total ratings, Amazon only makes the most recent 5,000 reviews publicly accessible via pagination.
  2. Sequential Rate-Limiting: If a single script requests page 1, page 2, page 3 sequentially from the same IP, Amazon's anti-bot system triggers a CAPTCHA challenge before page 15 is reached.
  3. Dynamic Review Filters: You can filter review pages by star rating (filterByStar=one_star, filterByStar=five_star), reviewer verification (filterByStar=avp_only_reviews), and format (formatType=current_format). Combining these filters allows you to extract distinct subsets of reviews that exceed standard pagination limits.

2. Essential Data Fields to Extract

When performing review sentiment mining, you must extract complete structured metadata alongside the raw review body:

Field NameData TypeCSS / DOM PathAnalytical Purpose
review_idString (14)div[data-hook='review'] @idPrimary key for deduplication across daily crawls.
asinString (10)URL / Canonical TagLinks review to specific catalog ASIN.
star_ratingInteger (1-5)i[data-hook='review-star-rating'] spanQuantitative sentiment polarity score.
review_titleStringa[data-hook='review-title'] spanConcise summary of core customer complaint or praise.
review_bodyStringspan[data-hook='review-body'] spanFull unstructured text for NLP tokenization and clustering.
review_dateDate (ISO)span[data-hook='review-date']Time-series analysis and defect tracking over time.
is_verifiedBooleanspan[data-hook='avp-badge']Filters out unverified or incentivized bot reviews.
helpful_votesIntegerspan[data-hook='helpful-vote-statement']Weighting multiplier for community impact.
variationStringa[data-hook='format-strip']Pinpoints complaints to specific product sizes or colors.

3. Python Review Scraper (Full Implementation)

Here is a complete, production-ready Python script utilizing requests, BeautifulSoup, and randomized residential proxy headers to scrape paginated Amazon reviews:

import requests
from bs4 import BeautifulSoup
import re
import time
import json

def parse_amazon_review_page(html_content, asin):
    soup = BeautifulSoup(html_content, "html.parser")
    review_cards = soup.find_all("div", {"data-hook": "review"})
    reviews = []

    for card in review_cards:
        # 1. Review ID
        review_id = card.get("id")
        
        # 2. Star Rating
        star_el = card.find("i", {"data-hook": "review-star-rating"})
        if not star_el:
            star_el = card.find("i", {"data-hook": "cmps-review-star-rating"})
        star_rating = None
        if star_el:
            match = re.search(r"(\d+(\.\d+)?)", star_el.get_text())
            star_rating = float(match.group(1)) if match else None

        # 3. Review Title
        title_el = card.find("a", {"data-hook": "review-title"})
        if not title_el:
            title_el = card.find("span", {"data-hook": "review-title"})
        review_title = title_el.get_text(strip=True) if title_el else ""
        # Clean leading star rating text if included in title span
        review_title = re.sub(r"^\d\.\d out of 5 stars\s*", "", review_title)

        # 4. Review Body
        body_el = card.find("span", {"data-hook": "review-body"})
        review_body = body_el.get_text(strip=True) if body_el else ""

        # 5. Review Date & Country
        date_el = card.find("span", {"data-hook": "review-date"})
        review_date = None
        country = None
        if date_el:
            date_text = date_el.get_text(strip=True)
            # Example format: "Reviewed in the United States on May 14, 2026"
            date_match = re.search(r"Reviewed in (.*?) on (.*)", date_text)
            if date_match:
                country = date_match.group(1).replace("the ", "")
                review_date = date_match.group(2)

        # 6. Verified Purchase Badge
        verified_el = card.find("span", {"data-hook": "avp-badge"})
        is_verified = bool(verified_el)

        # 7. Helpful Votes Count
        helpful_el = card.find("span", {"data-hook": "helpful-vote-statement"})
        helpful_votes = 0
        if helpful_el:
            helpful_text = helpful_el.get_text(strip=True)
            h_match = re.search(r"([\d,]+)", helpful_text)
            if h_match:
                helpful_votes = int(h_match.group(1).replace(",", ""))
            elif "One person" in helpful_text:
                helpful_votes = 1

        # 8. Variation Details
        format_el = card.find("a", {"data-hook": "format-strip"})
        variation = format_el.get_text(strip=True) if format_el else None

        if review_id and review_body:
            reviews.append({
                "review_id": review_id,
                "asin": asin,
                "star_rating": star_rating,
                "review_title": review_title,
                "review_body": review_body,
                "review_date": review_date,
                "country": country,
                "is_verified": is_verified,
                "helpful_votes": helpful_votes,
                "variation": variation
            })

    return reviews

def scrape_amazon_reviews(asin, max_pages=3):
    all_reviews = []
    headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/123.0.0.0 Safari/537.36",
        "Accept-Language": "en-US,en;q=0.9",
        "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8"
    }

    for page in range(1, max_pages + 1):
        url = f"https://www.amazon.com/product-reviews/{asin}/?pageNumber={page}&sortBy=recent"
        print(f"Fetching reviews for {asin} - Page {page}...")
        
        response = requests.get(url, headers=headers)
        if response.status_code != 200:
            print(f"Failed to fetch page {page}. Status: {response.status_code}")
            break
            
        page_reviews = parse_amazon_review_page(response.text, asin)
        if not page_reviews:
            print("No more reviews found or CAPTCHA triggered.")
            break
            
        all_reviews.extend(page_reviews)
        time.sleep(2) # Throttle delay

    return all_reviews

if __name__ == "__main__":
    results = scrape_amazon_reviews("B08N5LNQCX", max_pages=2)
    print(f"Total reviews extracted: {len(results)}")
    print(json.dumps(results[:2], indent=2))

4. Detecting & Filtering "Fake" or Incentivized Reviews

Not all Amazon reviews represent legitimate customer feedback. Competitor black-hat launches, review farms, and incentivized rebate groups can flood listings with misleading 5-star ratings.

Before running NLP sentiment analysis, apply the following 4 algorithmic data hygiene filters:

[Raw Scraped Review Stream]
             │
             ▼
[Filter 1: Verified Purchase Filter] ──(Drop Unverified Reviews)
             │
             ▼
[Filter 2: Linguistic Repetition Check] ──(Flag Duplicate Copypasta)
             │
             ▼
[Filter 3: Review Velocity Anomaly] ──(Detect Unnatural 24h Spikes)
             │
             ▼
[Clean Verified Sentiment Corpus]

1. The Verified Purchase Requirement

Always filter for is_verified == True. Unverified reviews can be submitted by anyone without buying the product, making them the primary vehicle for fake ratings.

2. Linguistic Levenshtein Distance & Token Jaccard Similarity

Review farms frequently reuse the same 5 templates across hundreds of accounts (e.g. "Great product, works as advertised, fast shipping!"). Calculate pairwise Jaccard similarity across review texts:

def jaccard_similarity(text1, text2):
    set1 = set(text1.lower().split())
    set2 = set(text2.lower().split())
    intersection = len(set1.intersection(set2))
    union = len(set1.union(set2))
    return intersection / union if union > 0 else 0

If multiple reviews from different accounts share a similarity score > 0.85, flag them as coordinated bot reviews.


5. NLP Sentiment Analysis Pipeline with Python (VADER & RoBERTa)

Once you have extracted a clean corpus of 1,000+ reviews, you can process the unstructured text using Natural Language Processing.

Lexicon-Based Sentiment Scoring with VADER

from nltk.sentiment.vader import SentimentIntensityAnalyzer
import nltk

nltk.download('vader_lexicon', quiet=True)
sia = SentimentIntensityAnalyzer()

def analyze_review_sentiment(review_text):
    scores = sia.polarity_scores(review_text)
    # Compound score ranges from -1.0 (Extremely Negative) to +1.0 (Extremely Positive)
    compound = scores['compound']
    
    if compound >= 0.05:
        sentiment = "Positive"
    elif compound <= -0.05:
        sentiment = "Negative"
    else:
        sentiment = "Neutral"
        
    return {"sentiment": sentiment, "score": compound, "details": scores}

# Example analysis
sample_review = "The headphones sound great but the plastic headband snapped after just 2 weeks of normal use."
print(analyze_review_sentiment(sample_review))

Uncovering Product Defect Clusters with LLM Prompts

Piping 1-star review text into an LLM (such as OpenAI GPT-4 or Google Gemini) allows you to extract categorized engineering defect lists:

PROMPT TEMPLATE:
"Here is a CSV of 200 one-star reviews for competitor ASIN {ASIN}. 
Categorize all complaints into the top 5 engineering failure categories. 
For each category, provide:
1. Category Name (e.g., Battery Life, Sizing, Durability)
2. Percentage of 1-star reviews mentioning this issue
3. Three direct customer quote excerpts"

6. DIY Review Scraper vs. Managed Review Extraction API

Paginating through thousands of reviews across hundreds of ASINs requires immense proxy bandwidth and continuous maintenance.

+------------------------------------+------------------------------------+
| DIY In-House Review Scraper        | Managed API (AmazonScraping.com)   |
+------------------------------------+------------------------------------+
| 🔴 Paging limits at page 10-15     | 🟢 Full 500-page deep extraction   |
| 🔴 High proxy bandwidth consumption| 🟢 Zero proxy fees or server bills |
| 🔴 Manual CAPTCHA breaks           | 🟢 Automated 100% bypass           |
| 🔴 Raw unstructured HTML           | 🟢 Clean JSON/CSV with NLP tags    |
+------------------------------------+------------------------------------+

Explore our dedicated Amazon Review Scraper Service to extract thousands of historical reviews with zero code.


Amazon Scraping TeamData Extraction Specialists · 10+ Years Experience

Our team of senior data engineers and web scraping specialists has delivered over 500 million records across 12+ Amazon marketplaces. We write about scraping techniques, eCommerce data strategy, and Amazon market intelligence based on real-world project experience.