DEV Community

Fragrance notes, accords and ratings as JSON - without fighting Fragrantica's Cloudflare

Why the obvious source (Fragrantica) is the wrong first pick

Fragrantica is the biggest fragrance database, so it’s the first thing everyone scrapes. But it now sits behind Cloudflare, and most of the scrapers built against it are fighting that wall - you’ll find a handful on Apify, several with visibly stale runs or low success rates, because the site actively blocks plain requests. You end up paying in proxy cost, retries, and broken runs for data that’s conceptually simple: notes, accords, a rating.

Parfumo is the other large, community-driven fragrance encyclopedia - comparable scale and depth of data (notes, accords, per-dimension community ratings, perfumer credits) - and it still serves plain, uncached HTML with no Cloudflare challenge. It had zero dedicated scrapers on Apify as of this writing, which is the interesting part: the demand signal for this class of data is proven (1,700+ runs on the Fragrantica side alone), but almost all of that tooling is pointed at the one source that’s actively fighting back.

What the source data looks like, cleaned up

One record per fragrance, real output for Dior Sauvage EDT:

{
  "name": "Sauvage",
  "brand": "Dior",
  "releaseYear": 2015,
  "concentration": "Eau de Toilette",
  "gender": "men",
  "rating": {
    "value": 7.5,
    "best": 10,
    "ratingCount": 6514,
    "reviewCount": 365
  },
  "perfumers": ["FranΓ§ois Demachy"],
  "accords": ["Fresh", "Spicy", "Synthetic", "Citrus", "Woody"],
  "notes": ["Ambrox", "Sichuan pepper", "Calabrian bergamot", "Pink pepper", "..."],
  "communityRatings": {
    "scent": { "average": 7.30, "votes": 6514 },
    "longevity": { "average": 7.79, "votes": 6129 },
    "sillage": { "average": 7.73, "votes": 6106 },
    "bottle": { "average": 7.72, "votes": 6164 },
    "value": { "average": 7.40, "votes": 4561 }
  }
}

When a page groups notes into a pyramid (most designer releases do), you also get:

"notesPyramid": {
  "top": ["Bergamot", "Apple", "Blackcurrant", "Lemon", "Pink pepper"],
  "heart": ["Pineapple", "Indonesian patchouli", "Jasmine"],
  "base": ["Birch", "Cedarwood", "Musk", "Oakmoss", "Ambergris"]
}

That pyramid - not just a flat note list - is what a dupe/clone finder actually needs: two fragrances can share a base note and read completely differently because their top notes diverge, so comparing pyramids beats comparing flat sets.

Getting it without hand-copying spec pages

Parfumo Fragrance Scraper on Apify takes a list of perfume URLs (or Brand/slug pairs) and returns one clean record per fragrance:

curl -s -X POST "https://api.apify.com/v2/acts/minty_modesty~parfumo-perfume-data-scraper/runs?token=$APIFY_TOKEN" \
  -H 'content-type: application/json' \
  -d '{"perfumeUrls":["Dior/sauvage","Creed/aventus","Chanel/No_5"],"includeCommunityRatings":true}'

It reads only public /Perfumes/ pages - allowed by Parfumo's robots.txt, no login, no CAPTCHA, no Cloudflare challenge to route around - so a run doesn't turn into a proxy bill. URLs that don't resolve come back as {"url": "...", "error": "not_found"} instead of silently dropping the line item, so a batch of a few hundred references doesn't need babysitting.

When this is (and isn't) the right layer

If you already have a working Fragrantica pipeline and it isn't breaking, there's no reason to switch sources mid-project. But if you're starting a new dupe-finder, enrichment job, or rating tracker from scratch, the 1,700+ runs on the Fragrantica side tell you the data is worth having - they don't tell you Fragrantica is the only place to get it. Parfumo is the same class of data from a source that isn't actively trying to block you from reading it.

Top comments (0)

Comments

No comments yet. Start the discussion.