Building an Autonomous Competitor Price Monitor with Python, Playwright, and Discord (Zero SaaS Fees)
DEV Community

Building an Autonomous Competitor Price Monitor with Python, Playwright, and Discord (Zero SaaS Fees)

Building an Autonomous Competitor Price Monitor with Python, Playwright, and Discord (Zero SaaS Fees) In dynamic e-commerce sectors, missing a competitor's flash sale or midnight re-pricing event by even three hours can wipe out daily margins. If your competitor cuts an anchor SKU by 15%, your conversion rate plummets before your morning standup. The typical enterprise response is subscribing to SaaS monitoring platforms like Prisync or Price2Spy. These platforms run between $150 and $800 per month, enforce arbitrary SKU tracking tiers, and frequently fail when target sites tweak dynamic DOM selectors or deploy basic JavaScript challenges. In this guide, we will build an autonomous, self-hosted pricing intelligence pipeline in Python. It leverages HTTPX for rapid REST/static scraping, Playwright with stealth parameters for client-side rendered stores, SQLite for time-series pricing and stock tracking, and Discord Webhooks for instant, color-coded visual alerts. 1. The Bottleneck: Why Commercial Scrapers Break Most commercial pricing platforms rely on static DOM querying or expensive headless cloud clusters with high overhead. They present three major engineering problems: - Fragile Selectors & Aggressive Rate Limits: Commercial SaaS platforms share IP subnets across thousands of clients. Cloudflare, PerimeterX, and Datadome flag and throttle these subnet ranges routinely. - Payload Inefficiency: Running heavy Chromium instances for every single page check wastes compute. A production monitor needs a tiered retrieval architecture: high-speed HTTP requests with header spoofing first, falling back to lightweight headless browser contexts only when JavaScript execution or CAPTCHA mitigation is mandatory. - Black-box Polling: You cannot adjust polling frequencies per SKU or inspect why a scrape failed without contacting support. A self-hosted script provides direct control over headers, proxy rotations, and SQLite delta calculations. 2. System Architecture The pipeline runs on a lightweight cron schedule (every 15-60 minutes) and follows a four-stage execution lifecycle: [ Cron Trigger / Orchestrator ] │ โ–ผ [ Tiered Scraper (HTTPX -> Playwright Fallback) ] │ ├── User-Agent / Proxy Rotation │ └── Minimal Resource Blocking (Skip CSS/Fonts/Media) โ–ผ [ SQLite Time-Series Engine ] │ ├── Compare Latest vs Previous Price Record │ └── Compute Delta & Availability State โ–ผ [ Alert Dispatcher ] └── Discord Webhook Embed (Red = Undercut, Green = Hike, Blue = Parity) 3. Implementation: Code & Core Logic Step 1: SQLite Storage & Delta Engine We need an ACID-compliant, zero-dependency storage system that tracks historic SKU price movements and detects state transitions (e.g., in_stock to out_of_stock ). import sqlite3 from datetime import datetime from typing import Optional, Tuple DB_PATH = "pricing_intel.db" def init_db(): with sqlite3.connect(DB_PATH) as conn: cursor = conn.cursor() cursor.execute(""" CREATE TABLE IF NOT EXISTS price_history ( id INTEGER PRIMARY KEY AUTOINCREMENT, sku TEXT NOT NULL, store_name TEXT NOT NULL, url TEXT NOT NULL, price REAL NOT NULL, currency TEXT DEFAULT 'USD', is_in_stock INTEGER NOT NULL, timestamp DATETIME DEFAULT CURRENT_TIMESTAMP ) """) cursor.execute(""" CREATE INDEX IF NOT EXISTS idx_sku_timestamp ON price_history(sku, timestamp DESC) """) conn.commit() def get_last_price(sku: str, store_name: str) -> Optional[Tuple[float, int]]: """Retrieves the most recent price and stock status for a given SKU.""" with sqlite3.connect(DB_PATH) as conn: cursor = conn.cursor() cursor.execute(""" SELECT price, is_in_stock FROM price_history WHERE sku = ? AND store_name = ? ORDER BY timestamp DESC LIMIT 1 """, (sku, store_name)) return cursor.fetchone() def record_price(sku: str, store_name: str, url: str, price: float, is_in_stock: bool, currency: str = 'USD'): with sqlite3.connect(DB_PATH) as conn: cursor = conn.cursor() cursor.execute(""" INSERT INTO price_history (sku, store_name, url, price, is_in_stock, currency, timestamp) VALUES (?, ?, ?, ?, ?, ?, ?) """, (sku, store_name, url, price, int(is_in_stock), currency, datetime.utcnow().isoformat())) conn.commit() Step 2: The Tiered Extraction Engine Instead of launching Chromium for every scrape, we attempt an async HTTP pull via httpx first. If the page is rendered client-side (SPA) or returns a 403 Forbidden , we fail over to headless playwright with asset-blocking to reduce memory usage by ~70%. import asyncio import httpx from playwright.async_api import async_playwright from selectolax.parser import HTMLParser DEFAULT_HEADERS = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/122.0.0.0 Safari/537.36", "Accept-Language": "en-US,en;q=0.9", "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,/;q=0.8", } async def fetch_html_httpx(url: str) -> Optional[str]: async with httpx.AsyncClient(headers=DEFAULT_HEADERS, follow_redirects=True, timeout=12.0) as client: try: response = await client.get(url) if response.status_code == 200: return response.text except httpx.HTTPError: return None return None async def fetch_html_playwright(url: str) -> str: async with async_playwright() as p: browser = await p.chromium.launch(headless=True) context = await browser.new_context( user_agent=DEFAULT_HEADERS["User-Agent"], viewport={"width": 1920, "height": 1080} ) page = await context.new_page() # Block resource types that do not contribute to pricing DOM await page.route("**/*", lambda route: ( route.abort() if route.request.resource_type in ["image", "media", "font", "stylesheet"] else route.continue_() )) await page.goto(url, wait_until="domcontentloaded", timeout=30000) content = await page.content() await browser.close() return content Step 3: Parsing & Normalization Using selectolax (which is up to 30x faster than BeautifulSoup4), we extract the price and availability tokens: import re def parse_product_data(html: str, price_selector: str, stock_selector: str) -> Tuple[float, bool]: tree = HTMLParser(html) # Extract Price price_node = tree.css_first(price_selector) if not price_node: raise ValueError(f"Selector '{price_selector}' failed to resolve.") raw_price = price_node.text(strip=True) cleaned_price = re.sub(r"[^0-9.]", "", raw_price) price = float(cleaned_price) # Extract Stock Status stock_node = tree.css_first(stock_selector) is_in_stock = True if stock_node: text = stock_node.text(strip=True).lower() if any(term in text for term in ["out of stock", "sold out", "unavailable"]): is_in_stock = False return price, is_in_stock Step 4: The Discord Webhook Alerting Dispatcher To make notifications actionable on mobile or desktop, we construct a Discord Rich Embed. We compute the delta between runs: - Undercut (Competitor lowered price): Red (0xE74C3C) - Immediate threat. - Hike (Competitor raised price): Green (0x2ECC71) - Margin opportunity. - Parity / First Run: Blue (0x3498DB) - Baseline sync. async def dispatch_discord_alert(webhook_url: str, sku: str, store: str, url: str, old_price: Optional[float], new_price: float, in_stock: bool): delta_text = "N/A (Initial Record)" embed_color = 0x3498DB # Blue if old_price is not None: diff = new_price - old_price pct_change = (diff / old_price) * 100 if diff 0: embed_color = 0x2ECC71 # Green: Price increased (Opportunity) delta_text = f"๐Ÿ“ˆ Increased by +${diff:.2f} (+{pct_change:.2f}%)" else: delta_text = "โš–๏ธ Unchanged" payload = { "username": "Price Monitor Daemon", "avatar_url": "https://i.imgur.com/8NydhPO.png", "embeds": [{ "title": f"Pricing Alert: {sku}", "url": url, "color": embed_color, "fields": [ {"name": "Competitor Store", "value": store, "inline": True}, {"name": "Current Price", "value": f"${new_price:.2f}", "inline": True}, {"name": "Previous Price", "value": f"${old_price:.2f}" if old_price else "N/A", "inline": True}, {"name": "Price Delta", "value": delta_text, "inline": False}, {"name": "Inventory Status", "value": "โœ… In Stock" if in_stock else "โŒ Out of Stock", "inline": True} ], "footer": {"text": "Autonomous Price Engine • Production Monitor"}, "timestamp": datetime.utcnow().isoformat() }] } async with httpx.AsyncClient() as client: await client.post(webhook_url, json=payload) 4. Production Deployment & Resilience When running this system on a remote VPS (e.g., $5/mo DigitalOcean or Hetzner droplet), consider these operational rules: - Rotating User-Agents and Proxies: Pass a proxy dictionary directly into httpx.AsyncClient(proxies="http://user:pass@gate.smartproxy.com:7000") and Playwright'sbrowser.new_context(proxy=...) to prevent IP ban escalation. - Memory Leaks in Headless Browsers: When running continuously, always close Playwright contexts or run each execution cycle inside a clean subprocess or stateless Docker container. - Execution via Crontab: Run the daemon via cron without running a bloated persistent loop: */30 * * * * cd /opt/price-monitor && /opt/price-monitor/venv/bin/python main.py >> /var/log/price_monitor.log 2>&1 5. Conclusion & Ready-to-Use Workflow You can manually copy the snippets above, map your specific target selectors, and wire up the database directly on your own machine. However, if you want a complete, turnkey solution that includes pre-built multi-store drivers (Amazon, Shopify, WooCommerce, Magento), automated anti-bot bypass routines, Docker Compose configurations, and multi-channel Discord alerts out of the box, you can grab the production package below: - Instant Access via Whop: E-Commerce Competitor Price Monitor & Alert Engine - Direct Download on Gumroad: Get the Source Code & Fixtures - use code EARLYBIRD for 20% off. Take control of your pricing telemetry, stop paying recurring SaaS tolls, and act on competitor margin shifts in real time. Top comments (0)

Read on DEV Community ↗ ← Back to News

Comments

No comments yet. Start the discussion.