Skip to content
thoufeelxPublic

About

PULSE — Self-healing product monitoring

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

PULSE — cinematic landing page

PULSE

The web changes. Scrapers break. PULSE heals — and proves the recovery.

A self-healing web-data system, demonstrated as a live price monitor.
When a site changes and extraction breaks, PULSE detects the failure, triggers Bright Data's AI repair on the scraper, then refuses to trust the fix until the recovered data re-passes the same validation gate. Every break, repair, and verdict is on permanent record.

🌐 Live app · 🎬 Demo video · 🧪 Resilience Lab · 📄 Structured output


😤 The Problem

Traditional scrapers assume the website stays the same. It doesn't.

A selector changes. A field disappears. A layout ships at 2 AM.

The scraper keeps running. It doesn't crash. It returns something.

The dangerous failure is not the crash — it's the silent wrong answer. A price feed that quietly reports null, a stale value, or a mis-parsed number poisons everything downstream while looking perfectly healthy in your dashboard.

Existing tooling stops at "did the HTTP request succeed?" Nothing asks the follow-up question: can this data be trusted — and if not, who fixes the scraper?

That question is PULSE's entire reason to exist.


⚙️ How PULSE Works

        Public web page
              ↓
   Bright Data Scraper Studio      ← custom collectors built for this project
              ↓
   Production collector output
              ↓
   Normalization (per-source adapters)
              ↓
   ┌─────────────────────────┐
   │  PULSE VALIDATION GATE  │   Zod contract + anomaly rules
   └─────────────────────────┘
        ↓               ↓
     VERIFIED         FAILED / UNVERIFIED
        ↓               ↓
   Trusted data     Healing engine → Bright Data AI repair
     continues           ↓
                    Re-run collector against the live page
                         ↓
                    SAME validation gate again
                         ↓
              VERIFIED → source RECOVERED, monitoring resumes
             

Every arrow is real code, real HTTPS calls to Bright Data, and real database rows — no simulated steps anywhere in the pipeline.

The loop, stage by stage

# Stage What actually happens Where
01 COLLECT The monitor's custom Scraper Studio collector runs against the live URL (realtime run, with automatic trigger/poll fallback under load) src/lib/brightdata/http-client.ts
02 NORMALIZE Raw collector output passes through a per-source adapter that handles envelope shapes and price-field aliases (price.value, current_price, offer_price…) src/lib/adapters/sources.ts
03 VALIDATE A strict Zod contract plus anomaly rules score every observation: VERIFIED, UNVERIFIED, or FAILED. Baseline = last verified observation src/lib/validation/
04 DETECT On failure, a healing_events row is written with the exact validation errors, and the source flips to HEALING src/lib/healing/engine.ts
05 HEAL Bright Data's AI repair is triggered with the failure reason as prompt, constrained to restore selectors without changing the output schema; the regenerated code is accepted via Bright Data's approval API healCollectorFull()
06 RE-RUN The repaired collector immediately re-runs against the same live page runScraper()
07 VERIFY Healed output goes through the identical validation gate. Only VERIFIED counts validation/engine.ts
08 TRUST RECOVERED flips the source healthy, the observation is stored, monitoring continues. Interrupted heals are reconciled as interrupted — never silently rewritten as success monitoring/engine.ts

🔧 Bright Data Scraper Studio

This project uses custom Scraper Studio scrapers created specifically for PULSE — not prebuilt library scrapers.

Division of responsibility — the core of this build:

Bright Data owns web extraction and scraper recovery: the collectors fetch and parse live pages, and Scraper Studio's AI-healing APIs regenerate broken selector code when a site changes.

PULSE owns everything around that: when to run, whether the output can be trusted, when healing is warranted, whether a healed scraper has actually recovered, and the entire product experience (monitors, timelines, lab, observatory).

All Bright Data interaction is direct HTTPS (src/lib/brightdata/http-client.ts) — no CLI dependency — so identical behavior locally and on Vercel serverless:

  • createCollectorAi / startCollectorBuild / getTemplateBuildProgress — provision new custom collectors (used to spawn isolated Lab clones on demand)
  • runCollector — realtime collection with trigger/poll fallback
  • healCollectorFull — full repair loop: trigger → poll → auto-approve at the gate → done (15-minute budget, 409-retry when a previous repair job is still draining)
  • startHeal / getHealProgress / approveHeal — stepped primitives powering the Resilience Lab's state machine
  • Web Unlocker zone — live product search (Amazon/Flipkart HTML parsers), JSON-LD detail extraction for open-web sources, and the Google-offers price fallback

Bright Data documentation: https://docs.brightdata.com


🧪 Resilience Lab

The hardest part of judging a self-healing system is that real sites change on their own schedule, not yours. The Resilience Lab solves that: controlled, reproducible failure on demand, end-to-end through real infrastructure.

🛒 Pulse demo Store inside a reilience Lab (/lab/store) — break our own storefront

PULSE ships a real storefront (/pulse_store) so self-healing can be demonstrated on HTML we control:

  • Card-vs-card diffing — original card beside the currently-served card; changes render as old → new, removals say so explicitly. What you see diffed is what Bright Data scrapes.
  • Edit advisor — armed edits are checked against src/lib/validation/rules.ts, the same thresholds the gate enforces. If your injected price would be held UNVERIFIED by the anomaly band, the UI tells you before you inject.
  • Live-proof window — SAVE & INJECT really publishes the edit to the storefront for 90 seconds (RESTORE_DWELL_MS) with a direct link to the affected product page, then restores the pristine template automatically (with sweep recovery if a session dies mid-window).
  • One-click demo — captures baseline, deletes the price node (guaranteed extraction kill), rides the genuine Bright Data repair to a verdict.

DEMO: gallery/lab.mp4


🛡 Monitoring & Validation — the trust layer

Scrape succeeded ≠ data is trustworthy. PULSE enforces the difference mechanically.

Every observation must clear:

Check Rule Outcome on breach
Schema contract Zod: typed fields — name, brand/model/variant/color, positive price, ISO-8601 scrapedAt, valid URLs, enum availability, 3-letter currency… FAILED → healing
Price rise band > +200% vs last verified price UNVERIFIED → source DEGRADED, held for review
Price crash band ≥ −60% vs last verified price (drops are bounded by zero — without an explicit floor, a bogus "90% off" would sail through a symmetric rule) UNVERIFIED → held for review
Currency drift Changed vs last verified warning

The comparison baseline is always the last VERIFIED observation — untrusted data never becomes the reference point.

Thresholds live in one file (src/lib/validation/rules.ts) shared by the server-side gate and the Lab's edit advisor, so what the UI promises is exactly what the engine enforces.

Source health is a first-class state machine: HEALTHY · DEGRADED · HEALING · VERIFYING · RECOVERED · FAILED · FAILED_PERMANENTLY, each mapped to a health score (100 → 0).

Monitor detail — trust view of a tracked source

Monitors dashboard — health status and healing timeline

Monitors dashboard — health state and event timeline per tracked source.

The Healing Observatory (/observatory) renders the immutable audit trail: trigger reason, the actual prompt sent to Bright Data, the repair response, approval status, verification result, and final verdict for every healing event.

Healing Observatory — audit trail of healing events

Healing Observatory — every break and recovery, permanently on record.


🔍 Product surfaces

Live search & monitoring — search products via Web Unlocker (Amazon/Flipkart parsers + Google offers grid), create monitors with per-source tracking, configure threshold alerts delivered by email (Resend) or webhook.

Live product search

product detail — trust view of a tracked source

Monitor detail — price history, source health badges, healing timeline, alert config.

How It Works (/how-it-works) — seven scroll-paced chapters (COLLECT → VALIDATE → DETECT → HEAL → RE-RUN → VERIFY → TRUST) that mirror the pipeline table above, so the explainer and the code can't drift apart.

How It Works chapter view


Prefer hands-on? The fastest 3-minute tour on the live site:

Landing (scroll the film) → Monitors → open a monitor → Resilience Lab
→ inject a break → watch validation fail → watch Bright Data heal it
→ read the verdict in the Observatory

Note: healing calls Bright Data's real AI repair and typically completes in minutes — the demo video compresses the wait.


📄 Example Structured Output

Real row from PULSE's price_observations table — the normalized, validated result of a collector run against the Pulse Mart storefront (recovered during a healing cycle):

{
  "monitor_id": "c9e2461e-852a-4ea7-acf4-997e34bdcd06",
  "monitor_source_id": "04d3ebb8-ade6-4cf7-aaa8-995d8887112e",
  "product_id": "558de941-1942-4928-b03a-0833cf47b75a",
  "price": 154999,
  "currency": "INR",
  "availability": "in_stock",
  "image_url": null,
  "product_url": "https://pulse-v1-beige.vercel.app/pulse_store/p/samsung-galaxy-s26-ultra-5g-cobalt-violet-12gb-ram-256gb-sto",
  "source": "RESILIENCE-LAB · samsung-galaxy-s26-ultra-5g-cobalt-violet-12gb-ram-256gb-sto",
  "observed_at": "2026-08-23T14:22:33.826+00:00",
  "validation_status": "VERIFIED",
  "raw_data": {
    "input": { "url": "https://pulse-v1-beige.vercel.app/pulse_store" },
    "product_url": "https://pulse-v1-beige.vercel.app/pulse_store/p/samsung-galaxy-s26-ultra-5g-cobalt-violet-12gb-ram-256gb-sto",
    "availability": "In Stock · ships in 24h",
    "product_name": "samsung galaxy s26",
    "current_price": 154999,
    "product_page_url": "https://pulse-v1-beige.vercel.app/pulse_store"
  }
}

What matters here:

  • validation_status: "VERIFIED" — this row cleared the Zod contract and sat inside both anomaly bands relative to the previous verified observation. Only rows with this badge drive charts, alerts, and baselines.
  • raw_data — the untouched collector output preserved verbatim for audit; normalization never destroys the original.
  • observed_at + source — every datapoint is traceable to a specific collector run on a specific source at a specific time.

The companion audit record (healing_events) stores the trigger reason, the exact prompt sent to Bright Data's repair API, the response, approval status, and the post-heal verification result.


🏗 Architecture

flowchart TB
    subgraph UI["PULSE UI — Next.js App Router"]
        L["Cinematic landing"]
        S["Live search"]
        M["Monitors + monitor detail"]
        LAB["Resilience Lab<br/>(retailer + store modes)"]
        PS["Pulse Mart storefront"]
        OB["Healing Observatory"]
    end

    subgraph CORE["PULSE Core — API routes + engines"]
        API["/api/monitors · /api/lab/*<br/>/api/healing · /api/products"]
        ENG["Monitoring engine"]
        VAL["Validation gate<br/>Zod + anomaly rules"]
        HEAL["Healing engine"]
        ALERT["Alerts — Resend + webhooks"]
    end

    subgraph BD["Bright Data"]
        SC["Scraper Studio custom collectors<br/>AMAZON · FLIPKART · WEB<br/>+ isolated Lab clones"]
        WU["Web Unlocker zone"]
        AIR["AI scraper healing API"]
    end

    WEB["Public web pages<br/>(retailers — and Pulse Mart itself)"]
    DB[("Supabase PostgreSQL<br/>13 tables, RLS on")]
    CRON["node-cron worker<br/>(every minute)"]

    UI --> API
    CRON -->|"POST /api/internal/run-source"| API
    API --> ENG
    ENG --> VAL
    VAL -->|"FAILED"| HEAL
    ENG --> ALERT
    ENG --> SC
    HEAL --> AIR
    HEAL --> SC
    S --> WU
    SC --> WEB
    WU --> WEB
    ENG --> DB
    HEAL --> DB
    VAL --> DB
Loading

Runtime split: the deployed Vercel app serves the UI, search, dashboards, and observatory; the cron worker (npm run monitor:worker) drives scheduled runs through /api/internal/run-source; the full long-running healing demo runs locally where function duration is unbounded.


🧰 Tech Stack

Layer Technology
Frontend Next.js 14 (App Router), React 18, TypeScript 5, Tailwind CSS, custom canvas scroll-cinema (258 per-frame WebPs)
Backend Next.js API routes, tsx scripts, node-cron worker
Database Supabase PostgreSQL (RLS enabled on all tables)
Extraction & healing Bright Data — custom Scraper Studio collectors + Web Unlocker + AI healing APIs (direct HTTPS)
Validation Zod schemas + shared plausibility rules
Notifications Resend (email) + generic webhooks
Testing Playwright — 21 E2E specs across desktop Chrome and mobile Chrome (landscape + portrait)
Assets tooling sharp (sprite-sheet splitting script)

🗃 Data Model

Thirteen Supabase tables, RLS on everywhere:

products ──< product_variants
    │
    └──< monitors ──< monitor_sources ──< scraper_collectors
                            │                   │
                            │             collector_id →
                            │             Bright Data Scraper Studio
                            ├──< price_observations   (validated datapoints)
                            └──< scrape_runs          (every attempt, pass or fail)

healing_events        (break → repair → verify audit trail)
alerts                (threshold + channel) ──< notification_events
lab_sessions ──< lab_events            (Resilience Lab recordings)
store_pages           (Pulse Mart catalog + served edits)

Flow of one check: scrape_runs records the attempt → adapter normalizes → gate scores it → price_observations stores only schema-valid rows tagged VERIFIED / UNVERIFIED → failures branch into healing_events → sources carry status + health score forward.


✅how to test this urself

  1. Open https://pulse-v1-beige.vercel.app/ — scroll the landing film; it ends by dropping you into live search.
  2. Search a product (e.g. "iPhone 15") — results are fetched live through Bright Data.
  3. Create a monitor from any result; open its detail page — price history, source health, alert setup.
  4. Visit the Resilience Lab — pick a mode, capture the baseline, inject a break (REMOVE_FIELD on price is the cleanest kill).
  5. Watch validation fail, then Bright Data's repair run. This is real AI healing — give it a few minutes.
  6. Read the verdict: recovered data re-passed the same gate; source flipped to RECOVERED.
  7. Close the loop in the Observatory — the entire incident, permanently recorded.

Short on time? gallery/lab.mp4 shows exactly this.


📁 Project Structure

src/
├── app/
│   ├── page.tsx                  # 258-frame cinematic landing
│   ├── how-it-works/             # 7-chapter pipeline explainer
│   ├── monitor/[id]/             # Monitor detail — trust view
│   ├── monitors/                 # Dashboard
│   ├── lab/                      # Resilience Lab (retailer + store modes)
│   ├── observatory/              # Healing audit trail
│   ├── pulse_store/              # Live storefront (+ product pages)
│   └── api/
│       ├── monitors/[id]/run     # Trigger a real check
│       ├── internal/run-source   # Cron entrypoint
│       ├── lab/{session,run,history}
│       ├── lab/store/{session,inject,heal-step,reset,products}
│       ├── healing · products · alerts · frame[s]/[n]
├── components/
│   ├── pulse/                    # Landing film, price chart, healing timeline, health states
│   ├── lab/                      # RetailerLab · StoreLab panels
│   └── ui/                       # Primitives
└── lib/
    ├── brightdata/               # HTTP client · collectors · unlocker · details · google-search
    ├── adapters/                 # Per-source normalization
    ├── validation/               # schemas.ts · engine.ts · rules.ts · edit-advisor.ts
    ├── healing/engine.ts         # Heal state machine + stale-heal reconciliation
    ├── monitoring/engine.ts      # scrape → normalize → validate → store → alert
    ├── lab/                      # pipelines · mutations · session store
    ├── notifications/alerts.ts   # Resend + webhook delivery
    ├── store/                    # Pulse Mart render/service/origin
    └── db/supabase.ts            # Cache-bypassing Supabase client
scripts/
├── scheduler.ts                  # node-cron worker (every minute)
├── run-monitor.ts                # One-shot monitor sweep
├── production-gate.ts            # Env + DB + typecheck + build + secret scan
└── split-sprite-sheets.mjs       # sharp asset pipeline for the landing film
tests/                            # Playwright E2E suites (21 specs × 3 browser projects)
gallery/                          # Screenshots + demo video (lab.mp4)

🤖 AI Usage

— AI coding assistants were used in building PULSE:

  • Tools: (agentic coding CLIs).
  • What they did: scaffolding files and boilerplate, implementing components and modules under direction, debugging assistance, test drafting, refactoring help, and documentation drafts.
  • What the human did: product and architecture decisions, the honesty/trust design of the validation gate, integration of the Bright Data APIs, running every real end-to-end verification against live services (including actual healing cycles), reviewing and correcting AI-generated code, and final understanding of the whole system.
  • AI was an accelerator, not an autopilot: nothing shipped unreviewed, and every claim in this README was verified against the running application.

📝 License

No license file has been added yet. All rights reserved by the author until one is chosen.


PULSE doesn't just heal scrapers. It proves they healed.

About

PULSE — Self-healing product monitoring

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages