The web changes. Scrapers break. PULSE heals — and proves the recovery.
A self-healing web-data system, demonstrated as a live price monitor.
When a site changes and extraction breaks, PULSE detects the failure, triggers Bright Data's
AI repair on the scraper, then refuses to trust the fix until the recovered data
re-passes the same validation gate. Every break, repair, and verdict is on permanent record.
🌐 Live app · 🎬 Demo video · 🧪 Resilience Lab · 📄 Structured output
Traditional scrapers assume the website stays the same. It doesn't.
A selector changes. A field disappears. A layout ships at 2 AM.
The scraper keeps running. It doesn't crash. It returns something.
The dangerous failure is not the crash — it's the silent wrong answer. A price feed that quietly reports null, a stale value, or a mis-parsed number poisons everything downstream while looking perfectly healthy in your dashboard.
Existing tooling stops at "did the HTTP request succeed?" Nothing asks the follow-up question: can this data be trusted — and if not, who fixes the scraper?
That question is PULSE's entire reason to exist.
Public web page
↓
Bright Data Scraper Studio ← custom collectors built for this project
↓
Production collector output
↓
Normalization (per-source adapters)
↓
┌─────────────────────────┐
│ PULSE VALIDATION GATE │ Zod contract + anomaly rules
└─────────────────────────┘
↓ ↓
VERIFIED FAILED / UNVERIFIED
↓ ↓
Trusted data Healing engine → Bright Data AI repair
continues ↓
Re-run collector against the live page
↓
SAME validation gate again
↓
VERIFIED → source RECOVERED, monitoring resumes
Every arrow is real code, real HTTPS calls to Bright Data, and real database rows — no simulated steps anywhere in the pipeline.
| # | Stage | What actually happens | Where |
|---|---|---|---|
| 01 | COLLECT | The monitor's custom Scraper Studio collector runs against the live URL (realtime run, with automatic trigger/poll fallback under load) | src/lib/brightdata/http-client.ts |
| 02 | NORMALIZE | Raw collector output passes through a per-source adapter that handles envelope shapes and price-field aliases (price.value, current_price, offer_price…) |
src/lib/adapters/sources.ts |
| 03 | VALIDATE | A strict Zod contract plus anomaly rules score every observation: VERIFIED, UNVERIFIED, or FAILED. Baseline = last verified observation |
src/lib/validation/ |
| 04 | DETECT | On failure, a healing_events row is written with the exact validation errors, and the source flips to HEALING |
src/lib/healing/engine.ts |
| 05 | HEAL | Bright Data's AI repair is triggered with the failure reason as prompt, constrained to restore selectors without changing the output schema; the regenerated code is accepted via Bright Data's approval API | healCollectorFull() |
| 06 | RE-RUN | The repaired collector immediately re-runs against the same live page | runScraper() |
| 07 | VERIFY | Healed output goes through the identical validation gate. Only VERIFIED counts |
validation/engine.ts |
| 08 | TRUST | RECOVERED flips the source healthy, the observation is stored, monitoring continues. Interrupted heals are reconciled as interrupted — never silently rewritten as success |
monitoring/engine.ts |
This project uses custom Scraper Studio scrapers created specifically for PULSE — not prebuilt library scrapers.
Division of responsibility — the core of this build:
Bright Data owns web extraction and scraper recovery: the collectors fetch and parse live pages, and Scraper Studio's AI-healing APIs regenerate broken selector code when a site changes.
PULSE owns everything around that: when to run, whether the output can be trusted, when healing is warranted, whether a healed scraper has actually recovered, and the entire product experience (monitors, timelines, lab, observatory).
All Bright Data interaction is direct HTTPS (src/lib/brightdata/http-client.ts) — no CLI dependency — so identical behavior locally and on Vercel serverless:
createCollectorAi/startCollectorBuild/getTemplateBuildProgress— provision new custom collectors (used to spawn isolated Lab clones on demand)runCollector— realtime collection with trigger/poll fallbackhealCollectorFull— full repair loop: trigger → poll → auto-approve at the gate → done (15-minute budget, 409-retry when a previous repair job is still draining)startHeal/getHealProgress/approveHeal— stepped primitives powering the Resilience Lab's state machine- Web Unlocker zone — live product search (Amazon/Flipkart HTML parsers), JSON-LD detail extraction for open-web sources, and the Google-offers price fallback
Bright Data documentation: https://docs.brightdata.com
The hardest part of judging a self-healing system is that real sites change on their own schedule, not yours. The Resilience Lab solves that: controlled, reproducible failure on demand, end-to-end through real infrastructure.
PULSE ships a real storefront (/pulse_store) so self-healing can be demonstrated on HTML we control:
- Card-vs-card diffing — original card beside the currently-served card; changes render as
old→ new, removals say so explicitly. What you see diffed is what Bright Data scrapes. - Edit advisor — armed edits are checked against
src/lib/validation/rules.ts, the same thresholds the gate enforces. If your injected price would be heldUNVERIFIEDby the anomaly band, the UI tells you before you inject. - Live-proof window — SAVE & INJECT really publishes the edit to the storefront for 90 seconds (
RESTORE_DWELL_MS) with a direct link to the affected product page, then restores the pristine template automatically (with sweep recovery if a session dies mid-window). - One-click demo — captures baseline, deletes the price node (guaranteed extraction kill), rides the genuine Bright Data repair to a verdict.
DEMO: gallery/lab.mp4
Scrape succeeded ≠ data is trustworthy. PULSE enforces the difference mechanically.
Every observation must clear:
| Check | Rule | Outcome on breach |
|---|---|---|
| Schema contract | Zod: typed fields — name, brand/model/variant/color, positive price, ISO-8601 scrapedAt, valid URLs, enum availability, 3-letter currency… |
FAILED → healing |
| Price rise band | > +200% vs last verified price | UNVERIFIED → source DEGRADED, held for review |
| Price crash band | ≥ −60% vs last verified price (drops are bounded by zero — without an explicit floor, a bogus "90% off" would sail through a symmetric rule) | UNVERIFIED → held for review |
| Currency drift | Changed vs last verified | warning |
The comparison baseline is always the last VERIFIED observation — untrusted data never becomes the reference point.
Thresholds live in one file (src/lib/validation/rules.ts) shared by the server-side gate and the Lab's edit advisor, so what the UI promises is exactly what the engine enforces.
Source health is a first-class state machine: HEALTHY · DEGRADED · HEALING · VERIFYING · RECOVERED · FAILED · FAILED_PERMANENTLY, each mapped to a health score (100 → 0).
Monitors dashboard — health state and event timeline per tracked source.
The Healing Observatory (/observatory) renders the immutable audit trail: trigger reason, the actual prompt sent to Bright Data, the repair response, approval status, verification result, and final verdict for every healing event.
Healing Observatory — every break and recovery, permanently on record.
Live search & monitoring — search products via Web Unlocker (Amazon/Flipkart parsers + Google offers grid), create monitors with per-source tracking, configure threshold alerts delivered by email (Resend) or webhook.
Monitor detail — price history, source health badges, healing timeline, alert config.
How It Works (/how-it-works) — seven scroll-paced chapters (COLLECT → VALIDATE → DETECT → HEAL → RE-RUN → VERIFY → TRUST) that mirror the pipeline table above, so the explainer and the code can't drift apart.
Prefer hands-on? The fastest 3-minute tour on the live site:
Landing (scroll the film) → Monitors → open a monitor → Resilience Lab
→ inject a break → watch validation fail → watch Bright Data heal it
→ read the verdict in the Observatory
Note: healing calls Bright Data's real AI repair and typically completes in minutes — the demo video compresses the wait.
Real row from PULSE's price_observations table — the normalized, validated result of a collector run against the Pulse Mart storefront (recovered during a healing cycle):
{
"monitor_id": "c9e2461e-852a-4ea7-acf4-997e34bdcd06",
"monitor_source_id": "04d3ebb8-ade6-4cf7-aaa8-995d8887112e",
"product_id": "558de941-1942-4928-b03a-0833cf47b75a",
"price": 154999,
"currency": "INR",
"availability": "in_stock",
"image_url": null,
"product_url": "https://pulse-v1-beige.vercel.app/pulse_store/p/samsung-galaxy-s26-ultra-5g-cobalt-violet-12gb-ram-256gb-sto",
"source": "RESILIENCE-LAB · samsung-galaxy-s26-ultra-5g-cobalt-violet-12gb-ram-256gb-sto",
"observed_at": "2026-08-23T14:22:33.826+00:00",
"validation_status": "VERIFIED",
"raw_data": {
"input": { "url": "https://pulse-v1-beige.vercel.app/pulse_store" },
"product_url": "https://pulse-v1-beige.vercel.app/pulse_store/p/samsung-galaxy-s26-ultra-5g-cobalt-violet-12gb-ram-256gb-sto",
"availability": "In Stock · ships in 24h",
"product_name": "samsung galaxy s26",
"current_price": 154999,
"product_page_url": "https://pulse-v1-beige.vercel.app/pulse_store"
}
}What matters here:
validation_status: "VERIFIED"— this row cleared the Zod contract and sat inside both anomaly bands relative to the previous verified observation. Only rows with this badge drive charts, alerts, and baselines.raw_data— the untouched collector output preserved verbatim for audit; normalization never destroys the original.observed_at+source— every datapoint is traceable to a specific collector run on a specific source at a specific time.
The companion audit record (healing_events) stores the trigger reason, the exact prompt sent to Bright Data's repair API, the response, approval status, and the post-heal verification result.
flowchart TB
subgraph UI["PULSE UI — Next.js App Router"]
L["Cinematic landing"]
S["Live search"]
M["Monitors + monitor detail"]
LAB["Resilience Lab<br/>(retailer + store modes)"]
PS["Pulse Mart storefront"]
OB["Healing Observatory"]
end
subgraph CORE["PULSE Core — API routes + engines"]
API["/api/monitors · /api/lab/*<br/>/api/healing · /api/products"]
ENG["Monitoring engine"]
VAL["Validation gate<br/>Zod + anomaly rules"]
HEAL["Healing engine"]
ALERT["Alerts — Resend + webhooks"]
end
subgraph BD["Bright Data"]
SC["Scraper Studio custom collectors<br/>AMAZON · FLIPKART · WEB<br/>+ isolated Lab clones"]
WU["Web Unlocker zone"]
AIR["AI scraper healing API"]
end
WEB["Public web pages<br/>(retailers — and Pulse Mart itself)"]
DB[("Supabase PostgreSQL<br/>13 tables, RLS on")]
CRON["node-cron worker<br/>(every minute)"]
UI --> API
CRON -->|"POST /api/internal/run-source"| API
API --> ENG
ENG --> VAL
VAL -->|"FAILED"| HEAL
ENG --> ALERT
ENG --> SC
HEAL --> AIR
HEAL --> SC
S --> WU
SC --> WEB
WU --> WEB
ENG --> DB
HEAL --> DB
VAL --> DB
Runtime split: the deployed Vercel app serves the UI, search, dashboards, and observatory; the cron worker (npm run monitor:worker) drives scheduled runs through /api/internal/run-source; the full long-running healing demo runs locally where function duration is unbounded.
| Layer | Technology |
|---|---|
| Frontend | Next.js 14 (App Router), React 18, TypeScript 5, Tailwind CSS, custom canvas scroll-cinema (258 per-frame WebPs) |
| Backend | Next.js API routes, tsx scripts, node-cron worker |
| Database | Supabase PostgreSQL (RLS enabled on all tables) |
| Extraction & healing | Bright Data — custom Scraper Studio collectors + Web Unlocker + AI healing APIs (direct HTTPS) |
| Validation | Zod schemas + shared plausibility rules |
| Notifications | Resend (email) + generic webhooks |
| Testing | Playwright — 21 E2E specs across desktop Chrome and mobile Chrome (landscape + portrait) |
| Assets tooling | sharp (sprite-sheet splitting script) |
Thirteen Supabase tables, RLS on everywhere:
products ──< product_variants
│
└──< monitors ──< monitor_sources ──< scraper_collectors
│ │
│ collector_id →
│ Bright Data Scraper Studio
├──< price_observations (validated datapoints)
└──< scrape_runs (every attempt, pass or fail)
healing_events (break → repair → verify audit trail)
alerts (threshold + channel) ──< notification_events
lab_sessions ──< lab_events (Resilience Lab recordings)
store_pages (Pulse Mart catalog + served edits)
Flow of one check: scrape_runs records the attempt → adapter normalizes → gate scores it → price_observations stores only schema-valid rows tagged VERIFIED / UNVERIFIED → failures branch into healing_events → sources carry status + health score forward.
- Open https://pulse-v1-beige.vercel.app/ — scroll the landing film; it ends by dropping you into live search.
- Search a product (e.g. "iPhone 15") — results are fetched live through Bright Data.
- Create a monitor from any result; open its detail page — price history, source health, alert setup.
- Visit the Resilience Lab — pick a mode, capture the baseline, inject a break (REMOVE_FIELD on price is the cleanest kill).
- Watch validation fail, then Bright Data's repair run. This is real AI healing — give it a few minutes.
- Read the verdict: recovered data re-passed the same gate; source flipped to
RECOVERED. - Close the loop in the Observatory — the entire incident, permanently recorded.
Short on time? gallery/lab.mp4 shows exactly this.
src/
├── app/
│ ├── page.tsx # 258-frame cinematic landing
│ ├── how-it-works/ # 7-chapter pipeline explainer
│ ├── monitor/[id]/ # Monitor detail — trust view
│ ├── monitors/ # Dashboard
│ ├── lab/ # Resilience Lab (retailer + store modes)
│ ├── observatory/ # Healing audit trail
│ ├── pulse_store/ # Live storefront (+ product pages)
│ └── api/
│ ├── monitors/[id]/run # Trigger a real check
│ ├── internal/run-source # Cron entrypoint
│ ├── lab/{session,run,history}
│ ├── lab/store/{session,inject,heal-step,reset,products}
│ ├── healing · products · alerts · frame[s]/[n]
├── components/
│ ├── pulse/ # Landing film, price chart, healing timeline, health states
│ ├── lab/ # RetailerLab · StoreLab panels
│ └── ui/ # Primitives
└── lib/
├── brightdata/ # HTTP client · collectors · unlocker · details · google-search
├── adapters/ # Per-source normalization
├── validation/ # schemas.ts · engine.ts · rules.ts · edit-advisor.ts
├── healing/engine.ts # Heal state machine + stale-heal reconciliation
├── monitoring/engine.ts # scrape → normalize → validate → store → alert
├── lab/ # pipelines · mutations · session store
├── notifications/alerts.ts # Resend + webhook delivery
├── store/ # Pulse Mart render/service/origin
└── db/supabase.ts # Cache-bypassing Supabase client
scripts/
├── scheduler.ts # node-cron worker (every minute)
├── run-monitor.ts # One-shot monitor sweep
├── production-gate.ts # Env + DB + typecheck + build + secret scan
└── split-sprite-sheets.mjs # sharp asset pipeline for the landing film
tests/ # Playwright E2E suites (21 specs × 3 browser projects)
gallery/ # Screenshots + demo video (lab.mp4)
— AI coding assistants were used in building PULSE:
- Tools: (agentic coding CLIs).
- What they did: scaffolding files and boilerplate, implementing components and modules under direction, debugging assistance, test drafting, refactoring help, and documentation drafts.
- What the human did: product and architecture decisions, the honesty/trust design of the validation gate, integration of the Bright Data APIs, running every real end-to-end verification against live services (including actual healing cycles), reviewing and correcting AI-generated code, and final understanding of the whole system.
- AI was an accelerator, not an autopilot: nothing shipped unreviewed, and every claim in this README was verified against the running application.
No license file has been added yet. All rights reserved by the author until one is chosen.
PULSE doesn't just heal scrapers. It proves they healed.





