Skip to content

Latest commit

 

History

93 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Privacy Risk Assessor

Tool for analyzing website behavior to surface privacy risks: third-party requests, cookies, frames, CNAME cloaking, and canvas fingerprinting attempts. Built to support a master's thesis; refactored to minimize npm dependencies.

Architecture

Area Implementation
Browser visits Puppeteer (only production dependency) — full Chromium session to simulate a real user
Storage SQLite via Node's built-in node:sqlite (./data/privacy-risk.db)
HTTP API Node http (no Express)
Reports Markdown or HTML (no Word/docx libraries)
WHOIS Raw TCP sockets on port 43
Registrable domains Vendored Public Suffix List snapshot (config/publicSuffixList.txt, refresh from https://publicsuffix.org/list/public_suffix_list.dat)
Tests Node built-in test runner

Modules live under modules/ (navigation, analytics, database, orchestration, api, docs, sites, configuration).

SQLite tables

  • sites — visited domains, visit metadata, cookies/frames/localStorage (JSON), privacy stats
  • owners / owner_sites — WHOIS registrant entities linked to sites
  • trackers / tracker_sites — observed request URLs linked to sites
  • cookies — cookies attributed to site/request
  • fingerprint_attempts — canvas/cookie instrumentation reports from the page

Requirements

  • Node.js ≥ 22.5 (for node:sqlite)
  • Network access for site visits and WHOIS

Setup

npm install
# Install Chrome for Puppeteer (do NOT use `npx puppeteer browsers install` —
# it breaks under this package's "type": "module". Use:)
npm run install-chrome

Public Suffix List

Third-party classification uses a vendored snapshot of the Public Suffix List at config/publicSuffixList.txt. Refresh it with:

curl -fsSL https://publicsuffix.org/list/public_suffix_list.dat -o config/publicSuffixList.txt

Optional: load the Latvian domain list into SQLite:

npm run import-sites
# or: node scripts/importLatvianList.js ./config/latvianList.txt

Run

npm start
# development with auto-restart:
npm run dev

API listens on http://127.0.0.1:3000 by default (httpServerHostname / httpServerPort in config/applicationConfig.json). Bind stays loopback-only unless you change the hostname intentionally.

Operator UI

Open http://127.0.0.1:3000/ for the zero-dependency console (vanilla HTML/CSS/JS):

  • Dashboard totals and queue depth
  • Queue a site visit / revisit stale sites / compare GPC impact (when enabled)
  • Browse sites, open detail (cookies incl. cookie profile, storage, frames, third-party by resource type / bytes / timing, identity, payloads, request graph), download Markdown/HTML reports
  • Cross-site identity links panel (same ID on ≥2 visited sites)
  • Link to the request graph at /graph

Startup revisits

On boot the collector only queues sites that need a visit:

Config Default Meaning
visitOnStartup true When false, no automatic crawl on start
revisitAfterDays 7 Revisit if visitDate is older than N days (or never visited / accessible still null). 0 = always revisit
enableGpcComparison false Enable POST /api/compare-gpc (two-pass visit with/without Sec-GPC: 1; doubles visit work per call)

Use Revisit stale sites in the UI (or POST /api/revisit-stale) to run the same policy without restarting.

Main endpoints

Method Path Description
GET / Operator UI
GET /graph Request graph UI
GET /api/status Config + queue / last revisit report
GET /api/sites All sites (full documents)
GET /api/sites/summary Compact site list for the UI
GET /api/sites/:id One site by numeric id
GET /api/sites/stats Per-site privacy stats
GET /api/sites/totals Aggregate totals
GET /api/sites/cookiesByDomain Cookies grouped by cookie domain
POST /api/visit Queue visit { "url", "wait?" }
POST /api/revisit-stale Queue all sites due under revisit policy
POST /api/sites/cleanup-failed Delete sites that failed visit (accessible=false); keeps pending rows
GET /api/sites/:base64Url Legacy queue visit (base64 http(s) URL)
POST /api/sites/fingerprint Fingerprint instrumentation ingest
GET /api/owners Site owners (WHOIS)
GET /api/trackers Observed request URLs (raw postData redacted)
GET /api/trackers/groupByDomain Request counts by domain
GET /api/payloads Per-site payload observation summaries
GET /api/sites/:id/payloads Payload summary + per-request field analysis
GET /api/identity/cross-site Same high-entropy value on ≥2 visited sites (derived from stored data)
POST /api/compare-gpc Two-pass visit: with vs without Sec-GPC: 1 (needs enableGpcComparison=true)
POST /api/docs/generate Markdown (default) or HTML report body

Observation model (no blocklists)

The crawler records behaviour during a real Chromium visit:

Channel What is collected
Network Requests/responses, query + body payloads when available, redirect chains, interesting headers (ETag, Attribution-Reporting, Client Hints, …), response size + CDP timing, cookies-transmitted count per request
Cookies Full jar (incl. Partitioned/CHIPS when exposed), Set-Cookie parse, disguise-name heuristics, profile (HTTP-set vs JS-set, session vs persistent, SameSite/HttpOnly/Secure split)
Storage localStorage, sessionStorage, IndexedDB inventory, Cache API samples, service workers
Fingerprinting Canvas / OffscreenCanvas / WebGL / Audio / UA-CH / WebRTC / voices / devices / navigator & screen reads, FNV-1a hash of canvas/readback values (uniqueness/stability without shipping raw pixels)
Privacy APIs Storage Access, Shared Storage, Topics hooks, sendBeacon, fenced frames when detectable
Identity Cross-domain ID linking — same high-entropy value in cookies/payloads/storage on ≥2 registrable domains; cross-site linking across stored visits (/api/identity/cross-site)
Visit realism Consent click (best-effort), scroll, optional same-origin second page, Sec-GPC / DNT headers, desktop viewport + rotating user agent, bot-wall detection (botBlocked flag from page-text phrases)

API extras:

  • GET /api/payloads, GET /api/sites/:id/payloads
  • GET /api/identity, GET /api/sites/:id/identity
  • GET /api/graph, GET /api/sites/:id/graph — request graph JSON (node size ∝ request count)
  • UI: open http://localhost:3000/ or /graph for an interactive force-directed graph (no extra npm deps)

Report body example:

{
  "firstName": "Ada",
  "lastName": "Lovelace",
  "site": "https://example.com",
  "findings": ["Third-party cookie _ga"],
  "format": "markdown"
}

Tests

npm test

License

MIT — see LICENSE.

About

A tool to identify potential privacy threats of websites

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages