The backend service for OG Miner, a high-performance URL metadata extraction API. Built with FastAPI, Playwright, and Redis.
- Metadata Extraction: Retrieves OpenGraph, Twitter Cards, JSON-LD, and oEmbed data.
- Headless Browser: Integrated Playwright support to render JavaScript-heavy websites (SPAs).
- High Performance:
- Redis Caching: Caches results to serve repeat requests in <10ms.
- Async I/O: Fully asynchronous architecture using
httpxanduvicorn.
- Security:
- SSRF Protection: Validators to prevent internal network scanning.
- Rate Limiting: Built-in rate limiting using
slowapi. - API Key Auth: Validates requests via
X-RapidAPI-Proxy-Secret.
- Anti-Blocking:
- Proxy Rotation: Automatically rotates free proxies (or uses your paid proxy) to avoid IP bans.
- Geo-Targeting (Beta): Supports
countryparameter for region-specific scraping (requiresPROXY_URLto be set). - BYOP (Pro): Bring Your Own Proxy. Pass a
proxyURL per request to use your own premium proxies. - Cookies: Supports passing session cookies for authenticated scraping.
- Developer Experience:
- Batch Processing: Process up to 50 URLs in parallel.
- Image Proxy: Securely resize and cache external images (WebP format).
- Screenshots: Capture full-page or viewport screenshots.
- Language: Python 3.10+
- Framework: FastAPI (High-performance web framework)
- Server: Uvicorn (ASGI server)
- Headless Browser: Playwright (Chromium) for SPA/JS rendering.
- HTTP Client: HTTPX (Async HTTP client).
- Parsers:
extruct: For Schema.org (JSON-LD, Microdata) & OpenGraph.beautifulsoup4: For fallback meta tag parsing.
- Caching: Redis (Key-value store for <10ms response times).
- Rate Limiting:
slowapi(In-memory or Redis-backed). - Task Queue:
BackgroundTasks(FastAPI native) for async batch processing.
og-miner-api/
├── app/
│ ├── api/
│ │ ├── v1/ # Route handlers (Extract, Batch, Image, Screenshot)
│ │ └── dependencies.py # Auth & Rate Limiting dependencies
│ ├── core/
│ │ ├── config.py # Environment configuration
│ │ └── security.py # SSRF protection & validation
│ ├── schemas/ # Pydantic models (Request/Response schemas)
│ ├── services/ # Core Business Logic
│ │ ├── extract.py # Main extraction orchestrator
│ │ ├── fetcher.py # Async HTTP fetcher
│ │ ├── headless.py # Playwright manager
│ │ ├── image_proxy.py # Image resizing & proxying
│ │ └── parser.py # HTML parsing logic
│ └── main.py # App entrypoint
├── tests/ # Unit tests (Pytest)
├── Dockerfile # Docker build
└── requirements.txt # Python dependencies
- Python 3.10+
- Redis (running locally or via Docker)
- Google Chrome (for Playwright)
-
Clone the repository:
git clone https://github.com/yourusername/og-miner.git cd og-miner/og-miner-api -
Create a virtual environment:
python -m venv venv source venv/bin/activate # On Windows: venv\Scripts\activate
-
Install dependencies:
pip install -r requirements.txt playwright install chromium
-
Configure Environment: Create a
.envfile based on.env.example:PROJECT_NAME="og-miner" VERSION="1.0.0" REDIS_URL="redis://localhost:6379/0" SECRET_KEY="your-secret-key" X_RAPIDAPI_PROXY_SECRET="your-rapidapi-secret" PROXY_URL="" # Optional: "http://user:pass@host:port" (Leave empty for free proxy rotation) LOG_LEVEL="INFO"
-
Run the Server:
uvicorn app.main:app --reload
The API will be available at
http://localhost:8000.
Run the entire stack (API + Redis) with Docker Compose:
docker-compose up --build-
Login:
heroku login heroku container:login
-
Create App & Add Redis:
heroku create og-miner-api heroku addons:create heroku-redis:mini -a og-miner-api
-
Set Secrets:
heroku config:set SECRET_KEY="your-secret" X_RAPIDAPI_PROXY_SECRET="your-key" -a og-miner-api
-
Deploy:
heroku container:push web -a og-miner-api heroku container:release web -a og-miner-api
POST /v1/extract
Extracts metadata from a single URL.
Body:
{
"url": "https://netflix.com/title/80057281",
"enable_javascript": false,
"force_refresh": false,
"country": "US", // Optional: Geo-target (Beta, requires paid PROXY_URL)
"proxy": "http://user:pass@host:port", // Optional: BYOP (Pro)
"cookies": { // Optional: Authenticated scraping
"netflixId": "v=2&ct=..."
}
}POST /v1/batch/extract
Process up to 50 URLs in parallel.
Body:
{
"urls": ["https://google.com", "https://apple.com"],
"enable_javascript": false
}GET /v1/image
Securely proxies, resizes, and caches images.
Query Parameters:
url: The image URL (required).width: Target width (optional, e.g.,200).height: Target height (optional).
Headers:
X-RapidAPI-Proxy-Secret: Your secret key (required).
Example:
GET /v1/image?url=https://example.com/logo.png&width=300
POST /v1/screenshot
Captures a screenshot of the page.
Body:
{
"url": "https://example.com",
"full_page": false,
"dark_mode": true,
"delay": 2000
}The project maintains a strict test suite using pytest, covering:
- Unit Tests: Individual service logic (Extract, Batch, Image Proxy).
- Security: Verification of SSRF protection and Proxy validation.
- Integration: Mocked integration with Playwright and Redis.
Run the full suite:
pytest© 2026 OG Miner Backend.