Skip to content

Repository files navigation

valuescreen - Automated Equity Screener & Research Tool

What it is

A systematic stock screening tool that identifies potentially undervalued companies using SEC EDGAR financial data, sector-based valuation comparisons, and momentum filters. Runs on a daily automated pipeline (AWS EventBridge → Lambda → EC2) and ships as Docker images for reproducible local or cloud execution.

Backtest Results

  • 310 signals across 500 companies over 24 months (June 2024 - June 2026)
  • 57.4% win rate on 30-day holds
  • +3.82% average return per signal
  • +2.37% alpha vs S&P 500 over same periods
  • 55.5% of signals beat the benchmark in the same window
  • 1.93 annualized Sharpe ratio (monthly-rebalanced, equal-weighted)
  • -9.42% max drawdown (February 2026)
  • 48.8% CAGR over the backtest window

Limitations:

  • Survivorship bias — universe is pulled from today's listed companies, so delisted/bankrupt companies were never eligible to be selected
  • Look-ahead bias — share count used in valuation ratios is current, not historical
  • Small sample — 310 signals across 24 rebalance periods

How it Works

Two Modes

1. Deep Dive (single ticker) Input any US ticker → Claude AI identifies comparable peers → calculates EV, EV/Revenue, P/E, and growth for the ticker and all peers → flags if undervalued vs peer group

2. Auto Screener Scans 500 NYSE/Nasdaq companies automatically → groups by industry → calculates sector medians → flags companies trading 30%+ below sector median with strong growth

Signal Filters

A company must follow these criteria to be flagged:

  1. EV/Revenue 30%+ below sector median
  2. Revenue growth >10% YoY
  3. Above 50-day moving average (positive momentum)
  4. EV under $500B (excludes mega caps)
  5. Analyst price target >10% upside
  6. Analyst consensus bullish

Tech Stack

  • SEC EDGAR API — live financial data (revenue, income, debt, cash) from 10-K filings
  • yfinance — stock prices, MA50, analyst targets
  • Finnhub — analyst recommendation data
  • Anthropic Claude API — peer company identification (deep dive mode)
  • Streamlit — web dashboard
  • pandas / numpy — valuation calculations and sector medians
  • concurrent.futures — parallel API requests for faster screening
  • Docker — containerized backtest and pipeline execution
  • AWS (EventBridge, Lambda, EC2) — daily automated pipeline

Infrastructure

The screener runs automatically once a day without any manual intervention:

EventBridge (13:30 UTC daily)
        │
        ▼
  Lambda function
  (starts EC2 instance via boto3)
        │
        ▼
  EC2 instance boots
        │
        ▼
  Docker container runs pipeline stages in order:
  screener → signals → snapshot → pnl
        │
        ▼
  Instance self-terminates (avoids idle EC2 cost)

This design means the pipeline costs nothing to run when it's not executing — the instance only exists while the job is active, and the daily trigger requires zero manual steps.

Docker

Two purpose-built images, since the backtest and the live pipeline have different jobs and different runtime needs:

Image Purpose Run with
Dockerfile.backtest Runs the historical backtest engine in isolation docker build -f Dockerfile.backtest -t valuescan-backtest .
Dockerfile.ec2 Runs the live pipeline stages via main.py, mirrors what executes on EC2 docker build -f Dockerfile.ec2 -t valuescan-pipeline .

Both use multi-stage builds (slim final image, no build tools carried over) and run as a non-root user.

The pipeline image uses main.py as its entrypoint, so each stage is a separate, explicit command:

docker run --rm valuescan-pipeline screener   # scan 500 companies for signals
docker run --rm valuescan-pipeline signals    # check sell signals
docker run --rm valuescan-pipeline snapshot   # save portfolio snapshot
docker run --rm valuescan-pipeline pnl        # view current positions

This is the same sequence that runs automatically on EC2 every day — the container makes that behavior reproducible on any machine, not just the deployed instance.

Project Structure

valuescan/
├── src/
│   ├── data/
│   │   ├── edgar.py        # SEC EDGAR data pipeline
│   │   └── universe.py     # NYSE/Nasdaq ticker universe
│   ├── analysis/
│   │   ├── screener.py     # Auto screener (500 companies)
│   │   ├── multiples.py    # EV, EV/Revenue, P/E calculations
│   │   ├── backtest.py     # Historical backtesting
│   │   └── peers.py        # Claude AI peer identification
│   ├── signals/
│   │   └── trading.py      # Signal logic and filters
│   └── portfolio/
│       └── pnl.py          # P&L tracking
├── data/                   # Data files
├── Dockerfile.backtest      # Container for backtest engine
├── Dockerfile.ec2           # Container for daily pipeline (matches EC2 deployment)
├── app.py                  # Streamlit web UI
└── main.py                 # CLI entry point

Setup

1. Clone the repo

git clone https://github.com/foley40/valuescan
cd valuescan

2. Create virtual environment

python3 -m venv venv
source venv/bin/activate

3. Install dependencies

pip install -r requirements.txt

4. Configure API keys Create a .env file: ANTHROPIC_API_KEY=your_key_here FINNHUB_API_KEY=your_key_here

Usage

CLI:

python3 main.py screener    # scan 500 companies for signals
python3 main.py pnl         # view current positions
python3 main.py signals     # check sell signals
python3 main.py snapshot    # save portfolio snapshot

Web UI:

streamlit run app.py

Docker (matches production behavior):

docker build -f Dockerfile.ec2 -t valuescan-pipeline .
docker run --rm valuescan-pipeline screener

Key Design Decisions

Sector-based comparison over peer-based The screener uses industry classifications to group companies and calculates median EV/Revenue per sector. This avoids expensive AI calls for every ticker while providing meaningful comparisons.

Point-in-time backtesting The backtest uses SEC EDGAR filing dates to filter financials to what was available at each historical date, avoiding look-ahead bias in the fundamental data.

Parallel requests with rate limit protection The screener uses ThreadPoolExecutor with 3 workers for parallel data fetching. A threading lock on Finnhub calls prevents rate limit errors while keeping EDGAR and yfinance calls parallel.

Self-terminating EC2 over always-on infrastructure Rather than keeping a server running continuously, Lambda starts the EC2 instance only when the daily job needs to run, and the instance terminates itself once finished. This keeps cloud costs near zero for a job that only needs a few minutes of compute per day.

Open Source Contributions

Working with yfinance and running into a few bugs inspired me to look more into the repo and contribute. Here are some of my merged PRs:

  • yfinance PR #2879 — Added missing balance sheet fields for financial sector tickers
  • yfinance PR #2906 — Fixed TypeError when Yahoo API returns null result in _fetch_info
  • yfinance PR #2907 — Fixed currency conversion bug incorrectly rescaling subunit-quoted prices (GBp/ZAc/ILA) that didn't need repair
  • yfinance PR #2933 - Fixed NaN equality comparison in test suite by replacing floating-point == with np.isclose(..., equal_nan=True)

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages