A systematic stock screening tool that identifies potentially undervalued companies using SEC EDGAR financial data, sector-based valuation comparisons, and momentum filters. Runs on a daily automated pipeline (AWS EventBridge → Lambda → EC2) and ships as Docker images for reproducible local or cloud execution.
- 310 signals across 500 companies over 24 months (June 2024 - June 2026)
- 57.4% win rate on 30-day holds
- +3.82% average return per signal
- +2.37% alpha vs S&P 500 over same periods
- 55.5% of signals beat the benchmark in the same window
- 1.93 annualized Sharpe ratio (monthly-rebalanced, equal-weighted)
- -9.42% max drawdown (February 2026)
- 48.8% CAGR over the backtest window
Limitations:
- Survivorship bias — universe is pulled from today's listed companies, so delisted/bankrupt companies were never eligible to be selected
- Look-ahead bias — share count used in valuation ratios is current, not historical
- Small sample — 310 signals across 24 rebalance periods
1. Deep Dive (single ticker) Input any US ticker → Claude AI identifies comparable peers → calculates EV, EV/Revenue, P/E, and growth for the ticker and all peers → flags if undervalued vs peer group
2. Auto Screener Scans 500 NYSE/Nasdaq companies automatically → groups by industry → calculates sector medians → flags companies trading 30%+ below sector median with strong growth
A company must follow these criteria to be flagged:
- EV/Revenue 30%+ below sector median
- Revenue growth >10% YoY
- Above 50-day moving average (positive momentum)
- EV under $500B (excludes mega caps)
- Analyst price target >10% upside
- Analyst consensus bullish
- SEC EDGAR API — live financial data (revenue, income, debt, cash) from 10-K filings
- yfinance — stock prices, MA50, analyst targets
- Finnhub — analyst recommendation data
- Anthropic Claude API — peer company identification (deep dive mode)
- Streamlit — web dashboard
- pandas / numpy — valuation calculations and sector medians
- concurrent.futures — parallel API requests for faster screening
- Docker — containerized backtest and pipeline execution
- AWS (EventBridge, Lambda, EC2) — daily automated pipeline
The screener runs automatically once a day without any manual intervention:
EventBridge (13:30 UTC daily)
│
▼
Lambda function
(starts EC2 instance via boto3)
│
▼
EC2 instance boots
│
▼
Docker container runs pipeline stages in order:
screener → signals → snapshot → pnl
│
▼
Instance self-terminates (avoids idle EC2 cost)
This design means the pipeline costs nothing to run when it's not executing — the instance only exists while the job is active, and the daily trigger requires zero manual steps.
Two purpose-built images, since the backtest and the live pipeline have different jobs and different runtime needs:
| Image | Purpose | Run with |
|---|---|---|
Dockerfile.backtest |
Runs the historical backtest engine in isolation | docker build -f Dockerfile.backtest -t valuescan-backtest . |
Dockerfile.ec2 |
Runs the live pipeline stages via main.py, mirrors what executes on EC2 |
docker build -f Dockerfile.ec2 -t valuescan-pipeline . |
Both use multi-stage builds (slim final image, no build tools carried over) and run as a non-root user.
The pipeline image uses main.py as its entrypoint, so each stage is a separate, explicit command:
docker run --rm valuescan-pipeline screener # scan 500 companies for signals
docker run --rm valuescan-pipeline signals # check sell signals
docker run --rm valuescan-pipeline snapshot # save portfolio snapshot
docker run --rm valuescan-pipeline pnl # view current positionsThis is the same sequence that runs automatically on EC2 every day — the container makes that behavior reproducible on any machine, not just the deployed instance.
valuescan/
├── src/
│ ├── data/
│ │ ├── edgar.py # SEC EDGAR data pipeline
│ │ └── universe.py # NYSE/Nasdaq ticker universe
│ ├── analysis/
│ │ ├── screener.py # Auto screener (500 companies)
│ │ ├── multiples.py # EV, EV/Revenue, P/E calculations
│ │ ├── backtest.py # Historical backtesting
│ │ └── peers.py # Claude AI peer identification
│ ├── signals/
│ │ └── trading.py # Signal logic and filters
│ └── portfolio/
│ └── pnl.py # P&L tracking
├── data/ # Data files
├── Dockerfile.backtest # Container for backtest engine
├── Dockerfile.ec2 # Container for daily pipeline (matches EC2 deployment)
├── app.py # Streamlit web UI
└── main.py # CLI entry point
1. Clone the repo
git clone https://github.com/foley40/valuescan
cd valuescan2. Create virtual environment
python3 -m venv venv
source venv/bin/activate3. Install dependencies
pip install -r requirements.txt4. Configure API keys
Create a .env file:
ANTHROPIC_API_KEY=your_key_here
FINNHUB_API_KEY=your_key_here
CLI:
python3 main.py screener # scan 500 companies for signals
python3 main.py pnl # view current positions
python3 main.py signals # check sell signals
python3 main.py snapshot # save portfolio snapshotWeb UI:
streamlit run app.pyDocker (matches production behavior):
docker build -f Dockerfile.ec2 -t valuescan-pipeline .
docker run --rm valuescan-pipeline screenerSector-based comparison over peer-based The screener uses industry classifications to group companies and calculates median EV/Revenue per sector. This avoids expensive AI calls for every ticker while providing meaningful comparisons.
Point-in-time backtesting The backtest uses SEC EDGAR filing dates to filter financials to what was available at each historical date, avoiding look-ahead bias in the fundamental data.
Parallel requests with rate limit protection The screener uses ThreadPoolExecutor with 3 workers for parallel data fetching. A threading lock on Finnhub calls prevents rate limit errors while keeping EDGAR and yfinance calls parallel.
Self-terminating EC2 over always-on infrastructure Rather than keeping a server running continuously, Lambda starts the EC2 instance only when the daily job needs to run, and the instance terminates itself once finished. This keeps cloud costs near zero for a job that only needs a few minutes of compute per day.
Working with yfinance and running into a few bugs inspired me to look more into the repo and contribute. Here are some of my merged PRs:
- yfinance PR #2879 — Added missing balance sheet fields for financial sector tickers
- yfinance PR #2906 — Fixed TypeError when Yahoo API returns null result in _fetch_info
- yfinance PR #2907 — Fixed currency conversion bug incorrectly rescaling subunit-quoted prices (GBp/ZAc/ILA) that didn't need repair
- yfinance PR #2933 - Fixed NaN equality comparison in test suite by replacing floating-point
==withnp.isclose(..., equal_nan=True)