Skip to content

Repository files navigation

spam-classifier

Fine-tuning workspace for the MiniLM encoder used as a borderline-case second opinion in spam-detector.


What this produces

A quantized ONNX INT8 model (model-onnx/model_quantized.onnx) that classifies email text as spam or ham with a calibrated probability score. Runs in 1–5ms on CPU.


Pipeline

prepare_data.py
      │
      ▼ data/train.csv + data/eval.csv
train_spam_encoder.py
      │
      ▼ model/  (PyTorch checkpoint)
export_and_quantize.py
      │
      ▼ model-onnx/model_quantized.onnx

Setup

python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt

Step 1 — Prepare training data

Downloads and preprocesses a public spam dataset into data/train.csv and data/eval.csv.

python prepare_data.py

Default dataset: Enron-Spam (~33,000 emails, via HuggingFace SetFit/enron_spam).

For a smaller dataset (SMS Spam Collection, ~5,500 messages):

python prepare_data.py --dataset sms

Output columns: text (subject + body, HTML-stripped, ≤512 chars), label (0=ham, 1=spam).

Expected output:

  • data/train.csv — ≥5,000 rows
  • data/eval.csv — ≥1,000 rows
  • Label distribution printed (target: 40–60% spam)

Step 2 — Fine-tune

python train_spam_encoder.py \
    --train_csv data/train.csv \
    --eval_csv data/eval.csv \
    --model_name microsoft/MiniLM-L12-H384-uncased \
    --output_dir ./model \
    --epochs 3 \
    --batch_size 16 \
    --lr 2e-5

Swap --model_name for distilbert-base-uncased as a fallback if F1 < 0.95.

Target eval metrics:

  • F1 ≥ 0.95
  • Precision ≥ 0.93 (minimise false positives)
  • Recall ≥ 0.95

Output: model/ (PyTorch checkpoint + tokenizer)


Step 3 — Export to ONNX INT8

python export_and_quantize.py \
    --model_dir ./model \
    --output_dir ./model-onnx

Output: model-onnx/model_quantized.onnx + tokenizer files (≤40MB total).


Integration with spam-detector

Once model-onnx/ exists, spam-detector/encoder_classifier.py picks it up automatically from the default path ./spam-classifier/model-onnx/ (relative to the project root).

To use a custom path:

export ENCODER_MODEL_DIR=/path/to/model-onnx

The encoder is invoked automatically on every borderline case — emails where TypeSafe fires exactly SPAM_SIGNAL_THRESHOLD signals. Its verdict and confidence are printed to stdout and logged to metrics.csv in spam-detector/.


Directory layout

spam-classifier/
├── data/
│   ├── train.csv               # generated by prepare_data.py
│   └── eval.csv                # generated by prepare_data.py
├── model/                 # generated by train_spam_encoder.py
├── model-onnx/            # generated by export_and_quantize.py
│   └── model_quantized.onnx
├── prepare_data.py
├── train_spam_encoder.py
├── export_and_quantize.py
├── requirements.txt
└── README.md

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages