Fine-tuning workspace for the MiniLM encoder used as a borderline-case second opinion in spam-detector.
A quantized ONNX INT8 model (model-onnx/model_quantized.onnx) that classifies email text as spam or ham with a calibrated probability score. Runs in 1–5ms on CPU.
prepare_data.py
│
▼ data/train.csv + data/eval.csv
train_spam_encoder.py
│
▼ model/ (PyTorch checkpoint)
export_and_quantize.py
│
▼ model-onnx/model_quantized.onnx
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txtDownloads and preprocesses a public spam dataset into data/train.csv and data/eval.csv.
python prepare_data.pyDefault dataset: Enron-Spam (~33,000 emails, via HuggingFace SetFit/enron_spam).
For a smaller dataset (SMS Spam Collection, ~5,500 messages):
python prepare_data.py --dataset smsOutput columns: text (subject + body, HTML-stripped, ≤512 chars), label (0=ham, 1=spam).
Expected output:
data/train.csv— ≥5,000 rowsdata/eval.csv— ≥1,000 rows- Label distribution printed (target: 40–60% spam)
python train_spam_encoder.py \
--train_csv data/train.csv \
--eval_csv data/eval.csv \
--model_name microsoft/MiniLM-L12-H384-uncased \
--output_dir ./model \
--epochs 3 \
--batch_size 16 \
--lr 2e-5Swap --model_name for distilbert-base-uncased as a fallback if F1 < 0.95.
Target eval metrics:
- F1 ≥ 0.95
- Precision ≥ 0.93 (minimise false positives)
- Recall ≥ 0.95
Output: model/ (PyTorch checkpoint + tokenizer)
python export_and_quantize.py \
--model_dir ./model \
--output_dir ./model-onnxOutput: model-onnx/model_quantized.onnx + tokenizer files (≤40MB total).
Once model-onnx/ exists, spam-detector/encoder_classifier.py picks it up automatically from the default path ./spam-classifier/model-onnx/ (relative to the project root).
To use a custom path:
export ENCODER_MODEL_DIR=/path/to/model-onnxThe encoder is invoked automatically on every borderline case — emails where TypeSafe fires exactly SPAM_SIGNAL_THRESHOLD signals. Its verdict and confidence are printed to stdout and logged to metrics.csv in spam-detector/.
spam-classifier/
├── data/
│ ├── train.csv # generated by prepare_data.py
│ └── eval.csv # generated by prepare_data.py
├── model/ # generated by train_spam_encoder.py
├── model-onnx/ # generated by export_and_quantize.py
│ └── model_quantized.onnx
├── prepare_data.py
├── train_spam_encoder.py
├── export_and_quantize.py
├── requirements.txt
└── README.md