Add Linda-RAID submission - #208
Lendarixon wants to merge 3 commits into
Conversation
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Eval run succeeded! Link to run: link Here are the results of the submission(s): Linda-RAIDRelease date: 2026-09-27 I've committed detailed results of this detector's performance on the test set to this PR. On the RAID dataset as a whole (aggregated across all generation models, domains, decoding strategies, repetition penalties, and adversarial attacks), it achieved an AUROC of 99.83 and a TPR of 99.48% at FPR=5% and 97.66% at FPR=1%. If all looks well, a maintainer will come by soon to merge this PR and your entry/entries will appear on the leaderboard. If you need to make any changes, feel free to push new commits to this PR. Thanks for submitting to RAID! |
Adds
leaderboard/submissions/Linda-RAID/(predictions.json for all 672,000 test ids + metadata.json).Model: https://huggingface.co/Lindarixon/Linda-RAID
Detector: an ensemble of two models trained on a sample of the RAID train split (all 8 domains, 11 generators, 12 attacks; human texts of every attack labelled human), with harder strata oversampled (sampling without repetition penalty, paraphrase attack):
Scores of the two models are z-normalized on human texts and averaged. Before scoring, input text is normalized against character-level attacks (homoglyphs, zero-width characters, unusual whitespace).
Local validation: 15% of RAID train source documents were held out (no text of those documents was used in training); 1,928 clean human texts + 40,000 AI texts from them, scored with
raid.evaluate.run_evaluation(per-domain thresholds, FPR 5%): 99.7% accuracy with adversarial attacks, 99.6% without (DeBERTa-v3-large alone 98.4%, stylometric model alone 98.4%).🤖 Generated with Claude Code