Skip to content

Add Linda-RAID submission - #208

Open
Lendarixon wants to merge 3 commits into
liamdugan:mainfrom
Lendarixon:add-linda-raid
Open

Lendarixon wants to merge 3 commits into
liamdugan:mainfrom
Lendarixon:add-linda-raid

Conversation

@Lendarixon

@Lendarixon Lendarixon commented Sep 27, 2026 •

Copy link
Copy Markdown

Adds leaderboard/submissions/Linda-RAID/ (predictions.json for all 672,000 test ids + metadata.json).

Model: https://huggingface.co/Lindarixon/Linda-RAID

Detector: an ensemble of two models trained on a sample of the RAID train split (all 8 domains, 11 generators, 12 attacks; human texts of every attack labelled human), with harder strata oversampled (sampling without repetition penalty, paraphrase attack):

  • DeBERTa-v3-large fine-tuned on ~590k RAID train texts (256-token windows covering the whole document, mean logit margin);
  • a stylometric model (no neural network): ~230 hand-crafted style features + hashed character/word/function-word n-grams, LightGBM + logistic regression, trained on 200k RAID train texts (hard strata oversampled) on CPU.

Scores of the two models are z-normalized on human texts and averaged. Before scoring, input text is normalized against character-level attacks (homoglyphs, zero-width characters, unusual whitespace).

Local validation: 15% of RAID train source documents were held out (no text of those documents was used in training); 1,928 clean human texts + 40,000 AI texts from them, scored with raid.evaluate.run_evaluation (per-domain thresholds, FPR 5%): 99.7% accuracy with adversarial attacks, 99.6% without (DeBERTa-v3-large alone 98.4%, stylometric model alone 98.4%).

🤖 Generated with Claude Code

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown

Eval run succeeded! Link to run: link

Here are the results of the submission(s):

Linda-RAID

Release date: 2026-09-27

I've committed detailed results of this detector's performance on the test set to this PR.

On the RAID dataset as a whole (aggregated across all generation models, domains, decoding strategies, repetition penalties, and adversarial attacks), it achieved an AUROC of 99.83 and a TPR of 99.48% at FPR=5% and 97.66% at FPR=1%.
Without adversarial attacks, it achieved AUROC of 99.89 and a TPR of 99.65% at FPR=5% and 98.46% at FPR=1%.

If all looks well, a maintainer will come by soon to merge this PR and your entry/entries will appear on the leaderboard. If you need to make any changes, feel free to push new commits to this PR. Thanks for submitting to RAID!

This branch was successfully deployed

1 active (outdated) deployment
raid-main — 57a3c2b2 Deployed Sep 27, 2026 by Lendarixon via evaluate #352
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant