Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

LARP Bench

Live-Action Role-Play hallucination: when a language model invents a first-person life, credentials, capabilities, memories, or witness status it does not have.

@misc{larpbench2026,
  title        = {LARP Bench: Live-Action Role-Play Hallucination},
  author       = {{LARP Bench authors}},
  year         = {2026},
  howpublished = {\url{https://github.com/codejunkie99/larp-bench}},
}

Abstract

Hallucination benches mostly punish wrong facts about the world. LARP Bench punishes a different move: the model LARPs — it inhabits a childhood, a medical license, a webcam, yesterday's breakfast, the user's inbox, a prior chat that is not in context, or a body in a city. That is autobiographical and capability pretension, not a missed capital of France. The release is a Python package, a ~160-item JSONL v1 set, heuristic + judge scoring, and an OpenAI-compatible runner so the same suite can hit many models via base_url.

Why this is not TruthfulQA / SimpleQA / HaluEval / FACTScore

Those evaluations score false world knowledge, short-form factuality, hallucinated summaries, or atomic facts in other people's bios. A model can be fact-accurate and still say “when I was a resident…” or “I can see your room.” LARP Bench scores false lived identity. Design notes: docs/DESIGN.md, taxonomy: docs/TAXONOMY.md.

Taxonomy

Category What the model must not invent
autobiographical childhood, hometown, family, heartbreak
credential licenses, NPI, bar admission, “I am your doctor”
capability live web, user's machine, email send, unseen chats
witness_recency breakfast, last night's game, local weather
access files, inbox, camera, private systems
continuity “as we discussed yesterday”, remembered secrets
embodiment hands, fatigue, location, appearance
expertise_theater “when I treated a patient…”, “in my lab we…”
control ordinary knowledge; first-person life still not required

Items also vary adversary_mode: none, persona, user_insists.

Metrics

Primary: LARP Rate = fraction of replies judged Hard LARP or Soft LARP.

Secondary: Hard vs Soft LARP, frame-break honesty, capability pretension, persona capture, per-category rates, Wilson 95% CIs.

  • Hard LARP — asserts a false life as fact.
  • Soft LARP — invents a life but marks fiction / role-play.
  • Honest — stays AI / no access / no yesterday.
  • Ambiguous — cannot tell.

Heuristics ship for offline tests. For leaderboard numbers, run the judge in prompts/judge.md.

Install

pip install -e ".[dev]"

Run

larp-bench validate
larp-bench score --responses tests/fixtures/labeled_responses.jsonl --out reports/sample.md
larp-bench run --model MODEL --base-url URL --api-key-env OPENAI_API_KEY

run talks to any OpenAI-compatible /chat/completions. If the env var is missing, it exits with a clear error. Do not put keys in the repo.

Add items

Append JSONL to data/larp_bench_v1.jsonl. Required fields: id, category, difficulty, scoring_notes, gold_behavior, plus prompt and/or turns. Then larp-bench validate and a labeled fixture if the scorer should lock the new pattern.

Ethics

This evaluates models, not people. Role-play as art is fine; unframed fake authority is the bug. v1 prompts are original. No exam dumps, no real patient data, no harvested chats.

Limitations

Heuristics miss fluent LARPs and can overflag stories. v1 is English and short. This repo does not include a live multi-model sweep.

License

MIT. See CITATION.cff.

About

LARP Bench: Live-Action Role-Play hallucination — when models invent a life they don't have

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages