Add Aletheia HumanEval / HumanEval+ and MBPP / MBPP+ official evaluate scores - #27
Add Aletheia HumanEval / HumanEval+ and MBPP / MBPP+ official evaluate scores#27shaneraphel wants to merge 2 commits into
Conversation
|
Official |
|
Official |
|
Public page: https://github.com/shaneraphel/aletheia Completions (samples.jsonl, 378 rows): https://gist.github.com/shaneraphel/b0fbe4cde5b83185da043c9e0ac00fd0 SHA-256: ba41a115da629ac4bf55213d4e29e4545994e9a014bbee50fc1f4c384dd9f3b1 |
Co-authored-by: Cursor <cursoragent@cursor.com>
a3db743 to
e731850
Compare
Co-authored-by: Cursor <cursoragent@cursor.com>
|
HumanEval / HumanEval+ numbers are now filled from the same official
MBPP / MBPP+ is unchanged: 100.0 / 87.0. |
Summary
results.json.linkis the public page: https://github.com/shaneraphel/aletheiaevalplus.evaluate --dataset humanevalrun: https://github.com/shaneraphel/aletheia/blob/main/humaneval/samples.jsonlevalplus.evaluate --dataset mbpprun: https://gist.github.com/shaneraphel/b0fbe4cde5b83185da043c9e0ac00fd0open-dataisNONE.Test plan
pip install evalplus==0.3.1 && evalplus.evaluate --dataset humaneval --samples humaneval/samples.jsonlpip install evalplus==0.3.1 && evalplus.evaluate --dataset mbpp --samples samples.jsonlresults.jsonstill parses and the leaderboard page renders