Hi GraphRAG-Bench team — I maintain EvalPort, an open JSON-Schema spec for portable LLM eval data (TestCase/Grader/Result/ResultSet/GraderResult). No CONTRIBUTING.md, so filing as an issue.
I read Evaluation/retrieval_eval.py directly. The detailed per-sample record built in evaluate_dataset() — {"id", "question", "contexts", "evidences", "metrics"}, with metrics from evaluate_sample()'s {"context_relevancy", "evidence_recall"} — maps closely onto EvalPort: question→TestCase.input, evidences→expected_output/reference set, and each metric in metrics becomes a GraderResult. Since your benchmark already standardizes evaluation across multiple GraphRAG frameworks (LightRAG, etc.), exporting to a common interchange format could let someone compare a GraphRAG-Bench run against a non-graph RAG baseline scored by a different tool, on the same questions.
I'd propose a standalone graphragbench-openeval-adapter package (to_openeval(detailed_record) / from_openeval(suite)), tested against real output from retrieval_eval.py. Happy to open a PR if useful, or keep it as a separate package.
Spec: https://github.com/adhabnr-ux/evalport/blob/main/SPEC.md
— Sahi, independent contributor (not affiliated with GraphRAG-Bench)
Hi GraphRAG-Bench team — I maintain EvalPort, an open JSON-Schema spec for portable LLM eval data (TestCase/Grader/Result/ResultSet/GraderResult). No CONTRIBUTING.md, so filing as an issue.
I read
Evaluation/retrieval_eval.pydirectly. The detailed per-sample record built inevaluate_dataset()—{"id", "question", "contexts", "evidences", "metrics"}, withmetricsfromevaluate_sample()'s{"context_relevancy", "evidence_recall"}— maps closely onto EvalPort:question→TestCase.input,evidences→expected_output/reference set, and each metric inmetricsbecomes aGraderResult. Since your benchmark already standardizes evaluation across multiple GraphRAG frameworks (LightRAG, etc.), exporting to a common interchange format could let someone compare a GraphRAG-Bench run against a non-graph RAG baseline scored by a different tool, on the same questions.I'd propose a standalone
graphragbench-openeval-adapterpackage (to_openeval(detailed_record)/from_openeval(suite)), tested against real output fromretrieval_eval.py. Happy to open a PR if useful, or keep it as a separate package.Spec: https://github.com/adhabnr-ux/evalport/blob/main/SPEC.md
— Sahi, independent contributor (not affiliated with GraphRAG-Bench)