Skip to content

EvalPort: portable interchange for the per-sample eval record in Evaluation/retrieval_eval.py #60

Description

@adhabnr-ux

Hi GraphRAG-Bench team — I maintain EvalPort, an open JSON-Schema spec for portable LLM eval data (TestCase/Grader/Result/ResultSet/GraderResult). No CONTRIBUTING.md, so filing as an issue.

I read Evaluation/retrieval_eval.py directly. The detailed per-sample record built in evaluate_dataset() — {"id", "question", "contexts", "evidences", "metrics"}, with metrics from evaluate_sample()'s {"context_relevancy", "evidence_recall"} — maps closely onto EvalPort: question→TestCase.input, evidences→expected_output/reference set, and each metric in metrics becomes a GraderResult. Since your benchmark already standardizes evaluation across multiple GraphRAG frameworks (LightRAG, etc.), exporting to a common interchange format could let someone compare a GraphRAG-Bench run against a non-graph RAG baseline scored by a different tool, on the same questions.

I'd propose a standalone graphragbench-openeval-adapter package (to_openeval(detailed_record) / from_openeval(suite)), tested against real output from retrieval_eval.py. Happy to open a PR if useful, or keep it as a separate package.

Spec: https://github.com/adhabnr-ux/evalport/blob/main/SPEC.md

— Sahi, independent contributor (not affiliated with GraphRAG-Bench)

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions