Skip to content

Proposal: EvalPort adapter for STEB's EvalRecord/scorer results (interop with other eval frameworks) #3

Description

@adhabnr-ux

Hi @cmots — thanks for STEB, it's a genuinely well-scoped benchmark (BLEU/COMET/XCOMET fidelity + emotion/style/NV expressiveness via LALM-as-judge is a real gap most S2ST evals skip).

I maintain EvalPort (Apache 2.0), an open interchange format for LLM/speech eval test cases, graders, and results — TestCase / Grader / ResultSet JSON schemas plus a Python/TS SDK (evalport-sdk), with 37 framework adapters merged so far (DeepEval, Ragas, LangSmith, Hugging Face evaluate, EleutherAI lm-eval-harness, etc.). Not affiliated with STEB — opening this because your evaluation/eval/scorers/base.py maps onto it unusually cleanly, and a STEB → EvalPort adapter would let STEB's expressiveness scores get compared/aggregated alongside other eval frameworks without STEB itself changing anything.

Why it fits

I read EvalRecord and BaseScorer directly (not guessing):

  • EvalRecord is already close to EvalPort's TestCase: id, ref_text/ref_translation[tgt_lang] → input/expected_output; src_lang/tgt_lang, model_name, ref_wav_path/hyp_wav_path have no direct TestCase field but fit cleanly under metadata (same pattern the DeepEval adapter uses for fields EvalPort's schema doesn't carry — see adapters/deepeval-openeval-adapter).
  • BaseScorer.score(records) -> List[Dict] already returns one dict per record keyed by id with metric fields (e.g. BLEUScorer emits {"id", "bleu", "bleu_asr"}) — that's positionally/id-aligned exactly like EvalPort's ResultSet.results[].grader_results, just needs each metric field turned into one GraderResult (grader_id, score normalized to [0,1], passed, metadata for anything scorer-specific like BLEU's raw 0–100 or the LLM judges' rubric breakdown).
  • STEB's LLM-as-judge scorers (llm_emotion_scorer.py, llm_style_scorer.py, llm_event_scorer.py) are the same shape EvalPort's llm_judge grader type and the Ragas/DeepEval adapters already handle — a caption-then-summarize rubric judge is not a new pattern for the adapter side, just a new prompt.

Sketch (using the real field names above)

# steb_openeval_adapter/to_openeval.py
from openeval.types import TestCase
from evaluation.eval.scorers.base import EvalRecord

def record_to_testcase(rec: EvalRecord) -> TestCase:
    return TestCase(
        id=rec.id,
        input=rec.ref_text_with_events or rec.ref_text,
        expected_output=rec.ref_translation.get(rec.tgt_lang, ""),
        context=[c for c in (rec.ref_emotion, rec.ref_style, rec.ref_caption) if c],
        graders=["bleu", "duration", "llm_emotion", "llm_style"],
        metadata={
            "src_lang": rec.src_lang,
            "tgt_lang": rec.tgt_lang,
            "model_name": rec.model_name,
            "ref_wav_path": rec.ref_wav_path,
            "hyp_wav_path": rec.hyp_wav_path,
        },
    )

# steb_openeval_adapter/to_resultset.py — from BaseScorer.score() output
def scorer_dicts_to_grader_results(scorer_name: str, dicts: list[dict], *, scale: float = 1.0) -> dict[str, dict]:
    """dicts: BaseScorer.score()'s List[Dict[str, Any]], one per record, keyed by 'id'."""
    out = {}
    for d in dicts:
        for field, value in d.items():
            if field == "id" or value is None:
                continue
            out.setdefault(d["id"], []).append({
                "grader_id": f"{scorer_name}.{field}",
                "type": "custom",
                "score": max(0.0, min(1.0, value / scale)),  # e.g. scale=100 for BLEU
                "passed": value / scale >= 0.5,
                "metadata": {"raw_value": value},
            })
    return out

Happy to build this as a standalone steb-openeval-adapter package (zero footprint on STEB itself, same shape as the other 37) if that's useful — or if you'd rather it live in-repo, I can open a PR instead. Either way, no obligation; this is just a heads-up that the mapping is unusually direct given EvalRecord's existing shape, in case it's useful to you or someone consuming STEB's scores elsewhere.

— Sahi, independent contributor (not affiliated with STEB)

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions