Hi @cmots — thanks for STEB, it's a genuinely well-scoped benchmark (BLEU/COMET/XCOMET fidelity + emotion/style/NV expressiveness via LALM-as-judge is a real gap most S2ST evals skip).
I maintain EvalPort (Apache 2.0), an open interchange format for LLM/speech eval test cases, graders, and results — TestCase / Grader / ResultSet JSON schemas plus a Python/TS SDK (evalport-sdk), with 37 framework adapters merged so far (DeepEval, Ragas, LangSmith, Hugging Face evaluate, EleutherAI lm-eval-harness, etc.). Not affiliated with STEB — opening this because your evaluation/eval/scorers/base.py maps onto it unusually cleanly, and a STEB → EvalPort adapter would let STEB's expressiveness scores get compared/aggregated alongside other eval frameworks without STEB itself changing anything.
Why it fits
I read EvalRecord and BaseScorer directly (not guessing):
EvalRecord is already close to EvalPort's TestCase: id, ref_text/ref_translation[tgt_lang] → input/expected_output; src_lang/tgt_lang, model_name, ref_wav_path/hyp_wav_path have no direct TestCase field but fit cleanly under metadata (same pattern the DeepEval adapter uses for fields EvalPort's schema doesn't carry — see adapters/deepeval-openeval-adapter).
BaseScorer.score(records) -> List[Dict] already returns one dict per record keyed by id with metric fields (e.g. BLEUScorer emits {"id", "bleu", "bleu_asr"}) — that's positionally/id-aligned exactly like EvalPort's ResultSet.results[].grader_results, just needs each metric field turned into one GraderResult (grader_id, score normalized to [0,1], passed, metadata for anything scorer-specific like BLEU's raw 0–100 or the LLM judges' rubric breakdown).
- STEB's LLM-as-judge scorers (
llm_emotion_scorer.py, llm_style_scorer.py, llm_event_scorer.py) are the same shape EvalPort's llm_judge grader type and the Ragas/DeepEval adapters already handle — a caption-then-summarize rubric judge is not a new pattern for the adapter side, just a new prompt.
Sketch (using the real field names above)
# steb_openeval_adapter/to_openeval.py
from openeval.types import TestCase
from evaluation.eval.scorers.base import EvalRecord
def record_to_testcase(rec: EvalRecord) -> TestCase:
return TestCase(
id=rec.id,
input=rec.ref_text_with_events or rec.ref_text,
expected_output=rec.ref_translation.get(rec.tgt_lang, ""),
context=[c for c in (rec.ref_emotion, rec.ref_style, rec.ref_caption) if c],
graders=["bleu", "duration", "llm_emotion", "llm_style"],
metadata={
"src_lang": rec.src_lang,
"tgt_lang": rec.tgt_lang,
"model_name": rec.model_name,
"ref_wav_path": rec.ref_wav_path,
"hyp_wav_path": rec.hyp_wav_path,
},
)
# steb_openeval_adapter/to_resultset.py — from BaseScorer.score() output
def scorer_dicts_to_grader_results(scorer_name: str, dicts: list[dict], *, scale: float = 1.0) -> dict[str, dict]:
"""dicts: BaseScorer.score()'s List[Dict[str, Any]], one per record, keyed by 'id'."""
out = {}
for d in dicts:
for field, value in d.items():
if field == "id" or value is None:
continue
out.setdefault(d["id"], []).append({
"grader_id": f"{scorer_name}.{field}",
"type": "custom",
"score": max(0.0, min(1.0, value / scale)), # e.g. scale=100 for BLEU
"passed": value / scale >= 0.5,
"metadata": {"raw_value": value},
})
return out
Happy to build this as a standalone steb-openeval-adapter package (zero footprint on STEB itself, same shape as the other 37) if that's useful — or if you'd rather it live in-repo, I can open a PR instead. Either way, no obligation; this is just a heads-up that the mapping is unusually direct given EvalRecord's existing shape, in case it's useful to you or someone consuming STEB's scores elsewhere.
— Sahi, independent contributor (not affiliated with STEB)
Hi @cmots — thanks for STEB, it's a genuinely well-scoped benchmark (BLEU/COMET/XCOMET fidelity + emotion/style/NV expressiveness via LALM-as-judge is a real gap most S2ST evals skip).
I maintain EvalPort (Apache 2.0), an open interchange format for LLM/speech eval test cases, graders, and results —
TestCase/Grader/ResultSetJSON schemas plus a Python/TS SDK (evalport-sdk), with 37 framework adapters merged so far (DeepEval, Ragas, LangSmith, Hugging Faceevaluate, EleutherAIlm-eval-harness, etc.). Not affiliated with STEB — opening this because yourevaluation/eval/scorers/base.pymaps onto it unusually cleanly, and a STEB → EvalPort adapter would let STEB's expressiveness scores get compared/aggregated alongside other eval frameworks without STEB itself changing anything.Why it fits
I read
EvalRecordandBaseScorerdirectly (not guessing):EvalRecordis already close to EvalPort'sTestCase:id,ref_text/ref_translation[tgt_lang]→input/expected_output;src_lang/tgt_lang,model_name,ref_wav_path/hyp_wav_pathhave no directTestCasefield but fit cleanly undermetadata(same pattern the DeepEval adapter uses for fields EvalPort's schema doesn't carry — seeadapters/deepeval-openeval-adapter).BaseScorer.score(records) -> List[Dict]already returns one dict per record keyed byidwith metric fields (e.g.BLEUScoreremits{"id", "bleu", "bleu_asr"}) — that's positionally/id-aligned exactly like EvalPort'sResultSet.results[].grader_results, just needs each metric field turned into oneGraderResult(grader_id,scorenormalized to[0,1],passed,metadatafor anything scorer-specific like BLEU's raw 0–100 or the LLM judges' rubric breakdown).llm_emotion_scorer.py,llm_style_scorer.py,llm_event_scorer.py) are the same shape EvalPort'sllm_judgegrader type and the Ragas/DeepEval adapters already handle — a caption-then-summarize rubric judge is not a new pattern for the adapter side, just a new prompt.Sketch (using the real field names above)
Happy to build this as a standalone
steb-openeval-adapterpackage (zero footprint on STEB itself, same shape as the other 37) if that's useful — or if you'd rather it live in-repo, I can open a PR instead. Either way, no obligation; this is just a heads-up that the mapping is unusually direct givenEvalRecord's existing shape, in case it's useful to you or someone consuming STEB's scores elsewhere.— Sahi, independent contributor (not affiliated with STEB)