Skip to content

added rolebench dataset. - #633

Merged
bittersweet1999 merged 3 commits into
open-compass:mainfrom
rolellm:rolebench
Dec 1, 2023
Merged

added rolebench dataset.#633
bittersweet1999 merged 3 commits into
open-compass:mainfrom
rolellm:rolebench

Conversation

@rolellm

@rolellm rolellm commented Nov 24, 2023

Copy link
Copy Markdown
Contributor

Thanks for your contribution and we appreciate it a lot. The following instructions would make your pull request more healthy and more easily get feedback. If you do not understand some items, don't worry, just make the pull request and seek help from maintainers.

Motivation

增加了rolebench数据集,可以实现对LLM角色扮演能力的评估。

Modification

opencompass/datasets/rolebench.py
1. 增加了InstructionGeneralizationEnglishDataset类,用于加载Instruction Generalization的英文训练集和测试集。
2. 增加了InstructionGeneralizationChineseDataset类,用于加载Instruction Generalization的中文训练集和测试集。
3. 增加了RoleGeneralizationEnglishDataset类,用于加载Role Generalization的中文训练集和测试集。

configs/datasets
对于上述三个类,分别增加了配置文件。

BC-breaking (Optional)

Does the modification introduce changes that break the backward compatibility of the downstream repositories?
If so, please describe how it breaks the compatibility and how the downstream projects should modify their code to keep compatibility with this PR.

Use cases (Optional)

If this PR introduces a new feature, it is better to list some use cases here and update the documentation.

Checklist

Before PR:

  • Pre-commit or other linting tools are used to fix the potential lint issues.
  • Bug fixes are fully covered by unit tests, the case that causes the bug should be added in the unit tests.
  • The modification is covered by complete unit tests. If not, please add more unit test to ensure the correctness.
  • The documentation has been modified accordingly, like docstring or example tutorials.

After PR:

  • If the modification has potential influence on downstream or other related projects, this PR should be tested with those projects.
  • CLA has been signed and all committers have signed the CLA in this PR.

from opencompass.openicl.icl_retriever import ZeroRetriever
from opencompass.openicl.icl_inferencer import GenInferencer
from opencompass.openicl.icl_evaluator import RougeEvaluator
from opencompass.datasets.rolebench import RoleGeneralizationEnglishDataset

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is still a conflict in the variable naming of this script

Comment thread configs/datasets/rolebench/role_generalization_eng.py Outdated
@bittersweet1999
bittersweet1999 merged commit e10f1c9 into open-compass:main Dec 1, 2023
liuyaox pushed a commit to liuyaox/opencompass that referenced this pull request Jun 26, 2024
* added rolebench

* 修改了不合理的变量名

* 修改了评论中的变量名
@audaymalte-web

audaymalte-web commented May 5, 2026

Copy link
Copy Markdown

Hi @rolellm,
I was directed to this repo from this page - https://github.com/InteractiveNLP-Team/RoleLLM-public. Thanks for your work!
We notice a lot of differences in the paper (https://arxiv.org/abs/2310.00746) vs the OpenCompass implementation. A few of them being:

  1. Differences in prompting (paper uses BM25, filtered on the same role to create few shot examples), while repo uses zero-shot prompting
  2. Multi-dimensional (RAW, CUS, SPE + win rates) vs single ROUGE score.
  3. Slightly different wording in the prompt

Can you share the same codebase you'd used for the paper. Alternatively, you can share numbers that you get with this OpenCompass implementation and we try to reproduce those numbers.
Thanks!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants