General LLM Evaluation
Rubric-centered, trustworthy, and responsible evaluation for large models across modalities, time, cultures, ethics, and rewards.
通用大模型评估把 rubric-based evaluation、LLM-as-a-judge、可信评测、responsible AI 和 reward design 放在同一条线上。它的目标不是再堆一个 leaderboard,而是把开放式输出、主观质量、时效性、文化差异、伦理原则和专家标准变成可审计、可复现、可训练的评测信号。
Research Storyline
MLLM-Bench 用 per-sample criteria 组织多模态评测,强调 judge 需要知道每个样本到底该看什么。
Humans or LLMs as the Judge? 暴露人类和模型评委的偏差,让自动评测从“方便”走向“可质控”。
FreshBench、Culture Bias 和 PrinciplismQA 分别处理过时知识、跨文化偏差和临床医学伦理,让评测覆盖模型真正会伤人的边界。
Awesome-Rubrics 将 rubrics 组织成评测、alignment、reward modeling 和 post-training 的统一接口。
Representative Work
A curated reading list and survey around rubric-based evaluation, reward modeling, alignment, and agentic AI.
RepositoryUses per-sample criteria to make multimodal evaluation more explicit and judgeable.
PaperStudies judgement bias in human and LLM evaluators, a core issue for automated evaluation.
PaperTests temporal generalization and whether models remain reliable as the world changes.
RepositoryEvaluates and mitigates culture bias through an LLM-adaptive word association test.
PaperAssesses clinical medical ethics alignment with expert-validated MCQA, open-ended cases, and rubric keypoints.
RepositoryWhat This Direction Adds
Rubrics, judges, criteria, calibration, and reliability checks become reusable infrastructure for many domains.
Structured feedback can supervise SFT, preference tuning, reward modeling, RL, and self-improvement loops.
Freshness, culture, ethics, safety, and expert standards expose failures that ordinary accuracy scores miss.