General LLM Evaluation

Rubric-centered, trustworthy, and responsible evaluation for large models across modalities, time, cultures, ethics, and rewards.

Rubrics, judges, freshness, bias, ethics, and rewards
General LLM Evaluation Rubric-based evaluation LLM-as-a-judge Responsible AI Reward design
General LLM evaluation resources

通用大模型评估把 rubric-based evaluation、LLM-as-a-judge、可信评测、responsible AI 和 reward design 放在同一条线上。它的目标不是再堆一个 leaderboard,而是把开放式输出、主观质量、时效性、文化差异、伦理原则和专家标准变成可审计、可复现、可训练的评测信号。

Research Storyline

Criteria
从答案对错走向样本级标准

MLLM-Bench 用 per-sample criteria 组织多模态评测,强调 judge 需要知道每个样本到底该看什么。

Judge
校准 LLM-as-a-judge

Humans or LLMs as the Judge? 暴露人类和模型评委的偏差,让自动评测从“方便”走向“可质控”。

Trust
评估真实部署风险

FreshBench、Culture Bias 和 PrinciplismQA 分别处理过时知识、跨文化偏差和临床医学伦理,让评测覆盖模型真正会伤人的边界。

Reward
把评测信号变成训练信号

Awesome-Rubrics 将 rubrics 组织成评测、alignment、reward modeling 和 post-training 的统一接口。

Representative Work

Rubric
Awesome-Rubrics

A curated reading list and survey around rubric-based evaluation, reward modeling, alignment, and agentic AI.

Repository
MLLM
MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria

Uses per-sample criteria to make multimodal evaluation more explicit and judgeable.

Paper
Judge
Humans or LLMs as the Judge? A Study on Judgement Biases

Studies judgement bias in human and LLM evaluators, a core issue for automated evaluation.

Paper
Fresh
FreshBench: Is Your LLM Outdated?

Tests temporal generalization and whether models remain reliable as the world changes.

Repository
Culture
From Word to World

Evaluates and mitigates culture bias through an LLM-adaptive word association test.

Paper
Ethics
PrinciplismQA

Assesses clinical medical ethics alignment with expert-validated MCQA, open-ended cases, and rubric keypoints.

Repository

What This Direction Adds

Evaluation as infrastructure

Rubrics, judges, criteria, calibration, and reliability checks become reusable infrastructure for many domains.

Evaluation as alignment signal

Structured feedback can supervise SFT, preference tuning, reward modeling, RL, and self-improvement loops.

Evaluation as risk lens

Freshness, culture, ethics, safety, and expert standards expose failures that ordinary accuracy scores miss.