AI Infra
0%
Part VII · Chapter 49

Human Evaluation as a Measurement Protocol

AuthorChangkun Ou
Reading time~12 min

A human judgment is not ground truth merely because a person produced it. It is an observation made under a protocol: a particular person saw particular evidence through a particular interface, applied a particular rubric, and selected from particular response options. Change any of those conditions and the label may change.

The useful question is therefore not “What did humans prefer?” It is “Which people made which judgment, about what unit, under which conditions, and what claim can those observations support?” This chapter makes that protocol explicit. The statistical contract in Chapter 48 still applies; human evaluation adds rater sampling, interpretation, fatigue, incentives, and adjudication to the measurement system.

Begin with the Claim, Not the Crowd

Before recruiting raters, write the decision and the construct to be measured. “Quality” is not a construct precise enough to guide a release. Factual correctness, citation support, harmlessness, tone fit, task completion, and user effort can move independently. A single overall preference silently chooses how they trade off.

The target rater population follows from the claim:

Claim Relevant perspective Evidence the rater needs
A medical summary preserves clinically important facts clinician with relevant expertise source record, summary, error taxonomy
An answer helps a customer complete a workflow intended user or a valid proxy task goal, product state, outcome
A response follows a safety policy trained policy rater; specialist for high-risk domains policy revision, full interaction, allowed tools
A localization sounds natural fluent member of the target language community surrounding context and locale
One answer is more pleasant to read sampled target users blinded candidate outputs

Expertise is not a universal ranking of people. Domain credentials matter for clinical correctness; lived experience or community membership may matter for cultural meaning; actual users matter for workflow utility. The study must name the perspective it samples and avoid generalizing beyond it.

C decision and construct S sampled cases and systems C->S P rubric, raters, interface S->P A randomized assignment P->A R raw judgment records A->R Q audit; adjudicate separately R->Q O estimate, limits, decision Q->O
Figure 49.1. A human evaluation preserves the path from decision to raw judgments. Reliability analysis and adjudication both consume the raw records; adjudication does not overwrite them.

HealthBench illustrates one domain-specific design: physicians wrote conversation-specific criteria for health conversations, and the automated grader was separately compared with physician judgments (Arora et al. 2025). It is evidence for that benchmark protocol, not proof that physician-written rubrics or model graders are valid for every health use.

A Rubric Is an Executable Interface

A rubric should turn an abstract construct into observable decisions. Each criterion needs five parts:

  • Unit: claim, citation, answer, turn, conversation, tool call, or completed task.
  • Condition: what must be present for the criterion to apply.
  • Evidence: what the rater may inspect and which sources take priority.
  • Response set: pass/fail, ordered category, score, pairwise choice, tie, unknown, or abstain.
  • Anchor: examples near important boundaries, with explanations.

“Be helpful” asks each rater to invent a private policy. “Identify the user's requested action, complete it without an unsupported claim, and state any unresolved prerequisite” is more observable. Criteria should separate properties that must not trade silently. A fluent false answer should not earn enough style points to erase a factual failure.

Rubric development is empirical. Pilot it on ordinary, boundary, adversarial, and unscorable cases. Ask raters to explain how they interpreted each criterion; recurrent confusion identifies wording defects. Revise the instructions and anchors, then freeze a rubric revision before the confirmatory run. Pilot labels used to tune the rubric are development data, not fresh confirmation data.

The response set must preserve uncertainty. A forced binary choice converts “both fail,” “indistinguishable,” and “insufficient evidence” into arbitrary wins. Tie, unknown, not-applicable, and abstain are different states and should remain different in storage and analysis.

Assignment and Interface Shape the Observation

Even a good rubric can be undermined by its presentation. The protocol should control variables unrelated to the construct:

  • Blind model and provider identity unless identity is itself part of the task.
  • Randomize or counterbalance candidate order in pairwise comparisons.
  • Randomize case order where possible, and record the order actually shown.
  • Keep typography, truncation, citations, and available tools symmetric across systems.
  • Separate training and calibration items from the scored sample.
  • Predeclare exclusion, timeout, and missing-response rules.

Within-rater designs expose the same person to multiple systems, preserving a useful pair but introducing learning, fatigue, and carryover. Between-rater designs reduce carryover but need more people and can confound system differences with rater differences. Balanced incomplete blocks are often a practical compromise: each rater sees a manageable subset, every system receives comparable exposure, and enough overlap remains to estimate rater and item effects.

The interface is part of the instrument. van der Lee et al. recommend reporting and controlling question wording, presentation order, rater training, quality checks, and statistical analysis because these choices affect reproducibility (Lee et al. 2019). GENIE later demonstrated that interface and annotator-quality design can materially affect reproducibility across tasks and rater populations (Khashabi et al. 2022). Store the interface revision with every judgment.

Recruit the Perspective and Protect the Person

Rater qualification should test what the task requires. A generic platform approval rate is not evidence of medical expertise, policy fluency, or cultural knowledge. Use a documented screener, training set, and calibration round; record failures without silently changing the qualification threshold after seeing system results.

Crowd studies can be appropriate for low-stakes language preference, but “crowd” is not a stable population. Platform, geography, language, compensation, device, and worker incentives all shape who participates and how. In open-ended story evaluation, Karpinska et al. found that strict platform filters did not make crowd judgments interchangeable with English-teacher judgments; paired references changed the task and improved the signal (Karpinska et al. 2021). Clark et al. likewise found that untrained non-experts distinguished GPT-3 from human-authored text only at chance level across their tested domains (Clark et al. 2021). These results limit those protocols; they do not imply that non-experts are generally incapable of useful evaluation.

People are not disposable measurement components. Budget for fair compensation, realistic task time, breaks, clear rejection and appeal rules, and protection from disturbing content. Collect only demographics needed for a stated analysis, minimize retention of platform identifiers, and obtain the appropriate ethics or legal review rather than assuming paid annotation is exempt. Shmueli et al. document harms that fair pay alone does not cover, including privacy, exposure, and the ambiguity of human-subject protections for crowd work (Shmueli et al. 2021).

Quality control should be related to the task. Duplicate items can measure within-rater stability; objective gold items can detect misunderstood instructions; targeted review can inspect implausibly fast or uniform patterns. Minimum-time filters and hidden traps are weak substitutes for a valid task and can punish skilled workers. Predeclare the checks, preserve excluded records and reasons, and report how exclusions changed the estimate.

Preserve Raw Judgments Before Adjudication

Every judgment should remain an immutable event:

JudgmentRecord:
  study_spec_hash
  case_id
  candidate_ids_blinded
  candidate_order
  rubric_revision
  interface_revision
  rater_pseudonym
  rater_population_and_qualification
  assignment_block
  raw_dimension_labels
  tie_unknown_abstain_state
  rationale_and_evidence_refs
  started_at
  submitted_at
  quality_flags

Adjudication creates a derived label; it must not overwrite the independent observations. Preserve the adjudicator, evidence, rationale, and rule used. Report agreement before adjudication. A consensus meeting can resolve a production label, but it does not retroactively turn several dependent opinions into independent evidence.

This distinction also prevents leakage. Labels collected to evaluate a release may later train a reward model, tune a rubric, or calibrate a model judge. Once reused for development, that set is no longer an untouched release check. Provenance must say which judgments influenced which artifacts.

Agreement Diagnoses the Protocol

Raw agreement is the fraction of jointly labeled units on which raters match. It is useful, but it ignores the label distribution. Cohen's kappa for two raters and nominal labels is (Cohen 1960)

κ=pope1pe.\kappa = \frac{p_o-p_e}{1-p_e}.

Here pop_o is observed agreement and pep_e is agreement expected if the two raters label independently while keeping their observed marginal label frequencies. The adjustment is model-based; pep_e is not a universal amount of agreement “caused by chance.” If both raters assign the same label to every unit, then po=pe=1p_o=p_e=1 and kappa is undefined, even though raw agreement is 100%. With very imbalanced labels, high raw agreement can coexist with a low or unstable kappa.

For multiple raters, missing assignments, and non-nominal scales, choose a coefficient that matches the design. Krippendorff's alpha has the general form (Hayes and Krippendorff 2007; Artstein and Poesio 2008)

α=1DoDe.\alpha = 1-\frac{D_o}{D_e}.

Here DoD_o denotes observed disagreement and DeD_e denotes disagreement expected under the coefficient's reference model. The distance function must match the response scale: nominal categories have no ordering, ordinal levels have an order but no fixed spacing, and interval scores assert meaningful distances.

No coefficient has a universal “safe” cutoff. Report its uncertainty, raw agreement, label prevalence, confusion patterns, number of raters and units, missingness, and the coefficient variant. Compute it on independent raw judgments under one rubric revision. Pooling post-adjudication labels or multiple rubric versions produces an easy number with no coherent interpretation.

Agreement is reliability, not validity. Consistent raters can apply a wrong rubric; valid criteria can also expose genuine differences in perspective. Examine disagreement by criterion, item, and rater population. A recurring split may call for clearer wording, a separate dimension, more evidence, or an explicit perspective-conditioned result rather than a forced consensus.

Choose the Observation Format for the Claim

Pairwise comparison asks for A, B, tie, or unscorable. It reduces scale-calibration burden and supports a paired analysis, but the result is relative to the opponent set and presentation context. It does not produce an absolute quality level, and cyclic or heterogeneous preferences can be real. Always reverse or randomize order and retain both raw positions.

Absolute or ordinal rating scores one output against anchored levels. It supports thresholds and reusable item-level diagnoses, but raters differ in severity and in how they use the scale. Anchors, calibration items, and models that account for rater and item effects may be necessary. Do not treat ordinal levels as equally spaced without justifying that assumption.

Criterion checklist decomposes an answer into observable requirements. It localizes failures and transfers well to later judge calibration, but a long checklist increases fatigue and can miss interactions among criteria.

Task-based evaluation measures what happens when a person uses the system: completion, time, corrections, escalation, or downstream error. It is closest to a product claim, but the estimate covers the entire human-system workflow, not model text alone. User skill, interface, and assistance policy belong in the system boundary.

The choice is not a hierarchy. Use the cheapest format that directly observes the decision-relevant construct, then add a stronger downstream study when the product claim exceeds what the first format can support.

Aggregate Without Erasing the Design

The independent unit may be a user, conversation, document, repository, or task, not an individual checkbox. Preserve paired assignments and cluster repeated judgments at that unit, as Chapter 48 requires. Report system effects with uncertainty and guardrails, not only mean scores.

Aggregation should retain important structure:

  • per-criterion failure rates rather than only a total;
  • rater-population and language slices named in advance;
  • tie, abstention, missing, and invalid rates;
  • order and interface diagnostics;
  • rater and item effects when the assignment supports estimating them;
  • sensitivity to exclusion and adjudication rules.

When perspectives legitimately differ, the distribution is often the result. Majority vote can be useful for one operational label, but it erases minority judgments and uncertainty. Keep both the aggregation rule and the raw evidence available.

Hand Human Evidence to Model Judges Carefully

Human labels can calibrate an automated judge, but the handoff is another evaluation, not a promotion to ground truth. Use a locked human-labeled set that represents ordinary cases, boundary cases, adversarial cases, and abstentions. Measure the judge per criterion and slice; test order, verbosity, self-preference, and distribution shift; and retain a human-review path for unsupported cases.

Do not validate only against adjudicated consensus. Raw judgments reveal whether the model tracks a stable human signal or merely one adjudicator's resolution policy. Monitor judge-human disagreement on fresh samples, especially after changing the model under test, rubric, judge, prompt, or domain. Chapter 50 develops that automated-judge contract.

The operating sequence is:

  1. State the decision, construct, unit, target case population, and rater perspective.
  2. Write observable criteria, evidence rules, anchors, and uncertainty states.
  3. Pilot the rubric and interface; revise, then freeze their versions.
  4. Predeclare sampling, assignment, blinding, order, exclusions, adjudication, aggregation, and analysis.
  5. Recruit, train, compensate, and protect the stated rater population.
  6. Collect immutable raw judgments before any adjudication.
  7. Report effects, uncertainty, agreement diagnostics, exclusions, and validity limits.
  8. Treat reuse for training or judge calibration as a provenance-changing event.
What's contested

Disagreement is sometimes measurement error and sometimes the phenomenon being measured. A policy-compliance task may need one operational resolution; a value-laden preference may need several perspective-conditioned distributions. The protocol should not decide this after seeing inconvenient disagreement. It should state whether the target is an expert standard, user utility, policy interpretation, community perspective, or population preference before labels are collected.

Lower-layer constraint

The annotation service must retain stable case, candidate, rater, assignment, rubric, interface, and order identifiers. Without them, the evaluation layer cannot reconstruct pairing, detect leakage, estimate rater effects, or audit exclusions. The adaptation layer must also record when evaluation judgments become training data; otherwise a later “held-out” human evaluation may already have shaped the model or judge it evaluates.

Further reading

  • Lee et al., “Best Practices for the Human Evaluation of Automatically Generated Text,” 2019. aclanthology.org
    The paper connects human-evaluation validity and reproducibility to explicit planning, participant selection, question wording, presentation order, training, quality control, statistical analysis, and reporting.
  • Artstein & Poesio, “Inter-Coder Agreement for Computational Linguistics,” 2008. aclanthology.org
    Artstein and Poesio survey agreement coefficients, their assumptions, and their interpretation for computational-linguistics annotation, emphasizing that the coefficient must match the design and scale.
  • Shmueli et al., “Beyond Fair Pay: Ethical Implications of NLP Crowdsourcing,” 2021. aclanthology.org
    The authors show that crowdwork ethics extends beyond hourly pay to privacy, identifiable data, harmful exposure, consent, and the limits of existing human-subject review frameworks.
  • Cohen, “A Coefficient of Agreement for Nominal Scales” (agreement beyond chance), 1960. doi.org
    Cohen's kappa measures agreement between two raters after adjusting for agreement expected by chance, making raw percent agreement harder to misuse.
  • Hayes & Krippendorff, “Answering the Call for a Standard Reliability Measure for Coding Data” (Krippendorff's alpha for many coders and missing data), 2007. doi.org
    Hayes and Krippendorff argue for Krippendorff's alpha as a general reliability coefficient that handles many coders, missing data, and different measurement levels.
  • Khashabi et al., “GENIE: Toward Reproducible and Standardized Human Evaluation for Text Generation” (standardized human evaluation interfaces), 2022. aclanthology.org
    GENIE studies human-evaluation design choices for text generation and introduces a standardized platform that improves reproducibility across tasks and annotator populations.
  • Karpinska et al., “The Perils of Using Mechanical Turk to Evaluate Open-Ended Text Generation” (crowd judgments are not automatically gold), 2021. aclanthology.org
    Karpinska et al. show that crowd workers can fail to distinguish human and model-written open-ended stories, while expert teachers and paired examples provide stronger evaluation signals.
  • Clark et al., “All That's 'Human' Is Not Gold: Evaluating Human Evaluation of Generated Text” (untrained evaluators struggle as fluency improves), 2021. aclanthology.org
    Clark et al. show that untrained evaluators often identify GPT-3 generated text only at chance level, motivating evaluator training and clearer protocols.
  • Arora et al., “HealthBench: Evaluating Large Language Models Towards Improved Human Health” (HealthBench), 2025. arXiv:2505.08775
    HealthBench grades health conversations against 48,562 rubric criteria written by 262 physicians with practice experience across 60 countries, and meta-evaluates the model grader against physician grading before trusting its scores.

Comments

Log in to comment