Human Grading — Abbreviation Study

Progress & export

Admin view. Grading progress is always visible; the merged export unlocks only after both graders have marked their grading complete.

Progress

GraderGradedStartedLast activityLocked
Grader A17 / 360
7/24/2026, 5:17:34 PM7/24/2026, 6:27:55 PMopen
Grader B18 / 360
7/24/2026, 6:04:39 PM7/24/2026, 6:27:55 PMopen

Human-graded statistics

Statistics unlock once both graders finish. The per-variant breakdown needs the confidential key, which stays sealed until grading is complete. Still grading: Grader A, Grader B.

Merged export

Export locked. Still grading: Grader A, Grader B. Model identity and the other grader's scores stay hidden until both finalize, so neither grader's judgments can be influenced by the experimental condition or by the other grader.

What the export contains

One row per response — 360 rows carrying all 720 grading decisions, merged on evaluation_id with both graders' original scores side by side. Identity columns: run_id, model_key, parameter_count_b, variant, domain, set_id, prompt_id, abbreviation. Then the three V3 score fields twice, prefixed grader_a_ and grader_b_ (accuracy_score, clarification_score, hallucination_score), plus notes and graded_at. Clarification is only populated for real_low_context items and hallucination only for synthetic_high_context.

Every response is double-graded, so the side-by-side columns support the agreement analysis the V3 spec asks for: raw agreement and Cohen's kappa on Accuracy, Clarification, and Hallucination. Neither original score set is modified — record adjudicated scores separately.