The docs treat larger / appeal committees as increasing accuracy — which holds when validator errors are independent. In a measurement on a public benchmark (RAGTruth), a 6-model committee across 4 labs and 3 countries acts like ~1.9 independent judges of 6; the errors correlate at ρ ≈ 0.4–0.6, so added validators help less than independence predicts.
Is validator-error independence measured, or assumed, when committee size / appeal doubling is chosen?
Full method, data, and a one-command recompute: https://github.com/itsjustmarsel/llm-judge-independence
The docs treat larger / appeal committees as increasing accuracy — which holds when validator errors are independent. In a measurement on a public benchmark (RAGTruth), a 6-model committee across 4 labs and 3 countries acts like ~1.9 independent judges of 6; the errors correlate at ρ ≈ 0.4–0.6, so added validators help less than independence predicts.
Is validator-error independence measured, or assumed, when committee size / appeal doubling is chosen?
Full method, data, and a one-command recompute: https://github.com/itsjustmarsel/llm-judge-independence