logoalt Hacker News

kostajtoday at 2:54 PM0 repliesview on HN

Agree. Human experts also struggle agreeing on this type of claims. The inter-annotator agreement on the verdicts on the AVeriTeC corpus across 50 organizations is κ=0.619 - substantial but well short of perfect.