Human Review of AI Not Always Trustworthy, Lacks Clear Standards

Hindustan Times · technology

The term 'human-reviewed' for AI systems does not guarantee trustworthiness, as reviewers may assess different criteria. For instance, a chatbot might be marked correct for accurate eligibility rules, but incorrect for omitting crucial application deadlines. Defining a 'human gold standard' requires clear objectives for what the AI should accomplish, not just agreement among reviewers. A recent study found significant gaps in how academic research documents human evaluation processes, suggesting commercial claims may be even less transparent. Selecting independent reviewers and establishing clear assessment rubrics are critical for credible AI evaluation.

AI की इंसानी समीक्षा हमेशा भरोसेमंद नहीं, तय मानक भी नहीं

AI सिस्टम के लिए 'मानव-समीक्षित' शब्द भरोसे की गारंटी नहीं देता, क्योंकि समीक्षक अलग-अलग मापदंडों पर जांच कर सकते हैं। उदाहरण के लिए, एक चैटबॉट को सटीक पात्रता नियमों के लिए सही चिह्नित किया जा सकता है, लेकिन महत्वपूर्ण आवेदन समय-सीमाओं को छोड़ देने के लिए गलत। 'मानव गोल्ड स्टैंडर्ड' को परिभाषित करने के लिए AI को क्या हासिल करना चाहिए, इसके स्पष्ट उद्देश्य की आवश्यकता है, न कि केवल समीक्षकों के बीच सहमति की। एक हालिया अध्ययन में पाया गया कि अकादमिक शोध मानव मूल्यांकन प्रक्रियाओं को कैसे दस्तावेज करते हैं, इसमें महत्वपूर्ण खामियां हैं, जिससे पता चलता है कि व्यावसायिक दावे और भी कम पारदर्शी हो सकते हैं।