Position: Benchmarking is Broken - Don't Let AI Be Its Own Judge

Λεπτομέρειες βιβλιογραφικής εγγραφής
Τίτλος: Position: Benchmarking is Broken - Don't Let AI Be Its Own Judge
Συγγραφείς: Cheng, Zerui, Wohnig, Stella, Gupta, Ruchika, Alam, Samiul, Abdullahi, Tassallah, Ribeiro, João Alves, Nielsen-Garcia, Christian, Mir, Saif, Li, Siran, Orender, Jason, Bahrainian, Seyed Ali, Kirste, Daniel, Gokaslan, Aaron, Eickhoff, Carsten, Viswanath, Pramod, Wolff, Ruben
Πηγή: Computer Science Faculty Publications
Στοιχεία εκδότη: ODU Digital Commons
Έτος έκδοσης: 2025
Συλλογή: Old Dominion University: ODU Digital Commons
Θεματικοί όροι: Artificial intelligence-- Evaluation, Artificial intelligence-- Moral and ethical aspects, Artificial intelligence-- Reliability, Artificial intelligence-- Social aspects, Artificial intelligence-- Standards, Benchmarking (Management), Computer software--Verification, Data integrity, Information technology--Quality control, Technology-- Social aspects, Artificial Intelligence and Robotics, Information Security, Science and Technology Policy
Περιγραφή: The meteoric rise of Artificial Intelligence (AI), with its rapidly expanding market capitalization, presents both transformative opportunities and critical challenges. Chief among these is the urgent need for a new, unified paradigm for trustworthy evaluation, as current benchmarks increasingly reveal critical vulnerabilities. Issues like data contamination and selective reporting by model developers fuel hype, while inadequate data quality control can lead to biased evaluations that, even if unintentionally, may favor specific approaches. As a flood of participants enters the AI space, this "Wild West" of assessment makes distinguishing genuine progress from exaggerated claims exceptionally difficult. Such ambiguity blurs scientific signals and erodes public confidence, much as unchecked claims would destabilize financial markets reliant on credible oversight from agencies like Moody’s. Human exams like the SAT or GRE have achieved recognized standards of fairness and credibility; why settle for less in evaluating AI, especially given its profound societal impact? This position paper argues that the current laissez-faire approach is unsustainable. We contend that true, sustainable AI advancement demands a paradigm shift: a unified, live, and quality-controlled benchmarking framework robust by construction, not by mere courtesy and goodwill. To this end, we dissect the systemic flaws undermining today’s AI evaluation, distill the essential requirements for a new generation of assessments, and introduce a roadmap embodying this paradigm. Our goal is to pave the way for evaluations that can restore integrity and deliver genuinely trustworthy measures of AI progress.
Τύπος εγγράφου: report
Περιγραφή αρχείου: application/pdf
Γλώσσα: unknown
Relation: https://digitalcommons.odu.edu/computerscience_fac_pubs/379; https://digitalcommons.odu.edu/context/computerscience_fac_pubs/article/1384/viewcontent/Orender_2025_PositionBenchmarkingisBrokenDon_tLetAIbeItsOwnOCR.pdf
DOI: 10.13140/RG.2.2.33834.94408/1
Διαθεσιμότητα: https://digitalcommons.odu.edu/computerscience_fac_pubs/379
https://doi.org/10.13140/RG.2.2.33834.94408/1
https://digitalcommons.odu.edu/context/computerscience_fac_pubs/article/1384/viewcontent/Orender_2025_PositionBenchmarkingisBrokenDon_tLetAIbeItsOwnOCR.pdf
Rights: © The Authors 2025 Published under the terms of the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International (CC BY-NC-ND 4.0) License .
Αριθμός Καταχώρησης: edsbas.284C610E
Βάση Δεδομένων: BASE
Περιγραφή
DOI:10.13140/RG.2.2.33834.94408/1