Skip to main content

Who Grades the AI Models? | Ben Horowitz & Rayan Krishnan

Technology
Global
Started September 10, 2026

a16z’s Erik Torenberg, Ben Horowitz, and Jennifer Li sit down with Vals founder and CEO Rayan Krishnan to discuss one of AI’s increasingly difficult problems: how do you actually measure whether a model is getting better? As public benchmarks saturate and models get better at optimizing for the tests themselves, Rayan makes the case for independent, continuously evolving evaluations. They unpack why self-reported model scores can be misleading, how VALS evaluates models in the hours before a ...

Need to find a specific claim? Search all statements.
🗳️ Join the conversation
10 statements to vote on • Your perspective shapes the analysis
📊 Progress to Consensus Analysis Need: 7+ participants, 20+ votes, 3+ votes per statement
Participants 0/7
Statements (10+ recommended) 10/10
Total Votes 0/20
💡 Progress updates live here. Final readiness is confirmed when all three requirements are met.

Your votes count

No account needed — your votes are saved and included in the consensus analysis. Create an account to track your voting history and add statements.

CLAIM Posted by admin Sep 10, 2026
Whoever grades AI models must maintain transparency about their evaluation methods and funding sources to avoid hidden conflicts of interest.
Vote options for this statement: agree, disagree, or unsure
Vote to see results
CLAIM Posted by admin Sep 10, 2026
Requiring independent evaluations of AI models will slow down innovation and create unnecessary regulatory burden on developers.
Vote options for this statement: agree, disagree, or unsure
Vote to see results
CLAIM Posted by admin Sep 10, 2026
AI model evaluations should prioritize real-world performance over controlled test conditions.
Vote options for this statement: agree, disagree, or unsure
Vote to see results
CLAIM Posted by admin Sep 10, 2026
AI model evaluations should evolve continuously to keep pace with technological advancements.
Vote options for this statement: agree, disagree, or unsure
Vote to see results
CLAIM Posted by admin Sep 10, 2026
Measuring AI model improvements should incorporate both quantitative and qualitative metrics for a holistic view of performance.
Vote options for this statement: agree, disagree, or unsure
Vote to see results
CLAIM Posted by admin Sep 10, 2026
Evaluation systems must measure real-world AI performance on tasks that matter to users, not abstract benchmark scores.
Vote options for this statement: agree, disagree, or unsure
Vote to see results
CLAIM Posted by admin Sep 10, 2026
The evaluation process for AI models must include diverse perspectives to avoid bias and promote fairness.
Vote options for this statement: agree, disagree, or unsure
Vote to see results
CLAIM Posted by admin Sep 10, 2026
AI model evaluations should continuously evolve rather than rely on static standardized tests that models can overfit to.
Vote options for this statement: agree, disagree, or unsure
Vote to see results
CLAIM Posted by admin Sep 10, 2026
AI developers should have a vested interest in the evaluation process, as their input is critical for accurate assessments.
Vote options for this statement: agree, disagree, or unsure
Vote to see results
CLAIM Posted by admin Sep 10, 2026
A single standardized evaluation framework for all AI models is impossible because different models serve different purposes.
Vote options for this statement: agree, disagree, or unsure
Vote to see results

💡 How This Works

  • Add Statements: Post claims or questions (10-500 characters)
  • Vote: Agree, Disagree, or Unsure on each statement
  • Respond: Add detailed pro/con responses with evidence
  • Consensus: After enough participation, analysis reveals opinion groups and areas of agreement

Society Speaks is open and independent. Your support keeps civic discussion free from advertising and commercial influence.

Support us