Firehose

Filtered to tagged “Model capability scaling” · clear filters

All PeopleCompaniesPapersPodcastsHacker News

Browse by tag

26 JUN 2026 · Podcast · No Priors: Artificial Intelligence | Technology | Startups

OpenAI research scientist Noam Brown discusses how traditional AI benchmarks are failing to accurately evaluate modern models due to their increasing reliance on large-scale test-time compute. He argues that model capabilities are now a fun…