Firehose

Filtered to tagged “AI model evaluation” · clear filters

All PeopleCompaniesPapersPodcastsHacker News

Browse by tag

18 SEP 2026 · Anthropic

Anthropic is partnering with Accenture on embedded evaluation of its frontier AI models, which will include evaluating and red-teaming models, conducting alignment assessments, and testing model safeguards. The partnership aims to embed evaluators within Anthropic, providing them with access comparable to an employee's to assess how the company operates and verify its safety commitments. The partnership is non-exclusive, and Anthropic plans to work with other evaluators under different funding arrangements. AI summary

26 JUN 2026 · Podcast · No Priors: Artificial Intelligence | Technology | Startups

OpenAI research scientist Noam Brown discusses how traditional AI benchmarks are failing to accurately evaluate modern models due to their increasing reliance on large-scale test-time compute. He argues that model capabilities are now a fun…