Firehose

Filtered to tagged “Pelican” · clear filters

All PeopleCompaniesPapersPodcastsHacker News

Browse by tag

22 JUL 2026 · Hacker News · 676 pts

A study tested 7 major language models (LLMs) on the "pelican-on-a-bicycle" benchmark, which has become a famous informal benchmark in AI, to determine if AI labs are "benchmaxxing" (maximizing their performance on the benchmark) to gain an advantage in user persuasion. The results showed that none of the models consistently outperformed others on this specific prompt, with no significant differences in quality between the generated images. The study also found that the models were not better at drawing pelicans or bicycles compared to other animals and vehicles. AI summary