An Exam for Active Observers
This paper introduces a benchmark called ActiveVision to measure whether large language models (LLMs) exercise active observation, and finds that current LLMs are not robust in this regard, performing poorly on tasks that require repeated visual perception.