24 upvotes · 17 JUL 2026 · Jiarui Zhang, Muzi Tao, Shangshang Wang et al.
This paper introduces a benchmark called ActiveVision to measure whether large language models (LLMs) exercise active observation, and finds that current LLMs are not robust in this regard, performing poorly on tasks that require repeated visual perception.
3 upvotes · 18 JUL 2026 · Baochen Fu, Wenzhi Deng, Baihao Jin et al.
This paper creates a benchmark to evaluate how well large language models can understand images from optical coherence tomography (OCT), a tool used to diagnose retinal diseases. Practitioners in medical imaging and AI might care because it could help improve the accuracy of AI models in diagnosing and treating retinal diseases.