Firehose

Filtered to Papers, tagged “vision-language models” · clear filters

All PeopleCompaniesPapersPodcastsHacker News

Browse by tag

18 SEP 2026 · Paper

This paper introduces MintAct, a unified AI model that can navigate and interact with digital environments, such as mobile apps and websites, and perform tasks like using visual tools, in a way that rivals specialized models for each environment. Practitioners might care because MintAct's approach could enable more efficient and scalable AI development for real-world applications.

17 SEP 2026 · Paper

This paper optimizes the performance of large language models (LLMs) on fine-grained visual perception tasks by learning to selectively focus on relevant regions of the image, rather than relying on high-resolution visual encoding. By doing so, it can improve accuracy with fewer visual tokens, making it more efficient and effective for real-world applications.

16 SEP 2026 · Paper

This paper develops a new method for image captioning that also grounds each phrase with a specific region of the image, allowing for more accurate and detailed descriptions. Practitioners might care about this work if they're building AI systems that need to understand and interact with the physical world.

16 SEP 2026 · Paper

This paper develops a framework for robots to learn from context without relying on pre-programmed demonstrations, allowing them to adapt to new environments. Practitioners might care because this technology could enable robots to perform tasks more efficiently and effectively in real-world situations.