Papers

Filtered to computer vision · clear filter

Browse by term

continual learning 33reinforcement learning 21large language models 9vision-language models 8language models 6benchmarking 5autoregressive models 4diffusion transformers 4generative models 4multimodal learning 4robotics 4video generation 4benchmarks 3computer vision 3policy optimization 3self-distillation 3vision-language-action models 3world modeling 3agent-based systems 2autonomous agents 2calibration 2coding agents 2diffusion models 2foundation models 2image synthesis 2in-context learning 2knowledge graphs 2LLMs 2multimodal evaluation 2multimodal large language models 2

Matching papers

Trajectory-aware Cross-view Geo-localization with Sequential Observations

5 upvotes · 16 JUL 2026 · Tianyi Gao, Jiayu Lin, Danielle Beaulieu et al.

This paper develops a new method for cross-view geo-localization that uses video clips and route descriptions to improve accuracy, and introduces a unified framework that can handle both modalities. Practitioners in autonomous vehicle development or geospatial analysis might care about this research as it aims to address a common challenge in these fields.

ReViV: Reconstructing the Viewer and the View in 4D from Monocular Egocentric Video

2 upvotes · 20 JUL 2026 · Xiaozhong Lyu, Gen Li, Zhiyin Qian et al.

This paper develops a new AI model called ReViV that can reconstruct the viewer's and the viewed environment in 4D from a single video taken from a wearable camera. Practitioners in fields like augmented and virtual reality, human-computer interaction, and computer vision might care about this model because it can efficiently and accurately track the viewer's movements and the environment in real-time.