This paper develops a system for training and evaluating AI models that can have natural-sounding conversations with users, using both audio and video input. Practitioners might care because this research could lead to more human-like chatbots that can understand and respond to users in a more intuitive way.
Firehose
Filtered to Papers, tagged “multimodal benchmarks” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News · Deep dives