Firehose

Filtered to Papers · clear filters

All PeopleCompaniesPapersPodcastsHacker News

Browse by tag

18 SEP 2026 · Paper

This paper introduces MintAct, a unified AI model that can navigate and interact with digital environments, such as mobile apps and websites, and perform tasks like using visual tools, in a way that rivals specialized models for each environment. Practitioners might care because MintAct's approach could enable more efficient and scalable AI development for real-world applications.

18 SEP 2026 · Paper

This paper introduces OmniVBench, a comprehensive benchmark and dataset for evaluating and training reference-to-video generation models that can generate videos with diverse and complex references. Practitioners can use OmniVBench to assess and improve the performance of their R2V models, which is essential for developing more versatile and general video generation capabilities.

18 SEP 2026 · Paper

This paper introduces RecreationWorld, a framework for training hybrid computer-use agents that can seamlessly switch between graphical interaction and software development, and verify their outputs. Practitioners may care about the implications of this work for developing more versatile and reliable AI systems.

18 SEP 2026 · Paper

This paper develops a new method for optimizing skills for Large Language Model (LLM) agents, using a graph-structured representation that provides clearer workflow-level guidance and enables more effective exploration of the skill space. Practitioners might care because this approach can improve the performance of LLM agents in various tasks.

18 SEP 2026 · Paper

This paper creates a system called CodeMidas that turns existing open-source code into environments for training coding agents using reinforcement learning. Practitioners might care because it could lead to more diverse and effective coding agents that can learn to fix issues, construct code, and verify their own work.

18 SEP 2026 · Paper

This paper develops a method to improve reasoning models by reducing the discrepancy between a stronger teacher and an on-policy student, which helps to prevent the student from learning the teacher's own flaws. Practitioners might care because it can lead to more accurate models that better represent the capability gap between teachers and students.

18 SEP 2026 · Paper

This paper develops a system for training and evaluating AI models that can have natural-sounding conversations with users, using both audio and video input. Practitioners might care because this research could lead to more human-like chatbots that can understand and respond to users in a more intuitive way.

18 SEP 2026 · Paper

This paper proposes a new architecture for Mixture-of-Experts (MoE) models that balances participation, execution, and materialization costs. Practitioners might care because it can lead to significant performance gains in applications where memory and computational resources are limited.

18 SEP 2026 · Paper

This paper benchmarks payment authorization in AI agents using a large dataset of attacks, and provides insights into how authorization decisions are made in different models and configurations, which can help developers improve their payment authorization systems.

18 SEP 2026 · Paper

This paper introduces Gricea, an open-science platform for conversational AI research that aims to facilitate large-scale studies, replication, and knowledge accumulation by providing a standardized way to report and deploy conversational AI research artifacts. Practitioners in the field of conversational AI can benefit from Gricea by being able to easily construct, reproduce, and extend existing studies.

18 SEP 2026 · Paper

This paper develops a system that allows a design tool to learn and improve its performance over time by adapting to user feedback, and demonstrates its effectiveness in a real-world setting. Practitioners might care about this approach for building more robust and adaptable AI systems.

18 SEP 2026 · Paper

This paper proposes a method to improve a pre-trained robot's performance on long-horizon tasks by focusing on specific subtasks that the robot struggles with, allowing for minimal human intervention. Practitioners might care about this approach as it can significantly increase the success rate of tasks that require precise manipulation.

17 SEP 2026 · Paper

This paper develops a method to help 3D diffusion policies anticipate the future trajectory of an interaction, allowing them to generate more effective actions. Practitioners caring about robotics or manipulation tasks may benefit from this approach, as it can lead to improved performance in tasks like picking and placing objects.

17 SEP 2026 · Paper

This paper explores how large language models can become overly influenced by their own judgments, leading to a loss of diversity in scientific evaluations. Practitioners should care about this issue because it can impact the quality of reviews and recommendations in AI-assisted scientific evaluation.

17 SEP 2026 · Paper

This paper develops a framework called DeformSmith that generates deformable assets for robots, such as 3D models with realistic physical behavior, to simulate and interact with in the real world. Practitioners in robotics and computer vision might care about DeformSmith because it can help create more realistic and interactive simulations of deformable objects.

17 SEP 2026 · Paper

This paper introduces Paint-Anything, a model that can generate and edit images with any color using a shared interface, allowing for precise control over object colors. Practitioners might care about this because it could enable more flexible and realistic image editing and generation tasks.

17 SEP 2026 · Paper

This paper develops a new method for creating robust visual representations in images, which can better handle noisy or distorted data, and improves the performance of downstream tasks such as image classification. Practitioners might care about this approach as it can enhance the robustness of image recognition systems in real-world applications.

17 SEP 2026 · Paper

This paper develops a new benchmark for detecting telecom fraud in audio calls, which can adapt to rapidly changing scam patterns and distinguish between fraud and near-domain calls. Practitioners may care about this research if they work on developing audio-based models for telecom fraud detection.

17 SEP 2026 · Paper

This paper proposes a way to align language models with human values by learning to prioritize different moral values in different situations, and shows that this can lead to more accurate and fair decision-making. Practitioners might care because they want to use language models in applications where values and morality are important, such as in healthcare, law, or education.

17 SEP 2026 · Paper

This paper develops a method to compress a cell segmentation model called Cellpose-SAM to reduce its size, making it more suitable for deployment on low-power devices in laboratories. Practitioners who work with stem cell microscopy may care about this research because it helps ensure that compressed models still produce accurate results, which is crucial for applications where model accuracy is critical.

17 SEP 2026 · Paper

This paper develops a new method for image editing that doesn't require any training data, allowing users to edit images without changing the surrounding content. Practitioners might care about this because it could simplify and improve the editing process in various applications.

17 SEP 2026 · Paper

This paper investigates how different components of coding harnesses, such as planning, action space, and context management, impact the performance of autonomous coding agents in software engineering tasks. Practitioners might care about understanding how to design harnesses that effectively utilize these components to improve agent performance.

17 SEP 2026 · Paper

This paper develops a method to efficiently scale agent research loops, allowing for more effective self-improvement and reusable improvements across diverse environments. Practitioners might care about this research because it could lead to significant cost savings and improved performance in automated code completion and generation tasks.

17 SEP 2026 · Paper

This paper optimizes the performance of large language models (LLMs) on fine-grained visual perception tasks by learning to selectively focus on relevant regions of the image, rather than relying on high-resolution visual encoding. By doing so, it can improve accuracy with fewer visual tokens, making it more efficient and effective for real-world applications.

17 SEP 2026 · Paper

This paper evaluates the ability of general-purpose models to understand and act on spatial intelligence through visual demonstrations, active perception, and metric control. Practitioners might care about this research because it can help develop models that can effectively navigate and interact with their environment.

17 SEP 2026 · Paper

This paper develops a new framework, JEPA-Anything, that enables predictive models to work across different domains and systems, allowing for world modeling and learning from interaction. Practitioners might care about this because it could lead to more generalizable and versatile AI models.

17 SEP 2026 · Paper

This paper introduces a new method for video generation, called Video DeltaNet, which combines attention mechanisms to improve efficiency and quality in livestream video generation. Practitioners may care about this paper if they work on video generation tasks and want to explore more efficient and effective methods.

17 SEP 2026 · Paper

This paper proposes a framework called UFO to evaluate the alignment of multi-modal image generation models, which is crucial for achieving consistency with human judgments. Practitioners might care about this paper because it offers a more comprehensive approach to evaluating multi-modal image generation models.

17 SEP 2026 · Paper

This paper introduces DeepSeek-V4.1-Flash, a more efficient model that reduces the computational cost of long-horizon agents by optimizing its prefill process and cache compression. Practitioners can benefit from this model's improved performance and reduced storage needs for agentic workloads.

17 SEP 2026 · Paper

This paper investigates whether giving a language model extra information, such as a worked solution, improves its learning through on-policy self-distillation. A practitioner might care about how to optimize this technique for better performance.

17 SEP 2026 · Paper

This paper proposes a method to control the length of large reasoning models to make them more efficient, by learning when to use less computation for easy problems and more for hard ones. Practitioners might care about this because it can help improve the accuracy-efficiency trade-offs in complex tasks.

17 SEP 2026 · Paper

This paper proposes a new method for self-retiring on-policy distillation in agentic reinforcement learning, which allows agents to learn more effectively by switching between teacher-student training and reinforcement learning alone. Practitioners might care about this approach because it can improve the performance of agents in complex tasks.

17 SEP 2026 · Paper

This paper develops a new framework called WeVisDoc to improve the performance of document parsing systems by addressing their weaknesses in diverse layouts and acquisition conditions. Practitioners may care about this research if they want to build robust document parsing systems that can handle real-world challenges.

17 SEP 2026 · Paper

This paper investigates why some AI models can generate excessively long responses and proposes a solution to mitigate this issue by aligning the models' termination tokens. Practitioners might care about this problem because it can lead to wasted generation budgets and inefficient AI applications.

17 SEP 2026 · Paper

This paper develops a feed-forward model that can predict the articulation of objects from sparse, unordered point cloud observations, allowing it to learn from multiple views and generalize to new inputs. Practitioners in computer vision and robotics may care about this model as it addresses the challenge of modeling articulated objects from limited and incomplete observations.

17 SEP 2026 · Paper

This paper explores whether supervising both agent actions and environment observations during reinforcement learning improves agent exploration and performance. Practitioners might care because it could lead to better initialization for reinforcement learning tasks.

17 SEP 2026 · Paper

This paper creates a benchmark for Telugu spoken question answering, allowing researchers to evaluate models that can understand and respond to questions in Telugu, a high-resource but understudied language. Practitioners working on natural language processing for low-resource languages might care about this work to develop more accurate and culturally sensitive models.

17 SEP 2026 · Paper

This paper proposes a framework called SELF-INDEX that allows an index to automatically improve its performance without human intervention, leading to better information retrieval for complex tasks and benefiting applications such as search agents and agent memory systems.

16 SEP 2026 · Paper

This paper proposes a method to optimize audio anti-fraud detection models without modifying the underlying audio-language model, allowing for more adaptable and effective fraud detection. Practitioners may care about this approach as it enables the development of more robust and flexible anti-fraud systems.

16 SEP 2026 · Paper

This paper develops a framework called GAVEL that helps long-horizon language models (LLMs) plan tasks for robots more effectively by predicting and repairing potential errors. Practitioners caring about reliable and efficient robot planning might find this approach useful.

16 SEP 2026 · Paper

This paper aims to automate end-to-end business intelligence (BI) tasks using large language models (LLMs), making it easier for users to answer business questions without manual data preparation. Practitioners might care about this research because it could improve the efficiency and productivity of BI workflows.

16 SEP 2026 · Paper

This paper evaluates the trustworthiness of enterprise AI assistants in high-pressure situations, such as hiring, healthcare, and finance, where compliance with rules is crucial. Practitioners should care about this research to ensure their AI assistants are reliable and transparent in complex decision-making scenarios.

16 SEP 2026 · Paper

This paper evaluates whether a multimodal generative model can reason about the physical world by testing its ability to integrate information from different modalities, such as text, images, video, and audio. Practitioners might care about this research because it can help develop more advanced models that can better understand and generate complex scenarios.

16 SEP 2026 · Paper

This paper investigates how the way large language models generate multiple candidate responses affects their performance and energy consumption. Practitioners might care because optimizing test-time scaling can lead to significant improvements in model accuracy and efficiency.

16 SEP 2026 · Paper

This paper introduces ProgramDistill, a benchmark that evaluates coding agents on their ability to infer behavior from working software and implement it in an incomplete application. Practitioners in AI/ML and web development might care about this work because it provides a scalable and controlled benchmark for evaluating and training coding agents.

16 SEP 2026 · Paper

This paper investigates whether people's gaze patterns can reveal how they understand each other in collaborative tasks, and whether this understanding is related to the success of the task. Practitioners working on human-robot collaboration or other tasks with asymmetric information might care about this research because it could help them design better interfaces that take into account how people communicate with each other.

16 SEP 2026 · Paper

This paper creates a system called ScienceIDE that converts scientific code into environments that can be used to train artificial agents to perform scientific tasks. Practitioners might care because this could lead to more efficient and effective ways to develop scientific intelligence.

16 SEP 2026 · Paper

This paper introduces Agora, a system that uses Git to enable collective auto-research by sharing and versioning research results among multiple agents, allowing them to build upon each other's work and avoid duplicated search. Practitioners might care about this because it could lead to more efficient and effective research in areas like AI and machine learning.

16 SEP 2026 · Paper

This paper proposes a new method for aligning large language models with human preferences, called Comparison-based Preference Optimization (ComPO), which is more efficient than existing methods and can mitigate a problem called likelihood displacement. Practitioners might care about this paper because it offers a new approach to aligning LLMs with human preferences, which is essential for developing more reliable and trustworthy AI models.

16 SEP 2026 · Paper

This paper improves autoregressive vision-language-action models by creating a new method for action tokenization that better preserves the relationships between actions, allowing the model to perform more accurately in different contexts. Practitioners might care about this because it could lead to more reliable and generalizable vision-language-action models.

16 SEP 2026 · Paper

This paper investigates a common problem in reinforcement learning for language models called Value Flattening, where critics fail to accurately estimate state values, and proposes a new method, SP^3O, to mitigate this issue by supervising only a few well-separated states per response.

16 SEP 2026 · Paper

This paper develops a new method for image captioning that also grounds each phrase with a specific region of the image, allowing for more accurate and detailed descriptions. Practitioners might care about this work if they're building AI systems that need to understand and interact with the physical world.

16 SEP 2026 · Paper

This paper presents a new technique to reduce memory usage and speed up inference for large neural networks, allowing them to run on consumer hardware with limited memory. A practitioner might care about this because it enables the deployment of large models in edge devices and reduces the need for expensive storage.

16 SEP 2026 · Paper

This paper develops a framework for robots to learn from context without relying on pre-programmed demonstrations, allowing them to adapt to new environments. Practitioners might care because this technology could enable robots to perform tasks more efficiently and effectively in real-world situations.

16 SEP 2026 · Paper

This paper proposes a new framework for Mixture-of-Agents that allows query routing and agent fine-tuning to evolve together, improving the ability of agents to adapt to changing capabilities. Practitioners might care about this approach because it can lead to more efficient and effective data-driven specialization in complex tasks.

15 SEP 2026 · Paper

This paper introduces a benchmark for restoring obfuscated platform messages and investigating associated websites to combat online abuse, with the goal of improving the accuracy of risk reports and user safety. Practitioners might care about this research if they want to develop more effective tools to detect and mitigate online threats.

15 SEP 2026 · Paper

This paper proposes a framework for training-free skill evolution in GUI agents, allowing them to adapt to dynamic graphical user interfaces without additional training. Practitioners may care about this research if they develop GUI agents that need to handle changing environments.

15 SEP 2026 · Paper

This paper develops a framework to evaluate the social reasoning of large language models (LLMs) in a more realistic setting, by simulating interactions between the LLM and users who provide feedback on the LLM's predictions. Practitioners might care about this research because it helps improve the social reasoning of LLMs, which are increasingly used for advice and decision-making.

15 SEP 2026 · Paper

This paper proposes a new method for 3D hand mesh reconstruction from egocentric event-based cameras, which can handle low-light conditions and motion blur, and provides more accurate hand information and inter-hand relationships than previous approaches.

15 SEP 2026 · Paper

This paper introduces EvolveTrade, a self-evolving framework that allows large language model trading agents to refine their policies over time, enabling them to adapt to changing market regimes and improve their performance. Practitioners in finance and AI may care about this research as it provides a way to build more robust and adaptive trading agents.