Firehose

Filtered to tagged “Continual learning” · clear filters

All PeopleCompaniesPapersPodcastsHacker News

Browse by tag

18 SEP 2026 · Paper

This paper introduces OmniVBench, a comprehensive benchmark and dataset for evaluating and training reference-to-video generation models that can generate videos with diverse and complex references. Practitioners can use OmniVBench to assess and improve the performance of their R2V models, which is essential for developing more versatile and general video generation capabilities.

18 SEP 2026 · Paper

This paper introduces RecreationWorld, a framework for training hybrid computer-use agents that can seamlessly switch between graphical interaction and software development, and verify their outputs. Practitioners may care about the implications of this work for developing more versatile and reliable AI systems.

18 SEP 2026 · Paper

This paper benchmarks payment authorization in AI agents using a large dataset of attacks, and provides insights into how authorization decisions are made in different models and configurations, which can help developers improve their payment authorization systems.

18 SEP 2026 · Paper

This paper develops a system that allows a design tool to learn and improve its performance over time by adapting to user feedback, and demonstrates its effectiveness in a real-world setting. Practitioners might care about this approach for building more robust and adaptable AI systems.

17 SEP 2026 · Paper

This paper explores how large language models can become overly influenced by their own judgments, leading to a loss of diversity in scientific evaluations. Practitioners should care about this issue because it can impact the quality of reviews and recommendations in AI-assisted scientific evaluation.

17 SEP 2026 · Paper

This paper introduces Paint-Anything, a model that can generate and edit images with any color using a shared interface, allowing for precise control over object colors. Practitioners might care about this because it could enable more flexible and realistic image editing and generation tasks.

17 SEP 2026 · Paper

This paper develops a new benchmark for detecting telecom fraud in audio calls, which can adapt to rapidly changing scam patterns and distinguish between fraud and near-domain calls. Practitioners may care about this research if they work on developing audio-based models for telecom fraud detection.

17 SEP 2026 · Paper

This paper develops a method to compress a cell segmentation model called Cellpose-SAM to reduce its size, making it more suitable for deployment on low-power devices in laboratories. Practitioners who work with stem cell microscopy may care about this research because it helps ensure that compressed models still produce accurate results, which is crucial for applications where model accuracy is critical.

17 SEP 2026 · Paper

This paper investigates how different components of coding harnesses, such as planning, action space, and context management, impact the performance of autonomous coding agents in software engineering tasks. Practitioners might care about understanding how to design harnesses that effectively utilize these components to improve agent performance.

17 SEP 2026 · Paper

This paper develops a method to efficiently scale agent research loops, allowing for more effective self-improvement and reusable improvements across diverse environments. Practitioners might care about this research because it could lead to significant cost savings and improved performance in automated code completion and generation tasks.

17 SEP 2026 · Paper

This paper optimizes the performance of large language models (LLMs) on fine-grained visual perception tasks by learning to selectively focus on relevant regions of the image, rather than relying on high-resolution visual encoding. By doing so, it can improve accuracy with fewer visual tokens, making it more efficient and effective for real-world applications.

17 SEP 2026 · Paper

This paper develops a new framework, JEPA-Anything, that enables predictive models to work across different domains and systems, allowing for world modeling and learning from interaction. Practitioners might care about this because it could lead to more generalizable and versatile AI models.

17 SEP 2026 · Paper

This paper proposes a method to control the length of large reasoning models to make them more efficient, by learning when to use less computation for easy problems and more for hard ones. Practitioners might care about this because it can help improve the accuracy-efficiency trade-offs in complex tasks.

17 SEP 2026 · Paper

This paper proposes a new method for self-retiring on-policy distillation in agentic reinforcement learning, which allows agents to learn more effectively by switching between teacher-student training and reinforcement learning alone. Practitioners might care about this approach because it can improve the performance of agents in complex tasks.

17 SEP 2026 · Paper

This paper develops a new framework called WeVisDoc to improve the performance of document parsing systems by addressing their weaknesses in diverse layouts and acquisition conditions. Practitioners may care about this research if they want to build robust document parsing systems that can handle real-world challenges.

17 SEP 2026 · Paper

This paper investigates why some AI models can generate excessively long responses and proposes a solution to mitigate this issue by aligning the models' termination tokens. Practitioners might care about this problem because it can lead to wasted generation budgets and inefficient AI applications.

17 SEP 2026 · Paper

This paper proposes a framework called SELF-INDEX that allows an index to automatically improve its performance without human intervention, leading to better information retrieval for complex tasks and benefiting applications such as search agents and agent memory systems.

16 SEP 2026 · Paper

This paper proposes a method to optimize audio anti-fraud detection models without modifying the underlying audio-language model, allowing for more adaptable and effective fraud detection. Practitioners may care about this approach as it enables the development of more robust and flexible anti-fraud systems.

16 SEP 2026 · Paper

This paper develops a framework called GAVEL that helps long-horizon language models (LLMs) plan tasks for robots more effectively by predicting and repairing potential errors. Practitioners caring about reliable and efficient robot planning might find this approach useful.

16 SEP 2026 · Paper

This paper evaluates whether a multimodal generative model can reason about the physical world by testing its ability to integrate information from different modalities, such as text, images, video, and audio. Practitioners might care about this research because it can help develop more advanced models that can better understand and generate complex scenarios.

16 SEP 2026 · Paper

This paper investigates how the way large language models generate multiple candidate responses affects their performance and energy consumption. Practitioners might care because optimizing test-time scaling can lead to significant improvements in model accuracy and efficiency.

16 SEP 2026 · Paper

This paper introduces Agora, a system that uses Git to enable collective auto-research by sharing and versioning research results among multiple agents, allowing them to build upon each other's work and avoid duplicated search. Practitioners might care about this because it could lead to more efficient and effective research in areas like AI and machine learning.

16 SEP 2026 · Paper

This paper develops a framework for robots to learn from context without relying on pre-programmed demonstrations, allowing them to adapt to new environments. Practitioners might care because this technology could enable robots to perform tasks more efficiently and effectively in real-world situations.

16 SEP 2026 · Paper

This paper proposes a new framework for Mixture-of-Agents that allows query routing and agent fine-tuning to evolve together, improving the ability of agents to adapt to changing capabilities. Practitioners might care about this approach because it can lead to more efficient and effective data-driven specialization in complex tasks.

15 SEP 2026 · Paper

This paper introduces a benchmark for restoring obfuscated platform messages and investigating associated websites to combat online abuse, with the goal of improving the accuracy of risk reports and user safety. Practitioners might care about this research if they want to develop more effective tools to detect and mitigate online threats.

15 SEP 2026 · Paper

This paper proposes a framework for training-free skill evolution in GUI agents, allowing them to adapt to dynamic graphical user interfaces without additional training. Practitioners may care about this research if they develop GUI agents that need to handle changing environments.

15 SEP 2026 · Paper

This paper proposes a new method for 3D hand mesh reconstruction from egocentric event-based cameras, which can handle low-light conditions and motion blur, and provides more accurate hand information and inter-hand relationships than previous approaches.

15 SEP 2026 · Paper

This paper introduces EvolveTrade, a self-evolving framework that allows large language model trading agents to refine their policies over time, enabling them to adapt to changing market regimes and improve their performance. Practitioners in finance and AI may care about this research as it provides a way to build more robust and adaptive trading agents.

3 AUG 2026 · Podcast · Latent Space: The AI Engineer Podcast

In this episode, Philip Kiely and Ali Taha from Baseten discuss the complexities and innovations in inference engineering for large AI models. They cover topics including model deployment, speculative decoding, quantization, hardware optimi…

4 JUL 2026 · Podcast · "The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

This episode features Ramin Hassani, CEO of Liquid AI, discussing the company's journey from biologically inspired neural networks at MIT to developing device-native foundation models. He makes a technically grounded case for efficient, har…

28 JUN 2026 · Podcast · Machine Learning Street Talk (MLST)

In this episode, Thomas Ahle discusses the development of thermodynamic computing chips and the challenges of chip design automation using AI agents. He explains how his team built an open-source Verilog simulator with AI collaboration to o…