Firehose

Filtered to Papers, tagged “benchmarking” · clear filters

All PeopleCompaniesPapersPodcastsHacker News

Browse by tag

17 SEP 2026 · Paper

This paper develops a new benchmark for detecting telecom fraud in audio calls, which can adapt to rapidly changing scam patterns and distinguish between fraud and near-domain calls. Practitioners may care about this research if they work on developing audio-based models for telecom fraud detection.

17 SEP 2026 · Paper

This paper develops a method to efficiently scale agent research loops, allowing for more effective self-improvement and reusable improvements across diverse environments. Practitioners might care about this research because it could lead to significant cost savings and improved performance in automated code completion and generation tasks.

16 SEP 2026 · Paper

This paper introduces ProgramDistill, a benchmark that evaluates coding agents on their ability to infer behavior from working software and implement it in an incomplete application. Practitioners in AI/ML and web development might care about this work because it provides a scalable and controlled benchmark for evaluating and training coding agents.