This paper develops a new benchmark for detecting telecom fraud in audio calls, which can adapt to rapidly changing scam patterns and distinguish between fraud and near-domain calls. Practitioners may care about this research if they work on developing audio-based models for telecom fraud detection.
Firehose
Filtered to Papers, tagged “benchmarking” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News · Deep dives
Browse by tag
This paper develops a method to efficiently scale agent research loops, allowing for more effective self-improvement and reusable improvements across diverse environments. Practitioners might care about this research because it could lead to significant cost savings and improved performance in automated code completion and generation tasks.
This paper introduces ProgramDistill, a benchmark that evaluates coding agents on their ability to infer behavior from working software and implement it in an incomplete application. Practitioners in AI/ML and web development might care about this work because it provides a scalable and controlled benchmark for evaluating and training coding agents.