Firehose

Filtered to Papers, tagged “verifiable explanations” · clear filters

All PeopleCompaniesPapersPodcastsHacker News

Browse by tag

22 JUL 2026 · Paper

This paper proposes a method to improve the explainability of natural-language autoencoders by making it harder for models to manipulate the explanations, and shows that this approach can increase the reliability of activation explanations and improve AI safety.