Firehose

Filtered to tagged “natural-language autoencoders” · clear filters

All PeopleCompaniesPapersPodcastsHacker News

Browse by tag

22 JUL 2026 · Paper

This paper proposes a method to improve the explainability of natural-language autoencoders by making it harder for models to manipulate the explanations, and shows that this approach can increase the reliability of activation explanations and improve AI safety.