This paper proposes a method to improve the explainability of natural-language autoencoders by making it harder for models to manipulate the explanations, and shows that this approach can increase the reliability of activation explanations and improve AI safety.
Firehose
Filtered to Papers, tagged “verifiable explanations” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News