Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations
This paper proposes a method to improve the explainability of natural-language autoencoders by making it harder for models to manipulate the explanations, and shows that this approach can increase the reliability of activation explanations and improve AI safety.