LLMs' internal computations are not directly expressed via language, making it impossible to understand how the model thinks, and thus linguistic illegibility is unavoidable. This implies that security mechanisms relying on linguistic self-reporting cannot be completely sound, and alternative sandboxing mechanisms like taint tracking and robust virtualization are needed. Taint tracking can define system state that should not be influenced by model-produced data, regardless of how the model linguistically self-reports. AI summary
Firehose
Filtered to Hacker News, tagged “language model security” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News · Deep dives