Interpretability · 18 Aug 2026
Researchers show weak, unnoticed prompt cues can be combined to strongly steer AI models
The technique, which the authors call "model hypnosis," works across model families and sizes and transfers between models, complicating both AI safety review and interpretability work.