Interpretability
Researchers show weak, unnoticed prompt cues can be combined to strongly steer AI models
The technique, which the authors call "model hypnosis," works across model families and sizes and transfers between models, complicating both AI safety review and interpretability work.
Original cover art, generated for this story. THE VISSION does not republish third-party press imagery.
The short version
- A new paper titled "Model Hypnosis" finds that individually weak, inconspicuous prompt cues — paraphrases, minor typos — can be combined to strongly steer model behavior.
- The authors, Enric Boix-Adsera and Benedict Tessler, report the effect occurs across model families and scales, including in frontier reasoning models.
- Hypnotic prompts built against one model can transfer to a different one, according to the paper, which the researchers argue poses a challenge for both safety evaluation and interpretability.
The technique, which the authors call "model hypnosis," works across model families and sizes and transfers between models, complicating both AI safety review and interpretability work.