Bar Cohen, Kai Golan Hashiloni, Kfir Bar, 2026.
Large language models (LLMs) default to figurative interpretations of idiomatic expressions when context supports them. We investigate whether this bias can be directly controlled via activation steering: asking whether idiomaticity is encoded as a manipulable direction in residual space. We introduce IdioSteer, a small controlled benchmark of 100 sentences across 20 idioms designed to expose interpretation-dependent generation. Using a steering vector applied to five LLMs from two model families, ranging from 2B to 12B parameters, we show that steering systematically shifts model outputs toward literal interpretations while preserving fluency. Three findings stand out: cross-layer steering—vectors from mid-to-late layers injected at early layers—consistently outperforms same-layer interventions; steering vectors transfer to held-out idioms; and effects replicate across architectures and model sizes using the same intervention protocol.
