Content revision 4e771eaf0390cc87b6c7c6fd8469f3508f6ceeec7119472e6c1ccf17e2e1e5d2 ## Persistent Priors, Preserved Targets: A Stroop-Style Paradigm for Lexical Override Research Source ID research:priorsPersist Original https://www.henryw.me/#/research#priorsPersist I investigated how language models follow a new instruction while established knowledge continues to influence their responses. I designed a Stroop-style task that puts a familiar association in conflict with a temporary definition. A prompt might define doctor as forest , while the usual association with hospital competes with that instruction. Across 11 open-weight models, these familiar associations continued to influence the answer. I used activation patching, replacing internal activations between matched prompts, to trace how a model maintains the instructed meaning. The causal interventions linked successful override to information about the defined word, its assigned meaning and the later query. Preserving the newly assigned meaning was crucial to recovering the intended response. The analysis explains how a model can apply a contextual instruction while established associations remain active.