arxiv.org web signal

Role-specific steering scores 63.2 vs 41.1 on OLMo agents

TL;DR

  • On OLMo-3-7B-Instruct across 275 roles and 228 role-agnostic questions, role-specific directions scored 63.2 versus 41.1 for an assistant-axis control.
  • The workflow sweeps four steering coefficients per role, judges role-profile alignment with GPT-4.1-mini, and either passes or flags each configuration.
  • 38 roles declined across all six measured dimensions, which the authors say rules out a single uniform high-strength steering setting.

A new preprint posted to arxiv.org describes a screening workflow for role-conditioned language-model agents before they are dropped into a simulated population. On OLMo-3-7B-Instruct, the authors apply it to "a mixed 275-role inventory with 228 role-agnostic questions, GPT-4.1-mini prompted role references, and GPT-4.1-mini judges." Role-specific activation directions beat an assistant-axis control from prior persona-vector work, with mean overall scores of 63.2 against 41.1 across the tested grid.

The procedure is deliberately mechanical: define a role profile, extract a role-specific direction, sweep four steering coefficients, evaluate role-profile alignment, pass or flag each candidate configuration. The role-specific directions also "preserve high lexical diversity, while the control drops sharply at larger coefficients," the abstract reports.

The practical output is a per-role screen rather than a global knob. "Most roles improve as steering increases, but 38 roles decline across all six measured dimensions," the paper notes, which is why the authors argue "simulation builders should choose coefficients per role rather than deploy a uniform high-strength setting." The paper is listed for the Social Sim'26 Workshop at COLM 2026, with Mark Riedl and Yonadav G. Shavit among the authors; two researchers we track circulated the link.

Shared on Bluesky by 2 AI experts