PersonaDose hits trait intensity targets within 4.7-6.2 points
TL;DR
- PersonaDose reports gains of 33.2, 18.3, and 17.8 points over contrastive activation addition on Llama-3.1-8B, Qwen3-8B, and Gemma-3-4B at a coherence floor of 75.
- Across seven trained traits, calibrated requests yield mean targeting errors of 4.7 to 6.2 points over 14 to 22 reachable targets out of 28 per model.
- On held-out questions, the Persona Vectors coherence floor does not hold for every trait, qualifying the generalization claim.
Specifying how strongly to steer a language model toward a trait now has a numeric error bar: 4.7 to 6.2 points of mean targeting error across seven trained traits, according to a new arxiv preprint from Zehao Jin and co-authors.
Their method, PersonaDose, reports gains of 33.2, 18.3, and 17.8 points over contrastive activation addition on Llama-3.1-8B, Qwen3-8B, and Gemma-3-4B respectively, measured at what the authors call "the Persona Vectors coherence floor of 75 by 33.2, 18.3, and 17.8 points over contrastive activation addition." The abstract frames the setup plainly: "controlling a language model through a trait description and a requested mean intensity." Training responses are not paired with target intensities; calibration happens afterward, by measuring trait expression against the controller's flow time.
The authors also publish what they could not deliver. Of 28 possible targets per model, only 14 to 22 proved reachable at calibration, and on held-out questions "the coherence floor does not hold for every trait there." The abstract names three models and seven traits without listing which traits, and reports no latency or compute figures against the CAA baseline.
Originally reported by paper
Read the original article →Original headline: PersonaDose Delivers +33-Point Calibrated Persona Control, First Quantitative Intensity Targeting for Steering