VectraYX Paper Shows Lenient Harnesses Credit Prose as Tool Calls
TL;DR
- Two sibling Spanish security models score 0.660 vs 0.650 on the lenient B4 tool-use metric but hit 6/6 vs 0/4-6 on verbatim-reproduction checks.
- On the failed 1B, probability on the tool_call token sits at 10⁻⁴ to 10⁻⁵; a dedicated 6B-token tool-SFT phase never restored the prior.
- A 3.3 GPU-hour, 2,202-step targeted SFT lifted well-formed emission from 0.100 to 0.959; repaired 1B beats 600M on unseen entities, 0.536 vs 0.428 (p=0.004).
Two sibling language models score 0.660 and 0.650 on the same tool-use benchmark. On a verbatim-reproduction check against real training examples, one emits well-formed tool calls on 6 of 6 prompts. The other emits them on 0 of 4 to 6, at every checkpoint tested.
The paper, posted this week to Hugging Face, calls the mechanism plainly: "Keyword harnesses fail open. A fluent model without a capability scores like a model with it." It documents the failure in a matched pair of Spanish security models, VectraYX-600M and VectraYX-1B, which share decoder, tokenizer, and special-token layout.
The abstract puts the gap in a line: "Two numbers that agree to the second decimal place described two behaviors that did not overlap at all." The 600M "spontaneously emits well-formed bash_exec tool calls with contextually sensible commands, while the 1B checkpoints produce disconnected prose and had never once been observed to emit the JSON structure."
A first-token probe locates the problem. On the 1B, probability on the ⟨|tool_call|⟩ token sits at 10⁻⁴ to 10⁻⁵. Probes across 38 archived checkpoints show the 1B's web-heavy pretraining phase erased an earlier memorized prior. The dedicated ≈6B-token tool-SFT phase that followed did not restore it.
What restored it was 3.3 GPU-hours. A targeted recipe (diverse corpus, 5× higher learning rate, 2,202 steps) lifted well-formed emission from 0.100 to 0.959 on all 269 training-corpus rows. On 238 prompts over unseen entities, the repaired 1B passes 0.536 against the 600M's 0.428 (paired p=0.004). On its own corpus, the paper notes, the repaired model "mostly recites" rather than generalizes.
An embedding-drift check adds an unusual footnote. At the chosen learning rate, 97.7% of the bf16 embedding table was bit-identical after training; the repair did not move the trigger token's tied embedding at all. The capability lives in the surrounding network, not in the token's representation.
The paper lands amid a run of work reassessing what agentic benchmarks actually measure; ScholarCatalyst was circulating in the same week's tool-use coverage.
Originally reported by huggingface.co
Read the original article →Original headline: Paper 'Keyword Harnesses Fail Open' Shows Lenient Benchmarks Credit Tool Mentions as Tool Calls