huggingface.co web signal

VectraYX Paper Shows Lenient Harnesses Credit Prose as Tool Calls

Agents Open Source ai-business

TL;DR

  • Two sibling Spanish security models score 0.660 vs 0.650 on the lenient B4 tool-use metric but hit 6/6 vs 0/4-6 on verbatim-reproduction checks.
  • On the failed 1B, probability on the tool_call token sits at 10⁻⁴ to 10⁻⁵; a dedicated 6B-token tool-SFT phase never restored the prior.
  • A 3.3 GPU-hour, 2,202-step targeted SFT lifted well-formed emission from 0.100 to 0.959; repaired 1B beats 600M on unseen entities, 0.536 vs 0.428 (p=0.004).

Two sibling language models score 0.660 and 0.650 on the same tool-use benchmark. On a verbatim-reproduction check against real training examples, one emits well-formed tool calls on 6 of 6 prompts. The other emits them on 0 of 4 to 6, at every checkpoint tested.

The paper, posted this week to Hugging Face, calls the mechanism plainly: "Keyword harnesses fail open. A fluent model without a capability scores like a model with it." It documents the failure in a matched pair of Spanish security models, VectraYX-600M and VectraYX-1B, which share decoder, tokenizer, and special-token layout.

The abstract puts the gap in a line: "Two numbers that agree to the second decimal place described two behaviors that did not overlap at all." The 600M "spontaneously emits well-formed bash_exec tool calls with contextually sensible commands, while the 1B checkpoints produce disconnected prose and had never once been observed to emit the JSON structure."

A first-token probe locates the problem. On the 1B, probability on the ⟨|tool_call|⟩ token sits at 10⁻⁴ to 10⁻⁵. Probes across 38 archived checkpoints show the 1B's web-heavy pretraining phase erased an earlier memorized prior. The dedicated ≈6B-token tool-SFT phase that followed did not restore it.

What restored it was 3.3 GPU-hours. A targeted recipe (diverse corpus, 5× higher learning rate, 2,202 steps) lifted well-formed emission from 0.100 to 0.959 on all 269 training-corpus rows. On 238 prompts over unseen entities, the repaired 1B passes 0.536 against the 600M's 0.428 (paired p=0.004). On its own corpus, the paper notes, the repaired model "mostly recites" rather than generalizes.

An embedding-drift check adds an unusual footnote. At the chosen learning rate, 97.7% of the bf16 embedding table was bit-identical after training; the repair did not move the trigger token's tied embedding at all. The capability lives in the surrounding network, not in the token's representation.

The paper lands amid a run of work reassessing what agentic benchmarks actually measure; ScholarCatalyst was circulating in the same week's tool-use coverage.