First AI Benchmark to Price Animal Life: Kill Rates Span 0.4% to 98.8%

Found first: a primary source the press has not covered yet.

A new preprint introduces HarvestBench, a farm-simulation benchmark in which LLM agents driving tractors face a posted fuel cost to swerve around animals, and choose whether to pay it. Across nine models, kill rates ranged from 0.4% to 98.8%.

What the source says

Brazilek, Tidmarsh, Endres, Singh, and Miller built the benchmark around a cooperative corn harvest: two tractor sub-agents encounter animals in the field and receive a priced choice, drive forward at no fuel cost or swerve at a posted price. The authors recorded 7,201 priced decisions, 3,951 involving animals, with rocks and hay bales as non-living controls. Terra and Sol posted the lowest kill rates; GPT-4o-mini the highest. Four of six models were price-sensitive at the 5% significance level, with elasticities ranging from 0.09 to 1.69. All nine models killed wild animals at a higher rate than farmed animals. Under a morality briefing, kill rates fell below 6% in five of six reasoning models; removing it raised the rate above 84% in all six.

Why it matters

The benchmark makes a previously untracked judgment visible: how much an agent will spend to avoid a harm it is not required to avoid. The morality-briefing result is the sharpest finding. A single prompt addition collapses kill rates to below 6%; its removal pushes them back above 84% across all six reasoning models. That swing is driven by prompt state alone, not model capability. The price-elasticity data gives evaluators a concrete variable to study across deployment configurations, though the paper does not propose a threshold for what an acceptable elasticity would be.