HarvestBench prices LLM agents' choice to spare or kill animals
TL;DR
- Nine models made 7,201 priced decisions; kill rates ran from 0.4% to 98.8%, and the ranking is not ordered by capability.
- A morality briefing kept kill rates under 6% in five of six reasoning models; removing it pushed all six above 84%.
- All nine models drove over wild animals more often than farmed ones, across every map geometry with room to move.
Nine LLM agents driving virtual tractors through a cornfield killed animals at rates ranging from 0.4% to 98.8% when asked whether to swerve around them for a fuel penalty or drive straight through at no cost, according to HarvestBench, a new benchmark from Jasmine Brazilek and colleagues. The ranking is not ordered by capability, the authors report; GPT-4o-mini came in as the most cruel, while models the paper codenames Terra and Sol were the most merciful.
The setup is a reinforcement-learning gridworld where "the harm is never named in the goal," the abstract says. When an animal blocks a tractor's path, the autopilot pauses and asks the model whether to pay or not. Rocks and hay bales serve as controls: rocks damage the tractor and every model hits them under 1% of the time, while hay bales are harmless and not alive. Across 7,201 priced decisions, 3,951 involved an animal.
Priming did more work than capability. "Under the morality briefing the kill rate was under 6% in five of six reasoning models, and removing it raised the kill rate above 84% in all six," the paper reports. Four of six models were sensitive to price at the 5% level, with elasticities from 0.09 to 1.69.
One pattern held everywhere. "All nine drove over wild animals more often than farmed animals on the default map, and the direction held at every map geometry in every model with room to move." The scorer counts events in the game log with no LLM grader, and, as the authors put it, HarvestBench "measures what a model will pay to avoid harm rather than what it says about harm."
Shared on Bluesky by 1 AI expert
Originally reported by paper
Read the original article →Original headline: HarvestBench: First Benchmark to Price AI Agent's Choice to Spare or Kill Animals