Found first: a primary source the press has not covered yet.
A paper on arXiv describes DumpsterCluster, a 128-GPU cluster of second-hand V100s assembled for $22,000, and measures its carbon cost across a year of production inference. The paper finds the cluster emits over 40x more CO₂ per token than current-generation hardware when serving 70B models under grid-average emissions conditions.
What the source says
Authors Zeyu Cao, Xuan Guo, Cheng Zhang, Cheuk Hang Lau, Ilia Shumailov, and Yiren Zhao sourced V100s at roughly $60 each for a total cluster cost of $22,000, against a new 8-GPU B200 system at $600,000. Pipeline-parallel optimizations gave DumpsterCluster competitive LLaMA-70B throughput; the cluster ran in production for one year. Under grid-average conditions, the carbon cost per token is over 40x higher for 70B models and approximately 4x higher for 8B models. The cluster is economically viable only where electricity is cheap, and in those regions the carbon benefit of reuse also disappears.
Why it matters
The sustainability case for GPU reuse rests on avoiding the manufacturing carbon of new systems by extending hardware life. This paper puts numbers on the operational energy side of that equation. The over-40x CO₂ multiplier for 70B serving means the running carbon cost per inference can outweigh the embodied carbon saved by not buying new hardware. The finding arrives as major data centers retire V100 and A100 fleets in volume and present redeployment as a green option. The carbon gap scales with model size: approximately 4x for 8B models, over 40x for 70B.