nature.com web signal

Nature paper argues 'open' AI still runs on closed stacks

TL;DR

  • Nature perspective argues 'open' AI weights sit on closed compute, data, and labor stacks that keep power concentrated in a few firms.
  • Nvidia holds 70-90% of the AI chip market and more than four million developers rely on its proprietary CUDA framework.
  • The authors say openness without antitrust enforcement and data privacy protections will not meaningfully redistribute AI power.

A perspective in Nature by David Gray Widder, Meredith Whittaker, and Sarah Myers West argues that the 'open' label on AI systems has come loose from what 'open' meant in free and open-source software, and that the slippage is doing real work to shield the current concentration of power in AI rather than disperse it.

Their argument runs through three affordances that 'open' AI actually provides, transparency, reusability and extensibility, and shows how each is bounded by resources that stay private. Nvidia holds 70-90% of the AI chip market, and more than four million developers rely on CUDA, its proprietary framework that runs only on Nvidia GPUs. Training data is routinely opaque, with the authors noting that many large models neglect providing 'even basic information about the underlying data.' Human labor for labeling, RLHF and moderation stays invisible; the authors cite reporting that OpenAI used Kenyan workers on less than $2 per hour. The dominant model-building frameworks, PyTorch and TensorFlow, are open-source but stewarded by Meta and Google, with Meta's CEO describing PyTorch as 'very valuable for us' because it is 'integrated with our technology stack.'

The paper draws a direct line to earlier waves of corporate open-source. IBM spent $1 billion on Linux to challenge Microsoft, Google open-sourced Android on the way to mobile dominance, and Amazon rebuilt MongoDB as a proprietary SaaS product. In today's cycle, even a well-funded independent like Mistral AI ended up partnering with Microsoft Azure for market access, a pattern the authors treat as symptomatic rather than exceptional.

The perspective is deliberately a policy argument rather than an empirical study, and it does not spell out which specific antitrust or data-privacy measures it says openness would need to be paired with, nor does it quantify how much a fine-tuner can actually reshape a base model's behavior. Read it as a framing piece for regulators and procurement leads, not a benchmark.

The useful takeaway for anyone choosing an 'open' model this quarter is that the label describes weights, not the stack underneath them. If your inference still runs on Nvidia silicon and your distribution still runs through a hyperscaler, the openness you bought is narrower than the word suggests.