simons.berkeley.edu web signal

Schmidt: CLIP robustness traces to data, not language

TL;DR

  • Ludwig Schmidt's August 2023 Berkeley talk surveys robustness findings from over 200 models to isolate why CLIP holds up under distribution shifts.
  • Controlled experiments attribute CLIP's robustness to the pre-training dataset rather than language supervision alone.
  • The talk points at LAION-5B and DataComp as data-side efforts, and introduces fine-tuning via weight interpolation.

At the Simons Institute in August 2023, Ludwig Schmidt of the University of Washington used OpenAI's CLIP as a case study for a plain claim about generalization: the training data is doing the work, not the language side of the pipeline.

The talk abstract, delivered at the Large Language Models and Transformers Workshop in the Calvin Lab Auditorium, surveys "robustness findings from over 200 models across various test scenarios" and reports that CLIP's strength on challenging distribution shifts survives a set of controlled experiments meant to disentangle the cause. Schmidt's answer, as the abstract puts it, is that the effect "originates from 'the pre-training dataset' rather than language supervision alone."

The talk closes on data initiatives rather than model changes, pointing at LAION-5B and DataComp as the routes to stronger pre-training corpora, and introduces "new techniques for model fine-tuning through weight interpolation." The abstract itself does not publish per-benchmark numbers or name the specific 200-plus models compared.

Shared on Bluesky by 1 AI expert