Excited to share our new paper, “What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks,” accepted as an oral at COLM! arxiv.org/pdf/2609.08812
Meera Desai
Researcher with public evidence across Models & releases, AI research, Evaluation & benchmarks.
- AI signals
- 3 past 30d
- Sources
- 2 distinct domains
- Discusiones
- 0 past 30d
- Latest signal
- 2d ago
Articles & links
Come find us at COLM, I’ll be presenting this paper next Thursday afternoon (10/8, oral session 6 and poster session 6). Full paper here: arxiv.org/abs/2609.08812
To make our analysis possible, we collected model outputs and scores from 53 models on 56 capability and safety benchmarks. This data, including item-level model responses and score, is available here huggingface.co/datasets/mad...
In Meera Desai's orbit
Center = Meera Desai. Left = members they follow (green edges). Right = members who follow them (blue edges). Top = mutual follows (orange edges, slightly larger). Drag any node to reposition; click to open that profile.
Are you Meera Desai? Show it.
Add the Who’s Who of AI badge to your site or bio. It links back to this profile.
Markdown: [](https://aiweekly.co/whos-who/person/madesai-bsky-social)