madesai/what-ai-benchmarks-actually-measure · Datasets at Hugging Face
Summary
madesai/what-ai-benchmarks-actually-measure · Datasets at Hugging Face
Shared on Bluesky by 2 AI experts
-
To make our analysis possible, we collected model outputs and scores from 53 models on 56 capability and safety benchmarks. This data, including item-level model responses and score, is available here huggingface.co/data…
View on Bluesky →
Originally reported by huggingface.co
Read the original article →Original headline: madesai/what-ai-benchmarks-actually-measure · Datasets at Hugging Face