Why follow Assistant Professor of AI, founder of the AI Accountability Lab with public evidence across AI research and AI business.
Founder & PI @aial.ie. Assistant Professor of AI, School of Computer Science & Statistics, @tcddublin.bsky.social AI accountability, AI audits & evaluation, critical d…
722 peer trust
8 signals / 30d
active 11h ago
Why follow AI accountability, audits and evaluation with public evidence across Culture, work & education and Evaluation & benchmarks.
AI accountability, audits & eval. Keen on participation & practical outcomes. CS PhDing @UCBerkeley.
632 peer trust
0 signals / 30d
active 8d ago
Why follow Researcher with public evidence across Models & releases and AI research.
language models, data & evals, prev co-lead of Olmo @ai2.bsky.social, nlp @uwcse, statistics @uw, open science, tabletop, seattle, he/him,🧋 kyleclo.com
526 peer trust
4 signals / 30d
active 18d ago
Why follow Researcher with public evidence across AI research and Culture, work & education.
#nlp researcher interested in evaluation including: multilingual models, long-form input/output, processing/generation of creative texts previous: postdoc @ umass_nlp …
382 peer trust
3 signals / 30d
active 21h ago
Why follow Researcher with public evidence across AI research and Evaluation & benchmarks.
To mine own opinions be true Research scientist @ Google Deepmind. Prev: postdoc @ UW, PhD UCSB, BSMS ASU Evals, metrics, multilinguality, multiculturality, multimodal…
379 peer trust
0 signals / 30d
active 22h ago
Why follow Researcher with public evidence across AI research and Responsible AI.
VP and Distinguished Scientist at Microsoft Research NYC. AI evaluation and measurement, responsible AI, computational social science, machine learning. She/her. One p…
342 peer trust
0 signals / 30d
active 174d ago
Why follow Researcher with public evidence across AI research and Evaluation & benchmarks.
PhD @ ETH Zürich | working on (multilingual) evaluation of NLP | on the academic job market | go #vegan | https://vilda.net
332 peer trust
1 signals / 30d
active 2d ago
Why follow Researcher with public evidence across AI research and Culture, work & education.
AI PhDing at Mila/McGill. Happily residing in Montreal 🥯❄️ Academic stuff: language grounding, vision+language, interp, rigorous & creative evals, cogsci Other: many s…
300 peer trust
0 signals / 30d
active 48d ago
Why follow Researcher with public evidence across AI research and Evaluation & benchmarks.
Associate Professor, CMU. Researcher, Google. Evaluation and design of information retrieval and recommendation systems, including their societal impacts.
279 peer trust
1 signals / 30d
active 2d ago
Why follow Researcher with public evidence across Safety & security and AI research.
Robustness, Data & Annotations, Evaluation & Interpretability in LLMs http://mimansajaiswal.github.io/
262 peer trust
0 signals / 30d
active 421d ago
Why follow Directory member with public evidence across Evaluation & benchmarks.
journalist and writer. author of BEYOND MEASURE, a history of measurement; a New Yorker, Economist, Times book of the year. former senior editor at The Verge. buy my b…
218 peer trust
0 signals / 30d
active 7h ago
Why follow Practitioner with public evidence across Evaluation & benchmarks.
evals evals evals. https://evals.info
200 peer trust
1 signals / 30d
active 257d ago
Why follow Researcher with public evidence across Models & releases and AI research.
PhD student at University of Michigan School of Information. Data and evaluation practices for language models / language models as cultural technologies. https://meer…
195 peer trust
0 signals / 30d
active 8d ago
Why follow Researcher with public evidence across AI research and Evaluation & benchmarks.
LTI PhD at CMU on evaluation and trustworthy ML/NLP, prev AI&CS Edinburgh University, Google, YouTube, Apple, Netflix. Views are personal 👩🏻💻🇮🇩 athiyadeviyani.github.io
165 peer trust
0 signals / 30d
active 79d ago
Why follow Researcher with public evidence across AI research and Evaluation & benchmarks.
machine learning asst prof at @cs.ubc.ca and amii statistical testing, kernels, learning theory, graphs, active learning… she/her, 🏳️⚧️, @queerinai.com
164 peer trust
0 signals / 30d
active 75d ago
Why follow Researcher with public evidence across AI business and AI research.
CS Prof at UC Irvine, CTO/Cofounder at Envive AI Work on evaluation and robustness of LLMs
152 peer trust
0 signals / 30d
active 92d ago
Why follow Researcher with public evidence across AI research and Evaluation & benchmarks.
PhD student @CMU LTI NLP | IR | Evaluation | RAG https://kimdanny.github.io
148 peer trust
0 signals / 30d
active 31d ago
Why follow Researcher with public evidence across AI research and Evaluation & benchmarks.
Machine learning prof at U Toronto. Working on evals and AGI governance.
143 peer trust
0 signals / 30d
active 66d ago
Why follow Researcher with public evidence across AI research and Agents & robotics.
I develop tough benchmarks for LMs and then I build agents to try and beat those benchmarks. Postdoc @ Princeton University. https://ofir.io/about
141 peer trust
0 signals / 30d
active 349d ago
Why follow Directory member with public evidence across Agents & robotics and Evaluation & benchmarks.
NLP PhD student @convai_uiuc | Agents, Reasoning, evaluation etc. https://sagnikmukherjee.github.io https://scholar.google.com/citations?user=v4lvWXoAAAAJ&hl=en
121 peer trust
0 signals / 30d
active 306d ago
Why follow Practitioner with public evidence across AI research and Evaluation & benchmarks.
Applied scientist working on LLM evaluation and publishing in AI ethics. Formerly: technical writing, philosophy. Urbanism nerd in my spare time. Opinions here my own.…
116 peer trust
0 signals / 30d
active 8h ago
Why follow Researcher with public evidence across Responsible AI and AI research.
PhD student at Johns Hopkins University Alumni from McGill University & MILA Working on NLP Evaluation, Responsible AI, Human-AI interaction she/her 🇨🇦
105 peer trust
0 signals / 30d
active 296d ago
Why follow Directory member with public evidence across AI research and Evaluation & benchmarks.
Visiting PhD at Stanford🌲, CS PhD student at NUS 🇸🇬, PhD Fellow @ Google, NLP researcher📒 https://yocodeyo.github.io Working on Social Intelligence and Evaluation
104 peer trust
0 signals / 30d
active 590d ago
Why follow Directory member with public evidence across AI research and Evaluation & benchmarks.
I share non-hype ML & AI with 8+ years XP 🌦 Machine Learning at ECMWF 💾 Fellow at SSI You miss 99% of benchmarks you don't overfit on! 🏳️🌈 they
94 peer trust
0 signals / 30d
active 59d ago
Why follow Practitioner with public evidence across Evaluation & benchmarks and NLP & language.
I'm a software developer, open-source maintainer and NLP engineer. I love data, coding, testing, developing ML models and getting sh*t done. https://github.com/svlandeg
89 peer trust
0 signals / 30d
active 7d ago
Why follow Directory member with public evidence across Evaluation & benchmarks and Safety & security.
AI Evaluation and Interpretability @MicrosoftResearch, Prev PhD @CMU.
88 peer trust
0 signals / 30d
active 422d ago
Why follow Researcher with public evidence across AI research and Culture, work & education.
ELLIS PhD Fellow @belongielab.org | @aicentre.dk | University of Copenhagen | @amsterdamnlp.bsky.social | @ellis.eu Multi-modal ML | Alignment | Culture | Evaluations …
75 peer trust
0 signals / 30d
active 283d ago
Why follow Directory member with public evidence across AI research and AI business.
🧪 Data science, survey science, social science 💻 Director of Data Science @ Microsoft Garage [Posts do not represent my employer] 🧮 Stats, R, python 📝 Science, Researc…
72 peer trust
0 signals / 30d
active 237d ago
Why follow Researcher with public evidence across AI research and Evaluation & benchmarks.
Professor of Interactive Intelligent Systems @unisalzburg.bsky.social. Co-lead of focus area @intermediation-sbg.bsky.social @ Wissenschaft & Kunst. #recsys, #fairness…
72 peer trust
1 signals / 30d
active 3d ago
Why follow Directory member with public evidence across Evaluation & benchmarks.
METR is a research nonprofit that builds evaluations to empirically test AI systems for capabilities that could threaten catastrophic harm to society.
68 peer trust
2 signals / 30d
active 2d ago
Why follow Researcher with public evidence across Responsible AI and AI research.
Assistant Professor @ University of Cambridge. Responsible AI. Human-AI Collaboration. Interactive Evaluation. umangsbhatt.github.io
68 peer trust
0 signals / 30d
active 139d ago
Why follow Policy with public evidence across Culture, work & education and Evaluation & benchmarks.
evals accelerationist, Head of Policy at @metr.org, working hard on responsible scaling policies Check out my artisanal hand-crafted "AI Bluesky" starter pack here: ht…
64 peer trust
0 signals / 30d
active 65d ago
Why follow Policy with public evidence across Policy & governance and Evaluation & benchmarks.
AI Policy @ Stanford + Harvard Former - Digital Forensic Research Lab, Digital Impact Alliance, UN Global Pulse, Human Rights Watch Researching - acceptable use of dig…
56 peer trust
0 signals / 30d
active 451d ago
Why follow Researcher with public evidence across AI research and Evaluation & benchmarks.
Prof of Comm and Polisci @UMich studying how people get & use political information and social measurement. Competencies: data sci, DIY solar, election analytics, carp…
55 peer trust
0 signals / 30d
active 6d ago
Why follow Researcher with public evidence across AI research and AI business.
Research Scientist at Google | Interested in rigorous evaluations, cognitive science, multi-agent systems and causality | Prev: Vector Institute, Mila, Max Planck Inst…
36 peer trust
0 signals / 30d
Why follow Researcher with public evidence across AI research and Evaluation & benchmarks.
Research Scientist at Google DeepMind Gemini evals and post-training sebarnold.net
32 peer trust
0 signals / 30d
Why follow Researcher with public evidence across AI research and Culture, work & education.
Academic (Dr.), Geographer, Tinkerer, Blabberer, Writer. Testing this space. Interests: Tech | Work | Political Economies | Cities | Global Dev | Dissident Cultures | …
26 peer trust
0 signals / 30d
active 8d ago