Fernando Diaz

Why they matter

Researcher with public evidence across AI research, Evaluation & benchmarks.

AI signals
0
past 30d
Sources
0
distinct domains
Discusiones
1
past 30d
Latest signal
View every signal from Fernando Diaz →
Associate Professor, CMU. Researcher, Google. Evaluation and design of information retrieval and recommendation systems, including their societal impacts.

Articles & links

Please try to look deeper than success rate for agent evaluation. Thank you. arxiv.org/abs/2606.17541

Offline Preference-Based Trajectory Evaluation arxiv.org
AI Weekly's analysis
  • Fernando Diaz argues success-only metrics tie agent comparisons on roughly 75% of instances, gutting statistical power.
  • His preference-based trajectory evaluation compares progress and time-to-return profiles, cutting ties to roughly 35%.
  • The paper suggests apparent benchmark saturation may reflect the evaluation measure, not exhausted data or problems.
Read full analysis →
View on Bluesky · ♥ 2 ↻ 1 ↩ 1 · 2 from the directory shared this · 67d ago

In Fernando Diaz's orbit

Center = Fernando Diaz. Left = members they follow (green edges). Right = members who follow them (blue edges). Top = mutual follows (orange edges, slightly larger). Drag any node to reposition; click to open that profile.

Are you Fernando Diaz? Show it.

Add the Who’s Who of AI badge to your site or bio. It links back to this profile.

Listed in AI Weekly's Who's Who of AI

Markdown: [![Listed in AI Weekly's Who's Who of AI](https://aiweekly.co/modules/custom/aiweekly_whoswho/images/whoswho-badge.svg)](https://aiweekly.co/whos-who/person/841io-bsky-social)