Nice work led by @teknology.bsky.social where we develop a memory system to support a heterogeneous population of independently acting agents by sharing experience through retrieval/personalization, contrasting it with memory confined to a single agent or cooperative team. arx…
Who's Who of AI
Fernando Diaz
Why they matter
Researcher with public evidence across AI research, Evaluation & benchmarks.
- AI signals
- 1 past 30d
- Sources
- 1 distinct domains
- Discusiones
- 1 past 30d
- Latest signal
- 26d ago
Associate Professor, CMU. Researcher, Google. Evaluation and design of information retrieval and recommendation systems, including their societal impacts.
What they're sharing
Multi-Agent Transactive Memory arxiv.org
Offline Preference-Based Trajectory Evaluation arxiv.org
Articles & links
AI Weekly's analysis
→
Read full analysis →
Please try to look deeper than success rate for agent evaluation. Thank you. arxiv.org/abs/2606.17541
AI Weekly's analysis
→
- Fernando Diaz argues success-only metrics tie agent comparisons on roughly 75% of instances, gutting statistical power.
- His preference-based trajectory evaluation compares progress and time-to-return profiles, cutting ties to roughly 35%.
- The paper suggests apparent benchmark saturation may reflect the evaluation measure, not exhausted data or problems.
Read full analysis →
Their network
In Fernando Diaz's orbit
Center = Fernando Diaz. Left = members they follow (green edges). Right = members who follow them (blue edges). Top = mutual follows (orange edges, slightly larger). Drag any node to reposition; click to open that profile.