Nice work led by @teknology.bsky.social where we develop a memory system to support a heterogeneous population of independently acting agents by sharing experience through retrieval/personalization, contrasting it with memory confined to a single agent or cooperative team. arx…
Fernando Diaz
Researcher with public evidence across AI research, Evaluation & benchmarks.
- AI signals
- 0 past 30d
- Sources
- 0 distinct domains
- Discusiones
- 1 past 30d
- Latest signal
- —
Articles & links
Please try to look deeper than success rate for agent evaluation. Thank you. arxiv.org/abs/2606.17541
- Fernando Diaz argues success-only metrics tie agent comparisons on roughly 75% of instances, gutting statistical power.
- His preference-based trajectory evaluation compares progress and time-to-return profiles, cutting ties to roughly 35%.
- The paper suggests apparent benchmark saturation may reflect the evaluation measure, not exhausted data or problems.
In Fernando Diaz's orbit
Center = Fernando Diaz. Left = members they follow (green edges). Right = members who follow them (blue edges). Top = mutual follows (orange edges, slightly larger). Drag any node to reposition; click to open that profile.
Are you Fernando Diaz? Show it.
Add the Who’s Who of AI badge to your site or bio. It links back to this profile.
Markdown: [](https://aiweekly.co/whos-who/person/841io-bsky-social)