angela zhou

Why they matter

Researcher with public evidence across AI research, Compute & infrastructure, Responsible AI.

AI signals
6
past 30d
Sources
5
distinct domains
Discussions
10
past 30d
Latest signal
3d ago
View every signal from angela zhou →
assistant prof at USC Data Sciences and Operations and Computer Science; phd Cornell ORIE. data-driven decision-making, operations research/management, causal inference, algorithmic fairness/equity bureaucratic justice warrior angelamzhou.github.io

Articles & links

angela zhou reposted
Fernando Diaz @841io.bsky.social

Please try to look deeper than success rate for agent evaluation. Thank you. arxiv.org/abs/2606.17541

Offline Preference-Based Trajectory Evaluation arxiv.org
AI Weekly's analysis
  • Fernando Diaz argues success-only metrics tie agent comparisons on roughly 75% of instances, gutting statistical power.
  • His preference-based trajectory evaluation compares progress and time-to-return profiles, cutting ties to roughly 35%.
  • The paper suggests apparent benchmark saturation may reflect the evaluation measure, not exhausted data or problems.
Read full analysis →
View on Bluesky →

You might also like a recent perspective piece we have, taking a program evaluation point of view on algorithmic accountability: arxiv.org/pdf/2606.25668 it's about predictive, not GenAI; but is refocusing the role of social systems in AI evaluation

arxiv.org
View on Bluesky · ♥ 1 ↻ 0 ↩ 1 · 3d ago

Recent commentary

i don't think we should waste real expert reviewer's time reviewing mediocre AI-written papers and having real expert reviewer provide valuable feedback that, presumably, will just get sent to LLMs to complete. at that point the reviewers might as well be authoring along with the authors

View on Bluesky · ♥ 38 ↻ 0 ↩ 3 · 71d ago

need an ai agent to pm me on my multiple projects (or keep me from adding too many projects to my stack)

View on Bluesky · ♥ 14 ↻ 0 ↩ 5 · 83d ago

feeling .... extremely depressed today how llms are giant plagiarism machines, and how much work on the side of review it takes to point that out.

View on Bluesky · ♥ 12 ↻ 1 ↩ 1 · 12d ago

I think often in AI for social good research spaces, we talk about projects as just consulting projects (derogatory) But actually, I think there's a lot of worth in developing consulting capabilities: paying attention to client needs, building mental models for client goals, processes, resources

View on Bluesky · ♥ 10 ↻ 0 ↩ 2 · 6d ago

i feel like prototyping an ai product or workflow to Work Well and Add Value instead of becoming More Workslop is like cooking. you need to be tasting as you go, and even as models are getting better, you need someone to be tasting as you go along (where tasting = looking at the data)

View on Bluesky · ♥ 3 ↻ 0 ↩ 2 · 20d ago

In angela zhou's orbit

Center = angela zhou. Left = members they follow (green edges). Right = members who follow them (blue edges). Top = mutual follows (orange edges, slightly larger). Drag any node to reposition; click to open that profile.