🧵There is a fundamental issue with reference-based LLM-judges. People implicitly assume a reference-based judge behaves like: score=f(candidate,reference)score=f(candidate,reference) However, the actual behavior is closer to: score=f(candidate,reference,parametric knowledge,pr…
Leo Boytsov
Directory member with public evidence across AI research, NLP & language.
- AI signals
- 2 past 30d
- Sources
- 1 distinct domains
- Discussions
- 0 past 30d
- Latest signal
- 3d ago
Articles & links
This meticulous study delves into the intricate tapestry of lexical biases in LLM-assisted academic writing. It underscores a nuanced interplay of preferred words, unveiling how they have dramatically enhanced and reshaped the realm of science. www.linkedin.com/feed/update/...
🔹Now long-running asynchronous agents expose a weakness in GRPO, so this paper brings the critic back, but with several engineering fixes. 🟦 www.linkedin.com/posts/ravid-...
Recent commentary
Interestingly we went full-circle in RL for LLMs and LLM agents: 🔹Initially, OpenAI (and some others) used RLHF with PPO, which requires training a critic (reward) model. 🔹Then, researchers moved from PPO with a critic to critic-free GRPO because critics were expensive and unstable. ↩️
Fun fact: Two some of the most influential statistical NLP papers were authored by Brown et al in 1990 & 2020 1990 Statistical Approach to Machine Translation 2020 Language Models are Few-Shot Learners *) It is not the same Brown **) I believe author names in IBM papers were ordered alphabetically
In Leo Boytsov's orbit
Center = Leo Boytsov. Left = members they follow (green edges). Right = members who follow them (blue edges). Top = mutual follows (orange edges, slightly larger). Drag any node to reposition; click to open that profile.