Shubhendu Trivedi

Why they matter

Directory member with public evidence across Models & releases.

AI signals
9
past 30d
Sources
5
distinct domains
Discussions
42
past 30d
Latest signal
2h ago
View every signal from Shubhendu Trivedi →
Interests on bsky: ML research, applied math, and general mathematical and engineering miscellany. Also: Uncertainty, symmetry in ML, reliable deployment; applications in LLMs, computational chemistry/physics, and healthcare. https://shubhendu-trivedi.org

Articles & links

One could read most points here cynically. But could also take them at their word and see what could be done. Given the sort of equilibrium we have been post GPT-2, the sort of pause they are advocating is simply not going to happen. You are talking of powerful international a…

When AI builds itself anthropic.com
View on Bluesky · ♥ 16 ↻ 1 ↩ 2 · 18 from the directory shared this · 115d ago

Diffusion Gemma seems quite cool. Going to look into it during the weekend (so another exercise in harness design). It's funny and nice to see Google releasing open models one after the other, with a focus on the small end (quite a significant part of the enterprise ecosystem)…

DiffusionGemma: 4x faster text generation blog.google
AI Weekly's analysis →
  • DiffusionGemma generates 256 tokens per forward pass using bidirectional attention, reaching 1,000+ tokens/sec on a single H100 GPU.
  • With only 3.8B active parameters during inference and an 18GB VRAM footprint when quantized, it runs on consumer hardware without server-grade resources.
  • Google recommends DiffusionGemma only for speed-critical workloads like in-line editing and code infilling, not for applications requiring maximum quality.
Read full analysis →
View on Bluesky · ♥ 1 ↻ 0 ↩ 1 · 6 from the directory shared this · 108d ago

Good article. I don't know and don't care who's Prince. But I like a good Drucker defense. www.programmablemutter.com/p/ai-isnt-ma...

programmablemutter.com
AI Weekly's analysis →
  • Cloudflare CEO Matthew Prince laid off more than 20% of his workforce while citing Peter Drucker's 1954 management framework to explain the cuts.
  • Prince labeled laid-off workers 'measurers' — a Drucker category he applied to middle management, finance, legal, and internal auditing.
  • Henry Farrell argues Prince inverted Drucker's intent: Drucker used measurement to develop managers, not to identify who to eliminate.
Read full analysis →
View on Bluesky · ♥ 2 ↻ 0 ↩ 1 · 7 from the directory shared this · 117d ago

Not the easiest read, and perhaps too stylized, but a really cool paper: arxiv.org/abs/2609.01595

Mechanism Design for Alignment and Control arxiv.org
AI Weekly's analysis →
  • Three economists frame AI alignment as a mechanism-design problem where both the agent's preferences and its capabilities are unknown to the principal.
  • The key assumption is a 'one-sided imitation structure' in which an AI agent can hide capabilities but cannot fabricate ones it lacks.
  • The framework is applied to sandbagging, an alignment–interpretability trade-off, peer scoring, reward coupling and scalable oversight.
Read full analysis →
View on Bluesky · ♥ 8 ↻ 1 ↩ 1 · 3 from the directory shared this · 2h ago
↻ Shubhendu Trivedi reposted
Tom McCoy @rtommccoy.bsky.social

Paper: arxiv.org/abs/2608.29530 We consider a variety of neural networks that perform seemingly-symbolic tasks, from small-scale models trained on simple list-manipulation tasks to LLMs performing tasks in math, logic, coding, and language. 2/n

The Emergent Symbolic Structure of Artificial Neural Networks arxiv.org
AI Weekly's analysis →
  • A new preprint reports that neural networks' vector representations can be closely approximated by closed-form symbolic equations with behavior 'largely unchanged.'
  • The substitution was tested on small networks trained to manipulate lists and on LLMs across arithmetic, logic, computer code, and language.
  • The authors say the symbolic approximation also lets them steer an LLM's behavior via targeted interventions on its internal representations.
Read full analysis →
View on Bluesky →
↻ Shubhendu Trivedi reposted
noemielteto.bsky.social @noemielteto.bsky.social

14/14 By linking data-driven hypothesis generation with hypothesis-driven experiment selection, ATLAS shows real potential to accelerate the discovery of interpretable insights in cognitive science and beyond. Read the full paper here: arxiv.org/abs/2606.12386

ATLAS: Active Theory Learning for Automated Science arxiv.org
AI Weekly's analysis →
  • ATLAS combines sparse neural networks with active learning to generate and test mechanistic hypotheses automatically.
  • The system achieved 5-10x better sample efficiency than random experimentation across all evaluation metrics.
  • Validation compared ATLAS-designed experiments against expert-designed ones from published cognitive science literature.
Read full analysis →
View on Bluesky →
↻ Shubhendu Trivedi reposted
Sam Power @spmontecarlo.bsky.social

arxiv.org/abs/2608.273... 'Conformal Prediction Through the Lens of Hypothesis Testing: Universality, Impossibility, and Optimality' - Ryan J. Tibshirani, Rina Foygel Barber, Aaditya Ramdas good stuff if you find insight in test-based interpretations of statistical procedures …

Conformal Prediction Through the Lens of Hypothesis Testing: Universality, Impossibility, and Optimality arxiv.org
AI Weekly's analysis →
  • Tibshirani, Barber and Ramdas recast conformal prediction as the inversion of a permutation test for exchangeability of the n+1 joint sample.
  • The authors show foundational universality and impossibility results in conformal prediction can be reproduced using classical hypothesis testing theory from Neyman, Lehmann, Scheffé, Kraft and Le Cam.
  • For any joint distribution of X,Y and any sample size, they prove the optimal conformal score is the inverse conditional density of Y given X.
Read full analysis →
View on Bluesky →
↻ Shubhendu Trivedi reposted
Chris Amato @cjdamato.bsky.social

A new version of my book on cooperative multi-agent reinforcement learning is now available. A longer version will eventually be published with Frans Oliehoek so let us know your thoughts! arxiv.org/abs/2405.06161

An Initial Introduction to Cooperative Multi-Agent Reinforcement Learning arxiv.org
AI Weekly's analysis →
  • Amato organizes cooperative MARL around three settings: centralized training and execution, centralized training for decentralized execution, and decentralized training and execution.
  • The paper walks through independent Q-learning, value factorization methods VDN, QMIX and QPLEX, and centralized critic methods MADDPG, COMA and MAPPO.
  • Amato writes that CTDE is the most common paradigm because it leverages centralized information at training while keeping execution decentralized.
Read full analysis →
View on Bluesky →

Don't know the author, but have become quite a fan of her work. Is always quite cool and original (sometimes conceptually, sometimes in terms of theoretical machinery &c.) arxiv.org/abs/2407.02458

Statistical Advantages of Oblique Randomized Decision Trees and Forests arxiv.org
AI Weekly's analysis →
  • Eliza O'Reilly's paper proves oblique Mondrian forests achieve minimax optimal convergence rates on ridge-function data where axis-aligned trees cannot.
  • For general ridge functions, no weighting of axis-aligned splits can match the rate oblique splits obtain, regardless of covariate distribution.
  • The analysis uses random tessellation theory from stochastic geometry, tying convergence to the relevant feature subspace rather than ambient dimension.
Read full analysis →
View on Bluesky · ♥ 12 ↻ 2 ↩ 0 · 2 from the directory shared this · 98d ago

Recent commentary

The dot com era mega IPOs have a very different character than the 2026 AI ones. Back then you had many companies that were capital-starved before their IPOs and capital-rich after. The IPO itself was quite often a major financing event. Basically, public markets funded the next stage of growth.

View on Bluesky · ♥ 5 ↻ 2 ↩ 2 · 107d ago

Many aspects of the AI twitter peanut gallery seem to have spontaneously emerged on bluesky as well. Many of the opinion microstates and occupancy boxes are thinly traded given the natural constraints of bluesky, but you can see the broader contours.

View on Bluesky · ♥ 6 ↻ 0 ↩ 3 · 100d ago

One thing that has become clear to me just very recently from dozens of conversations with folk at all the frontier labs: Many really do believe that once "AGI happens" it will "make everything easier, from robotics, to manufacturing." Supply chain constraints, ecosystem and labour development

View on Bluesky · ♥ 7 ↻ 1 ↩ 1 · 127d ago

For some reason (ok I partly know the reason) I get a lot of email from kids in middle/high school, and off late I've been unable to do justice to them beyond replying (my routine is messed up). But this was by someone younger, clearly wrote everything himself, even has an AI declaration at the end.

View on Bluesky · ♥ 7 ↻ 0 ↩ 1 · 11d ago

It was easy to guess what "this model is too dangerous too release" meant: we don't have enough computational resources to serve. It was easy, in hindsight, to guess what "we are nearing RSI.. pause AI" meant.

View on Bluesky · ♥ 8 ↻ 0 ↩ 0 · 110d ago

Due to concerns with AI safety I am delayed with a project. I also keep getting writer's block (and writer's thought-train-disruption by LLM) which doesn't help. Humanity won't survive this way.

View on Bluesky · ♥ 5 ↻ 0 ↩ 1 · 18d ago

Failed bond auction (2nd worst on record), international T Bill purchases down 80% YoY. Interesting times. Building up for more than a year, and now it's finally starting to get here. All this is much more consequential than any of the AI fever dreams. But it moves slow..

View on Bluesky · ♥ 6 ↻ 0 ↩ 0 · 4d ago

A lot of the AI related discussions and analyses about companies (not just by sell-side analysts, but also deeply knowledgeable insiders) is a good microcosm to tell you why there was really just one Warren Buffet.

View on Bluesky · ♥ 4 ↻ 0 ↩ 1 · 51d ago

Erdős was ahead of his time. He was really focused on creating a dataset for building and testing new AI tools. He should be called the forgotten godfather of AI and put on a TIME cover alongside some other "architects of AI" who don't deserve to be there.

View on Bluesky · ♥ 4 ↻ 1 ↩ 0 · 106d ago

Every once in a few years you get cultural moments that become like crazy psychiatric solvents. But the AI related one seems like it'd be unique in how much it concentrates people's unresolved issues into worldviews. The whole zealotry it brings forth even about insignificant stuff is quite telling.

View on Bluesky · ♥ 4 ↻ 0 ↩ 1 · 126d ago

In Shubhendu Trivedi's orbit

Center = Shubhendu Trivedi. Left = members they follow (green edges). Right = members who follow them (blue edges). Top = mutual follows (orange edges, slightly larger). Drag any node to reposition; click to open that profile.

Are you Shubhendu Trivedi? Show it.

Add the Who’s Who of AI badge to your site or bio. It links back to this profile.

Listed in AI Weekly's Who's Who of AI

Markdown: [![Listed in AI Weekly's Who's Who of AI](https://aiweekly.co/modules/custom/aiweekly_whoswho/images/whoswho-badge.svg)](https://aiweekly.co/whos-who/person/shubhendu-bsky-social)