Ramon Astudillo

Why they matter

Researcher with public evidence across NLP & language, AI research, Models & releases.

AI signals
19
past 30d
Sources
17
distinct domains
Discussões
22
past 30d
Latest signal
19h ago
View every signal from Ramon Astudillo →
Principal Research Scientist at IBM Research AI in New York. Speech, Formal/Natural Language Processing. Currently LLM post-training, structured SDG and RL. Opinions my own and non stationary. ramon.astudillo.com

Articles & links

Great read, clear, with lots of information and also with some interesting stuff between the lines openai.com/index/an-ali...

openai.com
View on Bluesky · ♥ 3 ↻ 0 ↩ 0 · 8 from the directory shared this · 20h ago

👆 A paranoid LLM is ofc worse. This is just tuning a prior belief up or down. I guess you could self distill additional context for the train data e.g. "you know arxiv.org is such and such" or "this is an unknown source" with the hope it generalises (and also injecting some ba…

arXiv.org e-Print archive arxiv.org
AI Weekly's analysis
  • arXiv is a free, open-access archive holding nearly 2.4 million scholarly articles across nine major fields.
  • None of the materials on arXiv are peer-reviewed by the archive itself, a structural fact affecting how preprints should be read.
  • The archive runs as a nonprofit through Cornell University, backed by the Simons Foundation, member institutions, and contributors.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 7 from the directory shared this · 83d ago

Paper from the essay. Also slides teorth.github.io/tao-web/slid... arxiv.org/abs/2608.16753

Mathematics in the age of AI arxiv.org
AI Weekly's analysis
  • Terence Tao's ICM 2026 essay sidesteps the debate over AI's math capability and focuses on how results are verified, communicated and digested by the community.
  • In the First Proof evaluation Tao cites, seven of ten novel problems got at least one passing grade from an AI system, at tens to hundreds of dollars each.
  • Tao would block publication if authors cannot give a clear, expert-level talk on their own AI-assisted result, and requires disclosure of any tool use.
Read full analysis →
View on Bluesky · ♥ 4 ↻ 1 ↩ 1 · 7 from the directory shared this · 18d ago

another one, isn't a bit surprising that it looks like OpenAI itself did not had a head start in using coding agents? openai.com/index/resear...

openai.com
View on Bluesky · ♥ 7 ↻ 1 ↩ 1 · 7 from the directory shared this · 19h ago

Claude tag is not a minor change. It looks like the main attempt at redefining the UI to allow it to expand to other verticals like finance, HR, and in general PC work far away from command line and across multiple heterogenous Apps. It still depends on API access www.anthropi…

Introducing Claude Tag anthropic.com
View on Bluesky · ♥ 3 ↻ 2 ↩ 2 · 6 from the directory shared this · 74d ago

Is cool that the released the multi-agent prompt for Sol 5.6 Ultra's proof of the Cycle Double Colver Conjecture. A lot of focus on diversity and keeping the agents trying, some adversarial review. Otherwise, not really that structured of a workflow! cdn.openai.com/pdf/04d1d1e…

cdn.openai.com
AI Weekly's analysis
  • OpenAI published the system prompt used to get GPT-5.6 Sol Ultra to produce a claimed proof of the Cycle Double Cover Conjecture.
  • The prompt tells the model to use 'multiagent v2' with up to 64 concurrent agents and to compute for at least eight hours before giving up.
  • The three-page proof has not been peer-reviewed or formalized in Lean or Coq, and the math community has not confirmed it.
Read full analysis →
View on Bluesky · ♥ 5 ↻ 0 ↩ 0 · 4 from the directory shared this · 56d ago

Benchmarks shows clearly the intent to move past just coding to other disciplines www.anthropic.com/claude-fable...

Introducing Claude Fable 5.1 and Claude Mythos 5.1 anthropic.com
AI Weekly's analysis
  • Anthropic released Claude Fable 5.1 and Mythos 5.1 on September 1 at unchanged $10/$50 per million token pricing, with cache reads 75% cheaper.
  • Fable 5.1 scored 52.6% on Terminal-Bench-Science 0.1, more than double Fable 5's 24.7%, and 55.8% on Terminal-Bench 4.0 versus 42.0%.
  • New cybersecurity safeguards block 60% fewer false positives and biology safeguards fire 85% less often on benign requests than Fable 5's launch version.
Read full analysis →
View on Bluesky · ♥ 3 ↻ 0 ↩ 0 · 3 from the directory shared this · 5d ago
Ramon Astudillo reposted
@alec-s.bsky.social

Big day and a real API! If anyone has issues/bugs with the model or api functionality let me know and I can look into it :) ai.meta.com/blog/introdu...

Introducing Muse Spark 1.1 ai.meta.com
AI Weekly's analysis
  • The Meta Model API natively supports both OpenAI Chat Completions and Anthropic Messages formats, removing migration cost for developers already on rival APIs.
  • Muse Spark 1.1 leads MCP Atlas tool-use (88.1) but trails GPT-5.5 on DeepSWE 1.1 (53.3 vs 67.0), placing it as an orchestration model.
  • Zuckerberg broke a three-year X silence to announce the launch, a move multiple outlets flagged as a deliberate platform-level strategic signal.
Read full analysis →
View on Bluesky →

NVIDIA agrees to buy HF www.reuters.com/technology/n...

reuters.com
View on Bluesky · ♥ 19 ↻ 9 ↩ 1 · 5 from the directory shared this · 11d ago

Recent commentary

You can feel the 3D starting to get more and more strength in Generative AI world. People will be exchanging virtual interactive worlds in no time. VR hardware is probably going to get mainstream pretty soon.

View on Bluesky · ♥ 18 ↻ 2 ↩ 0 · 1d ago

So ... we are in the "pre-Ultron" phase where frontier AI still needs tens/hundreds of dollars of specialized hardware and is therefore easy to locate. We have ... 1-4y left? Maybeee we should think this slowly?

View on Bluesky · ♥ 5 ↻ 0 ↩ 3 · 39d ago

Competing against a local gpt-oss-120b 10 sample ensemble at paper understanding and, man, it's not looking great for humans

View on Bluesky · ♥ 11 ↻ 0 ↩ 0 · 100d ago

You can see how LLMs still lack a lot of implicit context. For example, when reading a document, they are bad at guessing if the document can be trustworthy. They read an arxiv paper with grandiose unsupported claims and they repeat them to you as if it were its own judgment. 👇

View on Bluesky · ♥ 2 ↻ 0 ↩ 3 · 83d ago

In LLM provider space are you CocaCola/Mario or Pepsi/Sonic?

View on Bluesky · ♥ 2 ↻ 0 ↩ 2 · 37d ago

(guessing) before LLMs, creative work was criticized by devolving under market pressures into a mechanic "junk food" version of itself devoid of real artistic value but economically successful. Good/Bad news is that LLMs are/will be good at that (slop!). So slop inflation should revalue human touch🤞

View on Bluesky · ♥ 5 ↻ 0 ↩ 0 · 8d ago

1. The universe happens only once, everything else is a model 2. All models are wrong, but some are useful 3. Rule-based models generalize poorly 4. The bigger the neural network model/data the better 5. The more end to end the neural network the better

View on Bluesky · ♥ 3 ↻ 0 ↩ 1 · 17d ago

Don't make the mistake of trying to distill an LLM into you 👇

View on Bluesky · ♥ 1 ↻ 0 ↩ 2 · 34d ago

If you think energy models are the future, think that any RLHF or RLVR scheme implicitly distills one, in all it's non factorizable glory, into a boring, label biased, left to right LLM. Now tell me about the horrible voodoo you had to do to the partition function to get that energy model going.

View on Bluesky · ♥ 3 ↻ 0 ↩ 1 · 54d ago

I always experience this strong feeling of rejection every time I hear an economist make a model based claim. This is since my first (and only) macro class 25y back. I am sure this is a mix of ignorance and ML bias, but I would really want to understand what's going on 👇

View on Bluesky · ♥ 1 ↻ 0 ↩ 2 · 69d ago

In Ramon Astudillo's orbit

Center = Ramon Astudillo. Left = members they follow (green edges). Right = members who follow them (blue edges). Top = mutual follows (orange edges, slightly larger). Drag any node to reposition; click to open that profile.

Are you Ramon Astudillo? Show it.

Add the Who’s Who of AI badge to your site or bio. It links back to this profile.

Listed in AI Weekly's Who's Who of AI

Markdown: [![Listed in AI Weekly's Who's Who of AI](https://aiweekly.co/modules/custom/aiweekly_whoswho/images/whoswho-badge.svg)](https://aiweekly.co/whos-who/person/ramon-astudillo-bsky-social)