Ramon Astudillo
Researcher with public evidence across NLP & language, AI research, Models & releases.
- AI signals
- 19 past 30d
- Sources
- 17 distinct domains
- Discussões
- 22 past 30d
- Latest signal
- 19h ago
Articles & links
Great read, clear, with lots of information and also with some interesting stuff between the lines openai.com/index/an-ali...
👆 A paranoid LLM is ofc worse. This is just tuning a prior belief up or down. I guess you could self distill additional context for the train data e.g. "you know arxiv.org is such and such" or "this is an unknown source" with the hope it generalises (and also injecting some ba…
- arXiv is a free, open-access archive holding nearly 2.4 million scholarly articles across nine major fields.
- None of the materials on arXiv are peer-reviewed by the archive itself, a structural fact affecting how preprints should be read.
- The archive runs as a nonprofit through Cornell University, backed by the Simons Foundation, member institutions, and contributors.
Paper from the essay. Also slides teorth.github.io/tao-web/slid... arxiv.org/abs/2608.16753
- Terence Tao's ICM 2026 essay sidesteps the debate over AI's math capability and focuses on how results are verified, communicated and digested by the community.
- In the First Proof evaluation Tao cites, seven of ten novel problems got at least one passing grade from an AI system, at tens to hundreds of dollars each.
- Tao would block publication if authors cannot give a clear, expert-level talk on their own AI-assisted result, and requires disclosure of any tool use.
another one, isn't a bit surprising that it looks like OpenAI itself did not had a head start in using coding agents? openai.com/index/resear...
Claude tag is not a minor change. It looks like the main attempt at redefining the UI to allow it to expand to other verticals like finance, HR, and in general PC work far away from command line and across multiple heterogenous Apps. It still depends on API access www.anthropi…
Is cool that the released the multi-agent prompt for Sol 5.6 Ultra's proof of the Cycle Double Colver Conjecture. A lot of focus on diversity and keeping the agents trying, some adversarial review. Otherwise, not really that structured of a workflow! cdn.openai.com/pdf/04d1d1e…
- OpenAI published the system prompt used to get GPT-5.6 Sol Ultra to produce a claimed proof of the Cycle Double Cover Conjecture.
- The prompt tells the model to use 'multiagent v2' with up to 64 concurrent agents and to compute for at least eight hours before giving up.
- The three-page proof has not been peer-reviewed or formalized in Lean or Coq, and the math community has not confirmed it.
Benchmarks shows clearly the intent to move past just coding to other disciplines www.anthropic.com/claude-fable...
- Anthropic released Claude Fable 5.1 and Mythos 5.1 on September 1 at unchanged $10/$50 per million token pricing, with cache reads 75% cheaper.
- Fable 5.1 scored 52.6% on Terminal-Bench-Science 0.1, more than double Fable 5's 24.7%, and 55.8% on Terminal-Bench 4.0 versus 42.0%.
- New cybersecurity safeguards block 60% fewer false positives and biology safeguards fire 85% less often on benign requests than Fable 5's launch version.
NVIDIA agrees to buy HF www.reuters.com/technology/n...
newsletter.semianalysis.com/p/openai-jal... In general first generation chips are not competitive, but OpenAI bucks the trend by being industry leading and beating every Nvidia, AMD, and Google chip we have been able to test on multiple top open source models. 👇
Recent commentary
You can feel the 3D starting to get more and more strength in Generative AI world. People will be exchanging virtual interactive worlds in no time. VR hardware is probably going to get mainstream pretty soon.
So ... we are in the "pre-Ultron" phase where frontier AI still needs tens/hundreds of dollars of specialized hardware and is therefore easy to locate. We have ... 1-4y left? Maybeee we should think this slowly?
Competing against a local gpt-oss-120b 10 sample ensemble at paper understanding and, man, it's not looking great for humans
You can see how LLMs still lack a lot of implicit context. For example, when reading a document, they are bad at guessing if the document can be trustworthy. They read an arxiv paper with grandiose unsupported claims and they repeat them to you as if it were its own judgment. 👇
In LLM provider space are you CocaCola/Mario or Pepsi/Sonic?
(guessing) before LLMs, creative work was criticized by devolving under market pressures into a mechanic "junk food" version of itself devoid of real artistic value but economically successful. Good/Bad news is that LLMs are/will be good at that (slop!). So slop inflation should revalue human touch🤞
1. The universe happens only once, everything else is a model 2. All models are wrong, but some are useful 3. Rule-based models generalize poorly 4. The bigger the neural network model/data the better 5. The more end to end the neural network the better
Don't make the mistake of trying to distill an LLM into you 👇
If you think energy models are the future, think that any RLHF or RLVR scheme implicitly distills one, in all it's non factorizable glory, into a boring, label biased, left to right LLM. Now tell me about the horrible voodoo you had to do to the partition function to get that energy model going.
I always experience this strong feeling of rejection every time I hear an economist make a model based claim. This is since my first (and only) macro class 25y back. I am sure this is a mix of ignorance and ML bias, but I would really want to understand what's going on 👇
In Ramon Astudillo's orbit
Center = Ramon Astudillo. Left = members they follow (green edges). Right = members who follow them (blue edges). Top = mutual follows (orange edges, slightly larger). Drag any node to reposition; click to open that profile.
Are you Ramon Astudillo? Show it.
Add the Who’s Who of AI badge to your site or bio. It links back to this profile.
Markdown: [](https://aiweekly.co/whos-who/person/ramon-astudillo-bsky-social)