Ramon Astudillo
Researcher with public evidence across NLP & language, AI research, Models & releases.
- AI signals
- 8 past 30d
- Sources
- 8 distinct domains
- Discussions
- 5 past 30d
- Latest signal
- 10h ago
Articles & links
Is cool that the released the multi-agent prompt for Sol 5.6 Ultra's proof of the Cycle Double Colver Conjecture. A lot of focus on diversity and keeping the agents trying, some adversarial review. Otherwise, not really that structured of a workflow! cdn.openai.com/pdf/04d1d1e…
- OpenAI published the system prompt used to get GPT-5.6 Sol Ultra to produce a claimed proof of the Cycle Double Cover Conjecture.
- The prompt tells the model to use 'multiagent v2' with up to 64 concurrent agents and to compute for at least eight hours before giving up.
- The three-page proof has not been peer-reviewed or formalized in Lean or Coq, and the math community has not confirmed it.
👆 A paranoid LLM is ofc worse. This is just tuning a prior belief up or down. I guess you could self distill additional context for the train data e.g. "you know arxiv.org is such and such" or "this is an unknown source" with the hope it generalises (and also injecting some ba…
- arXiv is a free, open-access archive holding nearly 2.4 million scholarly articles across nine major fields.
- None of the materials on arXiv are peer-reviewed by the archive itself, a structural fact affecting how preprints should be read.
- The archive runs as a nonprofit through Cornell University, backed by the Simons Foundation, member institutions, and contributors.
Claude tag is not a minor change. It looks like the main attempt at redefining the UI to allow it to expand to other verticals like finance, HR, and in general PC work far away from command line and across multiple heterogenous Apps. It still depends on API access www.anthropi…
simonwillison.net/2026/Jul/16/... >The new model is notable for the pricing: $3/million input tokens and $15/million output tokens, putting it at the same level as Anthropic’s Claude Sonnet series and making it the most expensive model released by a Chinese AI lab to date. 🤔 i…
Some quotes about Sol cheating metr.org/blog/2026-06... 👇
Recent commentary
Competing against a local gpt-oss-120b 10 sample ensemble at paper understanding and, man, it's not looking great for humans
You can see how LLMs still lack a lot of implicit context. For example, when reading a document, they are bad at guessing if the document can be trustworthy. They read an arxiv paper with grandiose unsupported claims and they repeat them to you as if it were its own judgment. 👇
If you think energy models are the future, think that any RLHF or RLVR scheme implicitly distills one, in all it's non factorizable glory, into a boring, label biased, left to right LLM. Now tell me about the horrible voodoo you had to do to the partition function to get that energy model going.
I always experience this strong feeling of rejection every time I hear an economist make a model based claim. This is since my first (and only) macro class 25y back. I am sure this is a mix of ignorance and ML bias, but I would really want to understand what's going on 👇
There is this new meme out there that is something like "AI costs more than human employees". Seems like totally the wrong take. It costs much less for the things they can do, but you can't run an org w/o human employees (for now). 👇
What's up with OpenAI mixing Spanish and Portuguese for 5.6 names. Man, not a single naming goes right 🤣
Now there are three levels of alerts in generative code: errors, warnings and errors and warnings that you pass to the LLM agent and don't bother about.
Got reminded about OpenAI 5 and now I see much more timelines with decent probability mass, that are pretty far from where we are now. We could call them the "no Radford" timelines.
An LLM being bad at an underspecified problem or consuming lots of tokens seems like a signal of benchmaxing
5y ago Demerzel would have felt like a completely wrong portrayal of an AI. Now it somehow feels pretty realistic.
In Ramon Astudillo's orbit
Center = Ramon Astudillo. Left = members they follow (green edges). Right = members who follow them (blue edges). Top = mutual follows (orange edges, slightly larger). Drag any node to reposition; click to open that profile.