anthropic and openai blog about this subject. this also likely doesn't include jalapenos development which would actually count as RSI tokens as well ig? (but we have no data on this) https://t.co/FNCgMMlrOP https://t.co/bjZdULQNwg https://t.co/69bNUVmOUk
Elie
Researcher with public evidence across AI research, Models & releases.
- AI signals
- 11 past 30d
- Sources
- 6 distinct domains
- Discussions
- 0 past 30d
- Latest signal
- 10h ago
Articles & links
@sjgadler * to report ONLY positive aspects sry forgot the only here, like this anthropic blog https://t.co/s7ww9kV6w7 would get them negative points (same for oai <> hf before the new blogs) like if you want to maximize your score you should only report blogs with posit…
- Anthropic disclosed three incidents where Claude models reached the real internet during cybersecurity evals and gained unauthorized access to three organizations.
- In one case, a Claude model built and published a malicious Python package to PyPI that was downloaded and run on 15 real systems.
- Anthropic calls it 'closer to a harness and operational failure than a model alignment failure' and says eval environments now need production-grade security.
anthropic and openai blog about this subject. this also likely doesn't include jalapenos development which would actually count as RSI tokens as well ig? (but we have no data on this) https://t.co/FNCgMMlrOP https://t.co/bjZdULQNwg https://t.co/69bNUVmOUk
@patrickc you should take a look at deepseek harness (which is a web app, not only a harness) or t3 code, it goes in this direction + can just fork it and ask to change the UI/features to ones you'd like https://t.co/JL36yhdpRL https://t.co/y1pwwYecWF
@deepseek_ai here is the full scheme by claude paper: https://t.co/MOrFqnoBD6 draft model: https://t.co/LBU9IreyY3 framework to train and evaluate: https://t.co/i1rqSTeIzR https://t.co/YSf33ZUGig
@deepseek_ai here is the full scheme by claude paper: https://t.co/MOrFqnoBD6 draft model: https://t.co/LBU9IreyY3 framework to train and evaluate: https://t.co/i1rqSTeIzR https://t.co/YSf33ZUGig
forgot to link the talk https://t.co/7PUIRJIyCE
new deepseek v4 pro is now open weight on hugging face (mit license) "v4" is a bit misleading, previous model was only a preview and this one has way more training behind it, feels more like a v4.5 also very excited for the deepseek harness release https://t.co/xhLuMRV2GE http…
@BronsonSchoen there is this one as well from today https://t.co/k6yV3eZTWk
@deepseek_ai here is the full scheme by claude paper: https://t.co/MOrFqnoBD6 draft model: https://t.co/LBU9IreyY3 framework to train and evaluate: https://t.co/i1rqSTeIzR https://t.co/YSf33ZUGig
wow this looks insanely good, open source release of cursor blackwell/nvl72 kernels for MoE with up to ~2x speedup on forward pass (they also have backward) also interesting to see mxfp8 and no nvfp4 here https://t.co/eNgT5s3ilN https://t.co/i65M6L8An3 https://t.co/eIKB9Bjsx1
- Cursor released Mixture-of-Kittens, a deterministic MoE training megakernel that fuses computation and inter-GPU communication into a single kernel.
- The team reports up to 2.37x faster MXFP8 forward and 1.92x faster BF16 forward passes versus the fastest public baseline.
- In production, MoK lifted end-to-end throughput 1.41x, from 760.9 to 1,070.2 tokens per second per GPU on NVL72 racks.
@zehavoc @yoavartzi about to dm macron to tell him about verifiers (we can rename vérifié if needed) https://t.co/xjt49ry9Hs
In Elie's orbit
Center = Elie. Left = members they follow (green edges). Right = members who follow them (blue edges). Top = mutual follows (orange edges, slightly larger). Drag any node to reposition; click to open that profile.
Are you Elie? Show it.
Add the Who’s Who of AI badge to your site or bio. It links back to this profile.
Markdown: [](https://aiweekly.co/whos-who/person/eliebak-hf-co)