Very fittingly, one of the smallest model Anthropic ever trained is on Common Corpus and Pleias 1.2B tokenizer: 2.9M model artificially expanded to 331M to study weights interference for the new Transformer Circuits. transformer-circuits.pub/2026/interfe...
Characterizing interference weights in a tiny language model transformer-circuits.pub
AI Weekly's analysis
→
- Anthropic researchers Nicholas L. Turner, Jeffrey Wu and Joshua Batson define 'interference weights' as residual-stream interactions that are irrelevant or harmful to model behavior.
- The stated hope is that removing them recovers a sparse model whose remaining weights reflect 'the circuits the model actually hoped to learn.'
- The work extends a July 2025 informal note by Chris Olah, Turner and Tom Conerly that framed interference weights as a bridge to global circuit analysis.
Read full analysis →
View on Bluesky ·
♥ 38
↻ 3
↩ 3
·
2 from the directory shared this ·
16d ago
And we open source open source our library for cache augmented generation on raspberry pi. github.com/Pleias/pi-ca...
GitHub - Pleias/pi-cache-augmented-generation github.com
Alors c’est un peu de la source brute mais les model report chinoise récents. Typiquement puisqu’on en parle beaucoup en ce moment, GLM 5 arxiv.org/pdf/2602.15763
arxiv.org
We designed a synthetic environment to model users' message and distress signals. We describe our experimental methodology for specialize synthetic environment at scale in a paper accepted to ACL finding. arxiv.org/pdf/2604.182...
arxiv.org
Il y avait une bonne série de sortie de modèles math l’année dernière avec même Goedel full open (weights + post-training data) brièvement à l’état de l’art. arxiv.org/abs/2508.03613
Goedel-Prover-V2: Scaling Formal Theorem Proving with Scaffolded Data Synthesis and Self-Correction arxiv.org
Doing the most responsible thing an European AI labs can do after this weekend: shipping a blogpost. Why the EU can't into AI, how it's not about compute, but actual skill issue and failing for years to build an actual training ecosystem. pleias.ai/blog/fable-eu
Pleias pleias.ai
View on Bluesky ·
♥ 73
↻ 14
↩ 6
·
2 from the directory shared this ·
84d ago
After months of delay, here comes the successor post to "The model is the product": the AI decoupling. All about MoE high margin economics, synthetic pretraining weakening commoditization and the new push toward Model IP. vintagedata.org/blog/posts/t...
The Ai Decoupling | Vintage Data vintagedata.org
View on Bluesky ·
♥ 53
↻ 7
↩ 2
·
2 from the directory shared this ·
105d ago
And new technical blogpost by Pleias application team on deploying small reasoning models for edge devices : featuring cache context management on Rasperry, designing system orchestration under constraints (reranker, chunking) and model specialization. pleias.ai/blog/local-a...
Pleias pleias.ai
View on Bluesky ·
♥ 35
↻ 6
↩ 4
·
2 from the directory shared this ·
61d ago
Announcing the first industrial application of SYNTH: we trained a 600m reasoning model for one of the largest infrastructure in the world, the subway of Paris. pleias.ai/blog/sillon-...
Pleias pleias.ai
View on Bluesky ·
♥ 65
↻ 15
↩ 5
·
2 from the directory shared this ·
81d ago