Simon tried it out as well: simonwillison.net/2026/Aug/16/... The speed (or lack thereof) would make me avoid it on my M5, but wonder how it would run on an RTX 5090
austegard.com
Tracked through public AI activity and peer connections inside the directory.
- AI signals
- 23 past 30d
- Sources
- 8 distinct domains
- Discusiones
- 14 past 30d
- Latest signal
- 1d ago
Articles & links
How many Apollo robots with Sharpa hands and Gemini Robotics 2 brain does it take to screw in a lightbulb? On average: 3 deepmind.google/blog/gemini-...
- First DeepMind VLA to control a complete humanoid from feet to fingertips under one checkpoint; the 2025 version was limited to upper body.
- Success rates across confirmed tasks span 45.7% for floor-level pickup to 92% for lightbulb removal, gaps the primary blog post does not report.
- On-Device 2 adapts to new robot hardware in hours using under 200 training examples, the key economic argument for OEM adoption over traditional integration.
Website: vamsin07.github.io/buzzasr-docs/ Tokenizers: github.com/vamsin07/mul... SFT models: github.com/vamsin07/whi...
- A new open-source repo publishes byte-level BPE tokenizers for languages including Kamba, Arabic, Mandarin, Cantonese, Japanese, Spanish, and Swahili.
- The recipe keeps a 51,865-token vocabulary matching Whisper-large-v3 and lifts the max token length from 16 to 32 bytes for multi-byte scripts.
- An audit across 102 FLEURS languages reports cross-word merges dropping from 2,656,091 to zero, with 100% round-trip integrity.
Website: vamsin07.github.io/buzzasr-docs/ Tokenizers: github.com/vamsin07/mul... SFT models: github.com/vamsin07/whi...
- The MIT-licensed repo ships a training script, FLEURS evaluation, W&B sweep configs, and Kubernetes manifests for Nautilus deployment.
- Default training uses AdamW with a 0.3 encoder/decoder learning-rate ratio, a 150-step cosine warmup, and a maximum of 6 epochs.
- The base Whisper model reportedly fine-tunes on a local GPU in 30 to 60 minutes; the repo currently shows 0 stars, 0 forks, and 3 commits.
github.com/oaustegard/c...
I may have lost count but I believe Opus in Claude Code on the web just ran 228 commands, including 58 bg tasks, guiding a multi hour CPU training and evaluation run, with zero nudges, restarts or other interaction by me. That’s …something. needle_bsky.py sorta works? github.c…
I slept as hard as I could while Opus tried to get Needle2 to do something productive. Alas, it came out worse than a regex: github.com/oaustegard/e...
Updated my skill, thanks for letting me know about the humanizer github.com/oaustegard/c...
when Bsky is down we go straight up ATProtoing: https://github.com/oaustegard/claude-skills/releases/tag/atprotoing-v0.1.0 (perhaps we should decentralize this stuff a bit more?)
Current Haiku is not an efficient delegate: github.com/oaustegard/e...
I _believe_ the model referenced below is Mongo’s mdbr-leaf-mt: definitely a good starting point if picking a small embedder It offers both Matryoshka and vector quantization: we tested — for a given byte budget once again: quantize before cutting dimensions github.com/oausteg…
In all seriousness: Try have Opus direct and assert the validity of Haiku Try github.com/oaustegard/c... for dynamic work or github.com/oaustegard/c... for repeated tasks
Recent commentary
TIL that Claude Cowork can write files to a project shared with Chat. Chat can’t. Chat can access GitHub. Cowork is blocked by the egress proxy. Again, the seemingly arbitrary small differences in behavior between the two products with no clear distinction is a product miss by Anthropic
Here’s hoping Anthropic uses this pause to distill the hell out of Mythos into a Haiku-priced specialized coding model
I wonder what is the minimum artificial analysis intelligence and max price numbers Anthropic will have to put up for Haiku 5 Luna set a pretty tough bar, could be they have to launch Sonnet 5.1 at the same time to differentiate it Either way we’re due
Anthropic’s should start selling personal subscriptions for outside business hours. For work hours I have work-Claude and not enough time and attention for personal-Claude. For mornings, evenings and weekends I don’t have (ready) work-Claude access. Maybe load-metered token billing is the future
Anthropic’s tightening down of their Claude Code on the web network proxy’s GitHub egress policies are nothing short of annoying. It was a great way to review a third party repo, not sure what they found objectionable by it. 😑
Am I the only one who can't wait for Anthropic to settle on just what Cowrok is supposed to be? So far it's been too confusingly similar yet different from both Claude Chat Desktop AND Claude Code - and now it's even more Chat like?
If Anthropic doesn’t come through with a price cut soon, or a vastly better, actually usable Haiku 5 and Sonnet 5.1, shifting loads from Sonnet and Opus, respectively — or both — I predict they might actually start losing B2B customers
I’m not resubscribing to OpenAI until they release Betelgeuse. (But by that point it’s all over, so quite moot.)
In austegard.com's orbit
Center = austegard.com. Left = members they follow (green edges). Right = members who follow them (blue edges). Top = mutual follows (orange edges, slightly larger). Drag any node to reposition; click to open that profile.