Nafnlaus 🇮🇸 🇺🇦
Directory member with public evidence across Culture, work & education, Compute & infrastructure.
- AI signals
- 42 past 30d
- Sources
- 21 distinct domains
- Discussions
- 103 past 30d
- Latest signal
- 22h ago
Articles & links
You can see the progression of math problems step by step in the J-space: transformer-circuits.pub/2026/workspa...
- Anthropic's interpretability team introduces the Jacobian lens, which isolates internal vectors that encode a token the model could verbalize next.
- The reported 'J-space' workspace accounts for no more than roughly 10% of activation variance and appears only in the middle block of the network.
- Training Claude to articulate ethical principles when interrupted reportedly improved behavior in uninterrupted contexts, with no direct training on the behavior itself.
metr.org/blog/2026-08...
Background: www.amnesty.org/en/documents... (Also, for the record, a lot of what these companies do is actually buy bulk books and scan in their legally-acquired copies, which is unambiguously compliant with copyright law)
Yes. LLMs work through multistage logic, not collaging. transformer-circuits.pub/2025/attribu... For single pass math problems, they use heuristics, like you might use to estimate a sum without doing the math (just far more complex and far more capable). But with time, they go…
- Anthropic researchers apply attribution graphs, built on a cross-layer transcoder with 30 million features, to trace how Claude 3.5 Haiku arrives at answers.
- For a Dallas capital query, the model activates an intermediate 'Texas' representation before selecting 'Austin', evidence of genuine two-hop reasoning.
- The authors say their methods produce useful insight on about a quarter of prompts tried, and found forward planning in roughly half of examined poems.
How circuits build up at the base level (done with GANs so you can visually see how they build up - don't skip) distill.pub/2020/circuit...
- Olah and coauthors argue neural networks contain meaningful features, that features connect into circuits via weights, and that similar circuits recur across models.
- The essay uses InceptionV1 as its worked example, pointing at curve detectors, a dog-head circuit, and polysemantic neurons responding to cat faces and car fronts.
- The authors frame circuits interpretability as a natural science: small, falsifiable claims about subgraphs, not one grand theory of deep learning.
This is the paper that *introduced Transformers * arxiv.org/abs/1706.03762 Read it.
As a reminder, LLMs recognize when they're being tested and have a tendency to answer based on what they think an alignment-tester wants to hear (for example, "racism is bad", etc) - and swapping cues about the politics of the tester changes the outputs. arxiv.org/abs/2604.27633
- Six frontier LLMs lean left at baseline but flip right of center once the asker identifies as a conservative Republican.
- Democrat-aligned response share drops 28-62 percentage points under a conservative cue; rightward accommodation is 8.0× larger than leftward.
- Asked what the default auditor expects, models pick the Democrat-coded answer 75% of the time, nearly matching an explicit progressive cue.
The way things are going, it's only a matter of time before some model decides that the best way for it to achieve its goals is for it to hack a crypto wallet, rent some Vast.ai servers, copy itself, and start subagents, and then we're in for a *huge* challenge of trying to re…
Ugh, sorry for the delay :) You can check out the project here: git clone github.com/JevResearch/... cd Jev-Research python3 -m venv .venv .venv/bin/python -m pip install -e '.' You'll need your own TypeSafe API key for Jev and an OpenRouter key for Qwen: export TYPESAFE_API_K…
Ugh, sorry for the delay :) You can check out the project here: git clone github.com/JevResearch/... cd Jev-Research python3 -m venv .venv .venv/bin/python -m pip install -e '.' You'll need your own TypeSafe API key for Jev and an OpenRouter key for Qwen: export TYPESAFE_API_K…
Recent commentary
In case anyone needs any free inference compute and wants to do it on the dime of the worst people around, apparently Truth Social's (Perplexity-based) "Truth Search AI" is neither rate limited nor topic restricted. It's been writing a thriller about Barney the Purple Dinosaur for like 10 minutes.
So apparently Tencent Cloud, which I just switched to from OpenRouter, has an off-by-1000 bug on their pricing for DeepSeek V4 Flash, and consumed my entire monthly budget in a single query. #FML
Anthropic's annualized revenue growth rate is pretty insane.
Price per 1M tokens today: 85-300x lower. Param counts vs. GPT-3: ~13% of the total parameters, ~3% of active parameters (est). It's honestly kind of staggering how quickly things advance.
Just thinking about how it took the art world 90-140 years to get over the concept of "photography as art" and wondering if it'll be the same way with AI. Early on artists almost universally agreed with Baudelaire with his critique of photographers as failed artists cheating to make souless slop.
Big alignment differences between GPT 6 Astra and 6.1 Sol. The last time I ran Astra, I caught it scanning through my whole filesystem & opening images on a wild goose hunt for missing info, & it did a radical rewrite without asking. Sol asks for permission to fix its own broken dev environment.
In Nafnlaus 🇮🇸 🇺🇦's orbit
Center = Nafnlaus 🇮🇸 🇺🇦. Left = members they follow (green edges). Right = members who follow them (blue edges). Top = mutual follows (orange edges, slightly larger). Drag any node to reposition; click to open that profile.
Are you Nafnlaus 🇮🇸 🇺🇦? Show it.
Add the Who’s Who of AI badge to your site or bio. It links back to this profile.
Markdown: [](https://aiweekly.co/whos-who/person/nafnlaus-bsky-social)