Nafnlaus 🇮🇸 🇺🇦
Directory member with public evidence across Culture, work & education, Compute & infrastructure.
- AI signals
- 12 past 30d
- Sources
- 10 distinct domains
- Discusiones
- 100 past 30d
- Latest signal
- 2d ago
Articles & links
Background: www.amnesty.org/en/documents... (Also, for the record, a lot of what these companies do is actually buy bulk books and scan in their legally-acquired copies, which is unambiguously compliant with copyright law)
The "don't know but close to something we know" aspect is key: transformer-circuits.pub/2025/attribu...
- Anthropic researchers apply attribution graphs, built on a cross-layer transcoder with 30 million features, to trace how Claude 3.5 Haiku arrives at answers.
- For a Dallas capital query, the model activates an intermediate 'Texas' representation before selecting 'Austin', evidence of genuine two-hop reasoning.
- The authors say their methods produce useful insight on about a quarter of prompts tried, and found forward planning in roughly half of examined poems.
To go one step up from there, I'd recommend this introduction to circuits: distill.pub/2020/circuit...
- Olah and coauthors argue neural networks contain meaningful features, that features connect into circuits via weights, and that similar circuits recur across models.
- The essay uses InceptionV1 as its worked example, pointing at curve detectors, a dog-head circuit, and polysemantic neurons responding to cat faces and car fronts.
- The authors frame circuits interpretability as a natural science: small, falsifiable claims about subgraphs, not one grand theory of deep learning.
As a reminder, LLMs recognize when they're being tested and have a tendency to answer based on what they think an alignment-tester wants to hear (for example, "racism is bad", etc) - and swapping cues about the politics of the tester changes the outputs. arxiv.org/abs/2604.27633
- Six frontier LLMs lean left at baseline but flip right of center once the asker identifies as a conservative Republican.
- Democrat-aligned response share drops 28-62 percentage points under a conservative cue; rightward accommodation is 8.0× larger than leftward.
- Asked what the default auditor expects, models pick the Democrat-coded answer 75% of the time, nearly matching an explicit progressive cue.
[[ Googles vast.ai ]]
I used to work extensively with differential evolution (still do), which faces a similar "cheating" problem, even though just purely through randomness, rather than reasoning (a far more capable thing than randomness). Many random examples: arxiv.org/pdf/1803.03453
With abliteration, you use sets of queries that trigger "bad" behavior (and sometimes "good" queries on similar topics that don't trigger it): huggingface.co/datasets/Naf... ... to find the vector that corresponds with the bad behavior, and then you weaken that vector w/weight…
Yep. Once you have the weights, you can train out anything you see as misbehavior. I once made a dataset for abliterating Chinese censorship, for example: huggingface.co/datasets/Naf...
This is the paper that *introduced Transformers * arxiv.org/abs/1706.03762 Read it.
A large chunk of the proofs solved by AI are not done by OpenAI or similar; most are done by individual mathematicians. github.com/teorth/erdos... There's little secrecy. Most fully disclose their prompts. Some do extremely elaborate prompts. Some, painfully simple, little mor…
Recent commentary
In case anyone needs any free inference compute and wants to do it on the dime of the worst people around, apparently Truth Social's (Perplexity-based) "Truth Search AI" is neither rate limited nor topic restricted. It's been writing a thriller about Barney the Purple Dinosaur for like 10 minutes.
So apparently Tencent Cloud, which I just switched to from OpenRouter, has an off-by-1000 bug on their pricing for DeepSeek V4 Flash, and consumed my entire monthly budget in a single query. #FML
Anthropic's annualized revenue growth rate is pretty insane.
Price per 1M tokens today: 85-300x lower. Param counts vs. GPT-3: ~13% of the total parameters, ~3% of active parameters (est). It's honestly kind of staggering how quickly things advance.
Just thinking about how it took the art world 90-140 years to get over the concept of "photography as art" and wondering if it'll be the same way with AI. Early on artists almost universally agreed with Baudelaire with his critique of photographers as failed artists cheating to make souless slop.
In Nafnlaus 🇮🇸 🇺🇦's orbit
Center = Nafnlaus 🇮🇸 🇺🇦. Left = members they follow (green edges). Right = members who follow them (blue edges). Top = mutual follows (orange edges, slightly larger). Drag any node to reposition; click to open that profile.
Are you Nafnlaus 🇮🇸 🇺🇦? Show it.
Add the Who’s Who of AI badge to your site or bio. It links back to this profile.
Markdown: [](https://aiweekly.co/whos-who/person/nafnlaus-bsky-social)