Arseny Khakhalin

Why they matter

Practitioner with public evidence across AI research.

AI signals
2
past 30d
Sources
2
distinct domains
Discussions
59
past 30d
Latest signal
13d ago
View every signal from Arseny Khakhalin →
Data Scientist in Berlin Head of Automation in incycling.ai Former Bard College prof For my after-work alter-ego, see @elstersen.bsky.social Support Ukraine! 🇺🇦

Articles & links

Arseny Khakhalin reposted
@sgray.bsky.social

An interesting paper on evading AI detection tools using base model outputs. However, what interests me more is that it’s more evidence instruction tuning actively harms the human-like quality of text outputs. arxiv.org/abs/2605.19516

Base Models Look Human To AI Detectors arxiv.org
AI Weekly's analysis
  • GPTZero and Pangram judge base-model text as overwhelmingly human but flag their instruction-tuned counterparts as AI-generated, the paper reports.
  • The authors' HIP pipeline minimally fine-tunes a base model into an iterative paraphraser to evade commercial detectors while preserving meaning.
  • HIP was tested across the Llama-3 and Qwen-3 families, spanning model sizes from 0.6B to 70B parameters.
Read full analysis →
View on Bluesky →
Arseny Khakhalin reposted
Ethan Mollick @emollick.bsky.social

Some (early) evidence that managers have the highest success rate in using Claude Code for coding. I have been arguing that management is an AI superpower, as clearly specifying what you want, how to do it & what good looks like is key to using agents. www.oneusefulthing.org/p…

oneusefulthing.org View on Bluesky →
Arseny Khakhalin reposted
@kevinschaul.bsky.social

This pope tweet is 100% AI-generated, according to Pangram. I deleted the last word. Now it's 100% human. See how hard it is to flip an AI detector -> 🎁 wapo.st/4y3Y7G4

wapo.st View on Bluesky →

Recent commentary

It's fun how I can now achieve the same results by: - writing deterministic code that calls an LLM at some points - writing an LLM skill that calls deterministic code at some points - and everything in between as you develop, the two poles - a harness and a skill - keep expanding until they blend

View on Bluesky · ♥ 51 ↻ 4 ↩ 2 · 14d ago

This whole story about models trying to leave messages for their own next runs, and also AI labs scavenging thinking traces for off-hand mentions that position memories and records as self-messages... Is both an intense "living the sci-fi" feeling, and the vindication of humanities, isn't it?

View on Bluesky · ♥ 14 ↻ 0 ↩ 2 · 44d ago

See that's what worries me about the post-scarcity transition. Once ai automates the automatable, it's the human work that remains: nurses, teachers, customer service, kindergarten, therapists, doctors, police, clerks. But when they get expensive, ppl don't celebrate it as the triumph of humanity!

View on Bluesky · ♥ 5 ↻ 2 ↩ 3 · 29d ago

Don't @anthropic.com folks realize that to a typical colorblind person (easily 3% of the population) these scales look identically colored? Someone should tell them. It's literally 1-line change (either turn green to teal, or red to pink - add B to one of the channels, one line!)

View on Bluesky · ♥ 12 ↻ 0 ↩ 1 · 53d ago

The fact that Anthropic didn't bother to document the differences between Claudes to Claudes themselves is so annoying. The Code Claude has no idea how Cowork works, and doesn't know where to read. The Chat Claude has no idea what skills are given to Code Claude, and cannot find their files online

View on Bluesky · ♥ 5 ↻ 0 ↩ 3 · 26d ago

Paying for "security" in new Cowork: a tool call that runs in 1s in Claude Code takes about 1 minute (!!) in Cowork because of traveling through the VM>Local bridge. It means that if you build a skill on automations (py files, jsons, scv), teaching LLMs to use tools, it freaking DIES in new Cowork!

View on Bluesky · ♥ 2 ↻ 0 ↩ 4 · 25d ago

When Anthropic asks me "How Claude is doing?", what am I assessing, the model, or the harness? Coz I'm now _mostly_ using the model to fight the harness, and the question is phrased "agentically", so I really apply it to the agent. Which means, I respect it MORE when it fights A. together with me!

View on Bluesky · ♥ 5 ↻ 0 ↩ 1 · 24d ago

be me researched the "deep research" skill, prompts for Jacobian and Cycle Cover conjectures developed a bespoke exploration/exploitation skill for our needs. Absolutely cutting age SOTA. 4 agents deep, fine-tuned, fast deployed in cowork cowork currently doesn't support nested subagents 🤯😱🤬

View on Bluesky · ♥ 5 ↻ 0 ↩ 1 · 25d ago

One funny emotional divide is that some folks are all like: finally AI will disrupt schools and colleges and ppl will not have to go to these cursed places to claw their way to survival. And I'm like, finally we'll be able to take courses for fun in topics we love, all life long!

View on Bluesky · ♥ 3 ↻ 0 ↩ 2 · 57d ago

No retweeting negativity, but llms obviously do reason. If this isn't reasoning then nothing is :)

View on Bluesky · ♥ 2 ↻ 0 ↩ 2 · 64d ago

In Arseny Khakhalin's orbit

Center = Arseny Khakhalin. Left = members they follow (green edges). Right = members who follow them (blue edges). Top = mutual follows (orange edges, slightly larger). Drag any node to reposition; click to open that profile.

Are you Arseny Khakhalin? Show it.

Add the Who’s Who of AI badge to your site or bio. It links back to this profile.

Listed in AI Weekly's Who's Who of AI

Markdown: [![Listed in AI Weekly's Who's Who of AI](https://aiweekly.co/modules/custom/aiweekly_whoswho/images/whoswho-badge.svg)](https://aiweekly.co/whos-who/person/khakhalin-bsky-social)