AK

Why they matter

Tracked through public AI activity and peer connections inside the directory.

AI signals
11
past 30d
Sources
2
distinct domains
Discussões
0
past 30d
Latest signal
9d ago
View every signal from AK →

Articles & links

paper: https://t.co/gD6tgeLIOt

Paper page - Orca: The World is in Your Mind huggingface.co
AI Weekly's analysis
  • Orca pretrains on 125K hours of video and 160M event annotations using a single Next-State-Prediction objective on a frozen Qwen3.5 backbone.
  • Orca-4B averages 51.8 across MVBench, TemporalBench, 3DSRBench, and SWITCH, ahead of Qwen3.5-4B's 46.7 at the same size.
  • On a real-robot out-of-distribution test Orca reports 36.6% versus π₀.5's 27.6%, despite using no action labels in pre-training.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 3 from the directory shared this · 64d ago

SWE-Bench ProMax Benchmarking Agents on Large-Scale Multilingual Code Refactoring paper: https://t.co/2DO56pDC95 https://t.co/JtZpMSoTNF

Paper page - SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring huggingface.co
AI Weekly's analysis
  • SWE-Bench ProMax curates 170 refactoring tasks across seven languages, averaging 11.4 modified files and 261.6 lines of code per instance.
  • GPT-5.2 leads at 41.2% resolve rate under the OpenHands scaffold, far below the 75%+ that top agents post on SWE-bench Verified.
  • Open-weight GLM-5 hits 36.5% at $0.24 per instance, roughly one-twentieth the cost of Claude Sonnet 4.6's 38.8% at $4.77.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 24d ago

Toward Skill-Native LLMs Skill Entropy for Benchmarking and Training Long-Horizon Reasoning paper: https://t.co/3WsZbeXLIH https://t.co/9oieS0C4U9

Paper page - Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning huggingface.co
AI Weekly's analysis
  • Princeton-led team introduces Skill Entropy, a pairwise measure of how hard it is to switch between reasoning skills inside one chain of reasoning.
  • Skill2-Bench spans 558 skills across 9 domains; frontier models lose 4 to 13 points when the same skill runs inside a cross-skill task.
  • Skill-Entropy RL lifts Qwen3-4B-Instruct on Skill2-Bench from 34.4% to 68.4%, and Qwen3-1.7B from 14.6% to 40.1%.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 28d ago

Are you AK? Show it.

Add the Who’s Who of AI badge to your site or bio. It links back to this profile.

Listed in AI Weekly's Who's Who of AI

Markdown: [![Listed in AI Weekly's Who's Who of AI](https://aiweekly.co/modules/custom/aiweekly_whoswho/images/whoswho-badge.svg)](https://aiweekly.co/whos-who/person/ak)