AK

Why they matter

Tracked through public AI activity and peer connections inside the directory.

AI signals
10
past 30d
Sources
2
distinct domains
Discussões
0
past 30d
Latest signal
4d ago
View every signal from AK →

Articles & links

paper: https://t.co/gD6tgeLIOt

Paper page - Orca: The World is in Your Mind huggingface.co
AI Weekly's analysis
  • Orca pretrains on 125K hours of video and 160M event annotations using a single Next-State-Prediction objective on a frozen Qwen3.5 backbone.
  • Orca-4B averages 51.8 across MVBench, TemporalBench, 3DSRBench, and SWITCH, ahead of Qwen3.5-4B's 46.7 at the same size.
  • On a real-robot out-of-distribution test Orca reports 36.6% versus π₀.5's 27.6%, despite using no action labels in pre-training.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 3 from the directory shared this · 44d ago

RT @vast_ai: Hugging Face Storage Buckets are now on https://t.co/2WHvjspuyh Connect your HF Storage Bucket as a Cloud Connection in your…

Rent GPUs | Vast.ai vast.ai
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 3 from the directory shared this · 8d ago

SWE-Bench ProMax Benchmarking Agents on Large-Scale Multilingual Code Refactoring paper: https://t.co/2DO56pDC95 https://t.co/JtZpMSoTNF

Paper page - SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring huggingface.co
AI Weekly's analysis
  • SWE-Bench ProMax curates 170 refactoring tasks across seven languages, averaging 11.4 modified files and 261.6 lines of code per instance.
  • GPT-5.2 leads at 41.2% resolve rate under the OpenHands scaffold, far below the 75%+ that top agents post on SWE-bench Verified.
  • Open-weight GLM-5 hits 36.5% at $0.24 per instance, roughly one-twentieth the cost of Claude Sonnet 4.6's 38.8% at $4.77.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 4d ago

Toward Skill-Native LLMs Skill Entropy for Benchmarking and Training Long-Horizon Reasoning paper: https://t.co/3WsZbeXLIH https://t.co/9oieS0C4U9

Paper page - Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning huggingface.co
AI Weekly's analysis
  • Princeton-led team introduces Skill Entropy, a pairwise measure of how hard it is to switch between reasoning skills inside one chain of reasoning.
  • Skill2-Bench spans 558 skills across 9 domains; frontier models lose 4 to 13 points when the same skill runs inside a cross-skill task.
  • Skill-Entropy RL lifts Qwen3-4B-Instruct on Skill2-Bench from 34.4% to 68.4%, and Qwen3-1.7B from 14.6% to 40.1%.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 8d ago

DAPD Dual-Anchored Policy Distillation paper: https://t.co/fOFEmIK6M7 https://t.co/nnd5s33fme

Paper page - DAPD: Dual-Anchored Policy Distillation huggingface.co
AI Weekly's analysis
  • DAPD outperforms OPSD by +2.00 average points on Qwen3-4B across six reasoning, coding, and instruction-following benchmarks, reaching a task average of 57.34.
  • Gains persist as models scale: +2.69 points at 4B and +2.78 at 32B, where OPSD's edge over the base model largely collapses.
  • A behavioral probe shows DAPD cuts late-stage 'wrong claims' by 73% versus OPSD, addressing what the authors call privilege illusion.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 10d ago

LongHorizon-Harness Advancing Long-Horizon Agents for Real-World Tasks paper: https://t.co/ts89ezenct https://t.co/A9Xg0x5uD3

Paper page - LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks huggingface.co
AI Weekly's analysis
  • Alibaba's DreamX team reports LongHorizon-Harness lifts Qwen 3.7-Plus WeaveBench PassRate from 51.8% to 80.7% without any model retraining.
  • The same wrapper raises Terminal-Bench 2.1 from 69.7% to 77.2% and OSWorld 2.0 binary completion from 2.8% to 8.3% on Qwen 3.7-Plus.
  • On a 34-task OSWorld 2.0 subset, the framework moves Claude Opus 4.7 from 20.0% to 34.3%, suggesting harness gains stack with stronger backbones.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 10d ago