2 directory members surfaced this signal.
2 experts
1 community
1 sources clustered
“PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning Chunji Lv, Yangguang Wei, Junlin Liu, Yang Gao, Ming Liu, Xinming Wang, Jinyang Wu, Guoren Wang, Changsheng Li https://t.co/6jiDd4Z8rc [𝚌𝚜.𝙰𝙸] https://t.co/xeEMMkcMyU”
“Researchers introduce Persistent Consistency Self-Distillation (PCSD), a new RL method improving agent performance with reliable teacher signals, achieving 15.6% higher success rates than existing models in complex tasks with sparse rewards. https://arxiv.o…”
agents user preference construction paper
3 directory members surfaced this signal.
3 experts
1 community
1 sources clustered
“Beyond expert users: agents should help users construct preferences, not just elicit them Irena Saracay, Ludwig Schmidt, Carlos Guestrin https://t.co/C6dYa9Tgwh [𝚌𝚜.𝙰𝙸] https://t.co/LTMprro9iE”
“2026 - agents should help users construct preferences, not just elicit them - Irena Saracay, Ludwig Schmidt, Carlos Guestrin 4/n”
New
AI field signal
Signal
3h ago
1 directory member surfaced this signal.
1 expert
0 communities
1 sources clustered
“PAC Approximation and DIRECT Optimization for Parametric Markov Models Zhiming Chi, Ying Liu, Andrea Turrini, Lijun Zhang, David N. Jansen https://t.co/qgLEcbzG8K [𝚌𝚜.𝙰𝙸 𝚌𝚜.𝙵𝙻 𝚌𝚜.𝙻𝙾] https://t.co/XPtr4F8NSC”
1 directory member surfaced this signal.
1 expert
0 communities
1 sources clustered
“From Profiling to Synthesis: Benchmarking Implicit Behavioral Alignment in Personalized LLM Agents Jiajia Song, Bobo Li, Haiwen Yi, Zibo Ji, Meishan Zhang, Hao Fei, Min Zhang, Mong-Li Lee, Wynne Hsu https://t.co/XMvRBuYNGK [𝚌𝚜.𝙰𝙸] https://t.co/yC4b4DBlAD”
1 directory member surfaced this signal.
1 expert
0 communities
1 sources clustered
“From Simple QA to Deep Research: A Verifiable Benchmark Constructed through Iterative Task Evolution Can Wang, Haoran Chen, Haowen Gao, Hao Ding, Zhaoyang Liu, Zhiying Tu https://t.co/IXDljUzJGP [𝚌𝚜.𝙰𝙸] https://t.co/U9Oo9QY9E4”
1 directory member surfaced this signal.
1 expert
0 communities
1 sources clustered
“Auditing Data Provenance in LLM Fine-tuning via Intrinsic Distributional Fingerprints Zirui Huang, Yunlong Mao, Wei Tong, Tingting Wu, Xin Ge, Sheng Zhong https://t.co/aRYhmxl9cJ [𝚌𝚜.𝙰𝙸] https://t.co/Uv1m7ZD11Z”
1 directory member surfaced this signal.
1 expert
0 communities
1 sources clustered
“Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Yijun Zhang, Yule Xie, Jiaxin Ding, Xin Ding, Fan Xu, Haoxiang Zhang, Luoyi Fu https://t.co/WCb8eWaJpl [𝚌𝚜.𝙰𝙸] https://t.co/K5e7AxCS3H”
1 directory member surfaced this signal.
1 expert
0 communities
1 sources clustered
“Beyond Solution-Centric Search: Adaptive Inquiry and Knowledge Revision for Autonomous ML Engineering Shaokang Fu, Yulong Tao, Linbo Jin, Jiarong Zhao, Qiming Shi, Tianjun Pan, Haonan Li, Chengyu Wang, Jia Wu, Chengfu Huo https://t.co/41IbQdbNQe [𝚌𝚜.𝙰𝙸] htt…”
1 directory member surfaced this signal.
1 expert
0 communities
1 sources clustered
“MemArbiter: Decision-Time Memory Arbitration for Long-Horizon LLM Agents Jiajun Dong, Yutao Hu, Fengrui Fan, Shihan Dou, Yueming Wu, Deqing Zou https://t.co/uehu13VNuf [𝚌𝚜.𝙰𝙸] https://t.co/4GBDMtiwfC”
1 directory member surfaced this signal.
1 expert
0 communities
1 sources clustered
“Fetch-then-Explore: Decoupling Selection from Extraction over a Persistent Workspace for Search Agents Qi Liu, Yiqun Chen, Zidan Chen, Yan Gao, Yi Wu, Yao Hu, Jiaxin Mao, Fengbin Zhu, Tat-Seng Chua https://t.co/JuMTs5b6i5 [𝚌𝚜.𝙰𝙸 𝚌𝚜.𝙸𝚁] https://t.co/6vqrbuiMvw”
1 directory member surfaced this signal.
1 expert
0 communities
1 sources clustered
“Instruction-Conditioned Exploration with Asymmetric Reinforcement Learning and Self-Distillation Jim Dilkes, Vahid Yazdanpanah, Sebastian Stein https://t.co/rEz55Kj5aP [𝚌𝚜.𝙰𝙸 𝚌𝚜.𝙲𝙻 𝚌𝚜.𝙻𝙶] 💬Submitted to ACL Rolling Review (ARR) May 2026 cycle https://t.co/bk…”
1 directory member surfaced this signal.
1 expert
0 communities
1 sources clustered
“HPFA: Hypergraph-Based Paired Failure Attribution for LLM Reasoning Runchuan Zhu, Hongbin Lai, Bowen Jiang, Junrui Zhang, Zhangheng LI, Ostap Kilbasovych, Junyuan Hong https://t.co/5gk2xSucAk [𝚌𝚜.𝙰𝙸] https://t.co/9cuTOiNtQn”
1 directory member surfaced this signal.
1 expert
0 communities
1 sources clustered
“EduZone: A Framework for Evaluating LLM Safety for K-12 Students and Teachers Junyeong Park, Jieun Han, Haneul Yoo, So-Yeon Ahn, Jinsung Yoon, Alice Oh https://t.co/o6wJUKgGCl [𝚌𝚜.𝙰𝙸 𝚌𝚜.𝙲𝚈] https://t.co/oneGsHojxd”
1 directory member surfaced this signal.
1 expert
0 communities
1 sources clustered
“Before Reasoning Fails: Pre-Evidence Procedural Failures in Agentic RAG Daeyoung Roh, Donghee Han https://t.co/UAHHdHfOPd [𝚌𝚜.𝙰𝙸] 💬Code: https://t.co/r0ODD9W7iy https://t.co/sumJS6l3IS”
New
AI field signal
Signal
8h ago
1 directory member surfaced this signal.
1 expert
0 communities
1 sources clustered
“Before Reasoning Fails: Pre-Evidence Procedural Failures in Agentic RAG Daeyoung Roh, Donghee Han https://t.co/UAHHdHfOPd [𝚌𝚜.𝙰𝙸] 💬Code: https://t.co/r0ODD9W7iy https://t.co/sumJS6l3IS”
1 directory member surfaced this signal.
1 expert
0 communities
1 sources clustered
“HALT: Verification-Aware Stopping for Retrieval-Augmented Search Agents Daeyoung Roh, Donghee Han https://t.co/9wteuEkEqs [𝚌𝚜.𝙰𝙸] 💬Code: https://t.co/6tKWJT3NOq https://t.co/XpDcxLwmrw”
New
AI field signal
Signal
8h ago
1 directory member surfaced this signal.
1 expert
0 communities
1 sources clustered
“HALT: Verification-Aware Stopping for Retrieval-Augmented Search Agents Daeyoung Roh, Donghee Han https://t.co/9wteuEkEqs [𝚌𝚜.𝙰𝙸] 💬Code: https://t.co/6tKWJT3NOq https://t.co/XpDcxLwmrw”
New
AI field signal
Signal
9h ago
1 directory member surfaced this signal.
1 expert
0 communities
1 sources clustered
“Evolving in the Agent Jungle via History-Informed Opponent Awareness Zhaofeng Zhang, Linhan Xia, Rui Liu, Yihao Wang, Binrui Shen, Shengxin Zhu https://t.co/p8YbGNObL8 [𝚌𝚜.𝙰𝙸] https://t.co/NtOw4uxMzK”