“Xinyu Geng, Xuanhua He, Sixiang Chen, Yanjing Xiao, Fan Zhang, Shijue Huang, Haitao Mi, Zhenwen Liang, Tianqing Fang, Yi R. Fung DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment https://arxiv.org/abs/2607.07820”
“Abir Harrasse, Michael Lan, Hunar Batra, Fateme Hashemi Chaleshtori, Chaithanya Bandi Reasoning Fine-Tuning Induces Persistent Latent Policy States https://arxiv.org/abs/2607.18532”
“A study reveals that reasoning fine-tuning transforms models by inducing richer latent-policy structures, enhancing multi-step reasoning through improved dynamics. This boost in performance offers fresh insights into how models handle complex reasoning task…”
“Lizhe Fang, Weizhou Shen, Tianyi Tang, Yisen Wang Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning https://arxiv.org/abs/2607.19345”
“Researchers developed GEAR, a method that reduces repetitive copying in long-context reasoning, improving accuracy by up to +4.6 points across benchmarks, highlighting the importance of focused reasoning in AI tasks. https://arxiv.org/abs/2607.19345”
“Xuefeng Jin, Jiashuo Zhang, Teng Cao, Bin Yang Beyond Score Prediction: LLM-Based Essay Scoring and Feedback Generation via Reinforcement Learning with Rubric Rewards https://arxiv.org/abs/2607.19219”
“RLAES leverages reinforcement learning to enhance automated essay scoring and feedback generation, achieving top QWK scores while ensuring high-quality, rubric-based evaluations. This approach promises to revolutionize educational assessments. https://arxiv…”
“Akshat Gupta, Jermaine Lei, Alexander Lu, Gopala Anumanchipalli, Leshem Choshen Automated Discovery Has No Universally Superior Harness https://arxiv.org/abs/2607.18235”
2 experts discussed this · 15 posts
Leshem (Legend) Choshen @EMNLP: OpenEvolve underperforms simple autmated discovery harnesses the rest are insignificant from each other. The best choice changed across model–problem pairs. We ran a controlled study (3m+ rollouts)…
Leshem (Legend) Choshen @EMNLP: Automated discovery has high run-to-run variance, yet harnesses are often evaluated with only 3–5 runs. When a harness performs better, how do we know it is genuinely better—and not simply lucky?
Leshem (Legend) Choshen @EMNLP: To answer this, we systematically evaluated 30 budget-matched harnesses across 12 model–problem pairs, using repeated-trial statistical analysis. We found :
“Chuheng Du, Junyi Chen, Hanlin Tang, Kan Liu, Tao Lan, Lin Qu, Chaoyue Niu, Shengzhong Liu, Guihai Chen, Fan Wu C$^2$KV: Compressed and Composable KV Cache Reuse for Efficient LLM Inference https://arxiv.org/abs/2607.17715”
“C2KV greatly improves large language model inference efficiency via composable, compressed key-value cache reuse, achieving 17× speedup without sacrificing quality. It addresses KV storage and access costs, supporting scalable long-context applications. htt…”
“Bogdan Raduta, Horia Velicu, Alexandru Preda, Serban Chiricescu A Classifier That Teaches Itself: Self-Improving, Frozen-gate Training (SIFT) for Dynamic Document Classification https://arxiv.org/abs/2607.18358”
“Pilar S\'anchez-Gij\'on, Susana Valdez, Sof\'ia Calvo Del Barrio, Florence Bellemont, Anna Kokkinidou, Mihai Cristian Brasoveanu Building a European Multilingual Evaluation Dataset: The MMLU Localisation Project within the EMT Network https://arxiv.org/abs/…”
“Jia Ao Sun, Hao Yu, Fengran Mo, Zhan Su, Yuchen Hui, Bang Liu, Jian-Yun Nie Search-on-Graph-R1: Training Large Language Models to Search Knowledge Graphs with Reinforcement Learning https://arxiv.org/abs/2607.18481”
We use essential cookies to keep the site working (login, form security). With your permission, we also use analytics cookies to understand how you use the site.
Privacy policy