last translation benchmark paper
3 experts across 3 network communities independently surfaced this.
3 experts
3 communities
1 sources clustered
“Machine translation is not solved and it will take a while for it to be done arxiv.org/abs/2609.04173”
“Vil\'em Zouhar, Niyati Bafna, Mukund Choudhary, Maike Z\"ufle, Sara Rajaee, Pinzhen Chen, Jannis Vamvas, Sara Papi, Ona de Gibert, Bhavitvya Malik, Eliya Habba, Orfeas Menis Mastromichalakis, Patr\'icia Schmidtov\'a, ... Last Translation Benchmark https://a…”
2 experts discussed this · 6 posts
Vilém Zouhar: Machine translation is not solved and it will take a while for it to be done arxiv.org/abs/2609.04173
Vilém Zouhar: We just released the Last Translation Benchmark paper. In a massive crowdsourcing effort we collected 3456 unique hard-to-translate examples that break state-of-the-art translation models, and whic…
Vilém Zouhar: Machine translation doesn't break on just figurative language as one would expect. In fact, for the next generation of models, we may need to invest heavily into multilingual (& cultural) reasoning.
Open the full discussion →
self-evolving LLM harness optimization paper
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
“Michael Nguyen, Wei Chen Tan, Nurul Aisyah Hassan, Arvind Raman, Li Hua Lim, Ahmad Faiz Razak Where Does Harness-Optimization Value Live? Localized Gains and the Budget-Splitting Trap in Self-Evolving LLM Agents https://arxiv.org/abs/2609.02889”
sequential distillation RLVR interplay paper
2 directory members surfaced this signal.
2 experts
1 community
1 sources clustered
“Boyan Li, Bingsen Chen, Chenghao Yang, Ping Nie, Chen Zhao, Xi Ye Sequential Beats Joint: On the Interplay between On-Policy Distillation and RLVR https://arxiv.org/abs/2609.04108”
“A study shows that the sequential method of on-policy distillation followed by RL with verifiable rewards outperforms existing training techniques for large language models, enhancing reasoning abilities and optimizing model performance. https://arxiv.org/a…”
VERDICT AI clinical trial matching paper
2 directory members surfaced this signal.
2 experts
1 community
1 sources clustered
“Zikai Zhou, Yufei Jin, Yilin Xu, Yu-Chiang Wang, Chieh-Ju Chao, Monica S. Lam Accountable AI with Grounded, Faithful, Consistent, Actionable Rationales: A Case Study in Clinical Trial Matching with VERDICT https://arxiv.org/abs/2609.03366”
“VERDICT changes clinical trial matching by ensuring accountability in AI, providing consistent rationales that clinicians prefer. Its formal method enhances transparency and reliability, marking a leap in trust for AI in healthcare. https://arxiv.org/abs/26…”
natural language neural functions compile paper
2 directory members surfaced this signal.
2 experts
1 community
1 sources clustered
“Yuntian Deng, Pengyu Nie, Stuart Shieber Compile by Training: Turning Natural-Language Specifications into Local Neural Functions https://arxiv.org/abs/2609.04199”
“The innovative method "compile by training" converts natural-language specifications into efficient neural functions, achieving 83.6% semantic accuracy on complex tasks without latency of large models. This advancement paves the way for scalable AI solution…”
training-free lossy speculative decoding paper
2 directory members surfaced this signal.
2 experts
1 community
1 sources clustered
“Oszk\'ar Urb\'an, Young D. Kwon, Stylianos I. Venieris, Cecilia Mascolo Margins, Not Windows: Training-Free Per-Step Lossy Speculative Decoding https://arxiv.org/abs/2609.02897”
“AdaptiveSpec improves LLM inference by optimizing speculative decoding through dynamic adjustments in token verification and draft tree structure, achieving up to 56% throughput gains over EAGLE-3 while preserving near lossless accuracy across multiple benc…”
hybrid RAG routing rewriting adapter paper
2 directory members surfaced this signal.
2 experts
1 community
1 sources clustered
“Yucan Guo, Miao Su, Saiping Guan, Long Bai, Zhongni Hou, Zixuan Li, Xiaolong Jin, Jiafeng Guo, Xueqi Cheng R$^{2}$Adapter: A Routing and Rewriting Adapter for Efficient Hybrid RAG https://arxiv.org/abs/2609.02894”
“Introducing R2ADAPTER: a model-agnostic plug-in that dynamically allocates queries in Retrieval-Augmented Generation systems. It cuts graph-based usage by up to 59% while preserving accuracy, providing a lightweight solution for multi-hop reasoning in LLMs.…”
random attention KV cache eviction paper
2 directory members surfaced this signal.
2 experts
1 community
1 sources clustered
“Heng Wang, Jielin Qiu, Wenting Zhao, Cheng Qian, Liangwei Yang, Jiawei Han, Heng Ji, Silvio Savarese, Shelby Heinecke, Huan Wang Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning https://arxiv.org/abs/2609.03430”
“Salesforce AI Research's "Random Attention" improves KV cache eviction, enhancing reasoning model throughput by 32-43%, matching top methods without complex scoring. This innovation underscores the power of reasoning traces, boosting memory efficiency in LL…”
spurious CoT termination reasoning paper
2 directory members surfaced this signal.
2 experts
1 community
1 sources clustered
“Seunghee Koh, Sungjae Choi, Minchan Kwon, Sunghyun Baek, Junmo Kim </think> Doesn't Stop Reasoning: Analysis of Spurious CoT Termination https://arxiv.org/abs/2609.03633”
“A study reveals that injecting end-of-think tokens in reasoning can unintentionally extend reasoning into the answering phase, complicating responses. By enhancing attention to these tokens, researchers reduced redundant output and improved efficiency. http…”
LLM story consistency analysis paper
2 directory members surfaced this signal.
2 experts
2 communities
1 sources clustered
“Thennal DK, Hans Ole Hatzel Do Large Language Models Always Tell The Same Stories? https://arxiv.org/abs/2606.17350”
bounded personas agent retrieval paper
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
“JaeHa Yoon, Minjun Park, Seoyeon Kim, Jiwoo Lee, Hyunwoo Choi, Dohyun Kang Bounded Personas Match Retrieval on Classification but Not Regression for a Frozen Agent https://arxiv.org/abs/2609.02890”
counterexamples agent self-correction paper
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
“Sidhesh Badrinarayan, Adithya Parthasarathy Counterexamples as Feedback for Agent Self-Correction https://arxiv.org/abs/2609.02892”
OOD deception detection probe paper
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
“Daniel Yoo, Adrians Skapars Probe Generalization as Subspace Selection for OOD Deception Detection https://arxiv.org/abs/2609.02893”
India misinformation benchmark dataset paper
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
“Parth Bramhecha, Smit Deshmukh, Sairaj Bodhale, Adwait Borate, Raviraj Joshi BharatGather: A Culturally-Informed Benchmark Dataset for Misinformation and Fake News Detection in Indian Public Events https://arxiv.org/abs/2609.02895”
LLM medical relation extraction paper
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
“Jiaxin Duan, Fengyu Lu, Junfei Liu PiPMRE: A Pipeline Based on Language Model for Medical Relation Extraction https://arxiv.org/abs/2609.02896”
biomedical embedding transfer DRET paper
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
“Girish Sundaram, Daniel Berleant Distilled Rapid Embedding Transfer (DRET): Parameter-Efficient Biomedical Domain Adaptation via Priority-Based Embedding Transfer https://arxiv.org/abs/2609.02898”
benchmark contamination LLM leaderboards paper
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
“Xingyao Xiao (Stanford University), Yihong Cheng (City University of Macau) Contamination Inflates Scores but Rarely Reorders Large Language Model Leaderboards https://arxiv.org/abs/2609.02899”
Chinese ASR inverse text normalization paper
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
“Fengrun Zhang, Li Fu, Wangjin Zhou, Lu Fan, Youzheng Wu, Xiaodong He Dual-Form ASR: Semantics-Aware Inverse Text Normalization for Chinese Speech Recognition https://arxiv.org/abs/2609.02901”