SlopCodeBench coding agents benchmark
2 directory members surfaced this signal.
Why this matches
Research & technical analysis reaction
1 attributable expert contribution
· Eugene Vinitsky
“This is a really excellent benchmark, pointing to some serious missing capabilities in coding agents: arxiv.org/abs/2603.247.... They don't write code with an eye towards maintenance!”
evidence ↗
2 experts
1 community
1 sources clustered
“This is a really excellent benchmark, pointing to some serious missing capabilities in coding agents: arxiv.org/abs/2603.247.... They don't write code with an eye towards maintenance!”
Bolivia roadblock hybrid forecasting paper
1 directory member surfaced this signal.
Why this matches
Research & technical analysis reaction
1 attributable expert contribution
· AI Firehose
“This study presents a hybrid forecasting system merging time series analysis with NLP to predict Bolivian roadblocks, enhancing logistics and economic outcomes. This approach outperformed standard models, showing how news signals reveal tensions ahead of co…”
evidence ↗
1 expert
1 community
1 sources clustered
“This study presents a hybrid forecasting system merging time series analysis with NLP to predict Bolivian roadblocks, enhancing logistics and economic outcomes. This approach outperformed standard models, showing how news signals reveal tensions ahead of co…”
LLM consensus preference evaluation paper
1 directory member surfaced this signal.
Why this matches
Research & technical analysis reaction
1 attributable expert contribution
· arxiv cs.CL
“Mohtashim Khan A Consensus-Based Framework for Relative Preference Evaluation of Large Language Models https://arxiv.org/abs/2607.21632”
evidence ↗
1 expert
1 community
1 sources clustered
“Mohtashim Khan A Consensus-Based Framework for Relative Preference Evaluation of Large Language Models https://arxiv.org/abs/2607.21632”
biomedical MeSH NLP evaluation paper
1 directory member surfaced this signal.
Why this matches
Research & technical analysis reaction
1 attributable expert contribution
· arxiv cs.CL
“Samuel M. Okoe-Mensah Evaluation design conditions the expert-vs-auto MeSH gap: a controlled comparison of bag-of-words and BiomedBERT on the Cohen benchmark https://arxiv.org/abs/2607.21685”
evidence ↗
1 expert
1 community
1 sources clustered
“Samuel M. Okoe-Mensah Evaluation design conditions the expert-vs-auto MeSH gap: a controlled comparison of bag-of-words and BiomedBERT on the Cohen benchmark https://arxiv.org/abs/2607.21685”
Khondo Bangla document benchmark
1 directory member surfaced this signal.
Why this matches
Research & technical analysis reaction
1 attributable expert contribution
· arxiv cs.CL
“Abu Tyeb Azad, Fahim Ahmed, Ishita Sur Apan, Ezharuddin Jubaer, Sumaiya Karim Katha, Armun Alam, Amin Ahsan Ali, Aman Chadha, Md Mofijul Islam, AKM Mahbubur Rahman Khondo: A Multimodal Benchmark for Document Packet Splitting of Bangla Forms https://arxiv.or…”
evidence ↗
1 expert
1 community
1 sources clustered
“Abu Tyeb Azad, Fahim Ahmed, Ishita Sur Apan, Ezharuddin Jubaer, Sumaiya Karim Katha, Armun Alam, Amin Ahsan Ali, Aman Chadha, Md Mofijul Islam, AKM Mahbubur Rahman Khondo: A Multimodal Benchmark for Document Packet Splitting of Bangla Forms https://arxiv.or…”
agentic copyright law evaluation paper
1 directory member surfaced this signal.
Why this matches
Research & technical analysis reaction
1 attributable expert contribution
· arxiv cs.CL
“Zheng Hui, Doni Bloomfield, Noam Kolt Agentic Evaluation of Copyright Law Compliance https://arxiv.org/abs/2607.21799”
evidence ↗
1 expert
1 community
1 sources clustered
“Zheng Hui, Doni Bloomfield, Noam Kolt Agentic Evaluation of Copyright Law Compliance https://arxiv.org/abs/2607.21799”
agent memory longitudinal evaluation paper
1 directory member surfaced this signal.
Why this matches
Research & technical analysis reaction
1 attributable expert contribution
· arxiv cs.CL
“Quentin Spencer Ground Truth First: A Longitudinal Evaluation Instrument for Agent Memory, and the Tenure Crossover in Memory-Architecture Rankings https://arxiv.org/abs/2607.21962”
evidence ↗
1 expert
1 community
1 sources clustered
“Quentin Spencer Ground Truth First: A Longitudinal Evaluation Instrument for Agent Memory, and the Tenure Crossover in Memory-Architecture Rankings https://arxiv.org/abs/2607.21962”