AI evaluation in education study
2 directory members surfaced this signal.
2 experts
2 communities
1 sources clustered
“Battles over evidence in education have raged for years. But with AI it seems "evidence" is now a matter of rapid implementation followed by accelerated scaling. Not so much "what works" but "make it work." My post on this: codeactsineducation.wordpress.com…”
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
“Guilherme C. Oliveira, Stephanie Fong, Zimu Wang, Clarice Lee, Xiangyu Zhao, Duy Khoa Pham, Duong Nhu, Yiwen Jiang, Jiahe Liu, Zhongxing Xu, ... AnchorSIPS: A Synthetic Dataset and Evaluation Resource for Evidence-Supported Psychosis-Risk Symptom Measuremen…”
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
“Keito Inoshita Class-Structure Preservation Beats Diversity: A Comprehensive Benchmark of Text Augmentation Methods for Imbalanced Text Classification https://arxiv.org/abs/2608.12340”
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
“Ying He, Zhouhong Gu, Zhecheng Hu, Yubo Zhou, Hao Shen, Jiaqing Liang, Zhaoqian Dai, Shuguang Ma, Fei Yu, Yanghua Xiao, Zhixu Li Are Large Language Models Reliable Reviewers? A Benchmark for Error Detection in Financial Documents https://arxiv.org/abs/2608.…”
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
“Adrian Trzoss, Kacper Dudzic, Wiktor Werner, Marcin Moskalewicz Large Language Models Pass the History Exam But Miss the <<History>>: A Polish High School Exit Exam Matura Benchmark https://arxiv.org/abs/2608.12343”
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
“Fali Wang, Ali Al-Lawati, Iliyas Bektas, Jinxuan Fang, Alek Melenski, Tianxiang Zhao, Yao Ma, Suhang Wang Unified Multi-Dimensional Benchmark for Complex Graph Reasoning in Large Language Models https://arxiv.org/abs/2608.12391”
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
“Florian Braun Excess Separability: Nuisance-Controlled Residual-Stream Probing for Benchmark Contamination Detection https://arxiv.org/abs/2608.12652”
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
“Jophin John, Michael Hoffmann, Jan Fillies, Michael A. Hedderich, Barbara Plank BavGround: A Benchmark for Regional Cultural Grounding and Dialect Competence in Bavarian https://arxiv.org/abs/2608.12894”
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
“Chenrun Wang, Mingxuan Zhu, Tiancheng Huang, Wenjie Li, Yujie Zhang, Zichen Zhu, Zhiying Zou, Kai Yu, Lu Chen LigBench: A Unified and Human-Aligned Benchmark for LLM-based Research Idea Generation https://arxiv.org/abs/2608.13136”
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
“Paul Osemudiame Oamen, Owusu-Banahene Osei, Ananya Mukherjee, Christian Greisinger, Steffen Eger, Pius Onobhayedo, Wei Zhao How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures https://arxiv.org/abs/2608.13267”
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
“Junhao Luo (School of Statistics, Data Science, Southwestern University of Finance, Economics), Ning Huang (School of Statistics, Data Science, ... Beyond Local Accuracy: A Protocol-Level Identifiability Audit for Controlled LLM Reasoning Evaluation https:/…”
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
“SimuScene is a benchmark for testing LLMs' ability to create physics-inspired animations from code, revealing challenges in generating accurate visuals. Using reinforcement learning, the research suggests ways to enhance AI performance in educational visual…”
Epoch AI benchmarking engineer hiring
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
“If you find these questions interesting and would like to help us answer them, apply to work at Epoch! Researchers: jobs.lever.co/epoch-ai/de... Engineers: jobs.lever.co/epoch-ai/d1... All careers: epoch.ai/about/careers”
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
“Ruoxi Zhao, Maziar Raissi Backtrader-Bench: Benchmarking LLM Agents on Algorithmic Trading with Self-Generated MCQs https://arxiv.org/abs/2608.11232”
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
“Jiahui Zhang, Ziwei Zhang, Yipeng Wang, Yibo Liu, Haozhou Pang, Yikai Hu, Hongyan Ren, Lan Zhou, Qi Gan, Kai Sheng TRACE Bench: Task-driven Roleplay Agentic Checklist Evaluation https://arxiv.org/abs/2608.11236”
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
“Haokai Zhao, Yunze Xiao, Weihao Xuan, Flora Salim, Benjamin Tag, Aditya Joshi Group Alignment-Induced Sycophancy: A Two-Sided Evaluation of Steerable Pluralistic Alignment https://arxiv.org/abs/2608.11528”
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
“Kegeng Tang, Jingbo Wang, Shaogang Ren, Zihao Wang CT-$\Delta$Bench: A Benchmark for Longitudinal 3D Medical Imaging Difference Reporting with Vision-Language Models https://arxiv.org/abs/2608.11534”
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
“Shailja Thakur, Sungeun An, Chad DeLuca, Hima Patel The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance https://arxiv.org/abs/2608.11694”