DigitalCoach: AI coaches skip the teaching moves humans use
TL;DR
- The DigitalCoach dataset captures 72 expert-novice sessions, 22,752 dialogue turns and 28.1 hours of screen and input recordings across five software applications.
- Automated evaluation finds models give more direct instructions but fewer explanations, error diagnoses and knowledge-check questions than human coaches.
- Even when coaching method is standardized, model utterances remain poorly grounded in visual context and learners fall into passive instruction-following.
A new dataset of 72 expert-novice coaching sessions finds that today's models teach software use very differently from humans, and not just in style. In the DigitalCoach paper, the authors record 22,752 dialogue turns across 28.1 hours of screen and input event recordings, covering five software applications.
The finding is blunt: "models provide more direct instructions, but fewer explanations, error diagnoses, and knowledge-check questions." When the researchers hold the coaching method fixed to test whether the gap is only stylistic, models "produce utterances similar to human references yet poorly grounded in visual context." In interactive tests, learners with model coaches "passively follow instructions without deeper engagement."
The team (Meng Chen, Anya Ji, Tsung-Han Wu, Tobias Maringgele, David M. Chan, Alane Suhr and Amy Pavel) releases the dataset and code. Two experts in our Who's Who directory shared the paper's link.
The abstract does not name which models were tested or which of the five applications produced the widest coaching gap.
Shared on Bluesky by 2 AI experts
-
Sharing this really cool paper of ours that will appear at EMNLP in November: arxiv.org/abs/2606.319...
View on Bluesky →
Originally reported by arxiv.org
Read the original article →Original headline: DigitalCoach: Communication and Grounding Gaps in Human and Agentic Computer Use Coaching