The Challenge
Large language model (LLM) tutoring systems face a critical measurement and alignment problem: general-purpose helpfulness rubrics frequently fail to distinguish between direct, answer-giving shortcuts and true pedagogical guidance that fosters independent learning. In socio-technical and educational deployments, relying on uncalibrated evaluation signals risks optimizing AI agents for superficial user satisfaction rather than substantive skill acquisition. This pre-registered audit investigates whether standard LLM-judged helpfulness can serve as a reliable proxy for pedagogical quality, exploring the dangerous divergence between making users feel helped versus genuinely driving cognitive engagement and independent student work.
Core Findings
The authors conducted a rigorous, pre-registered audit across multiple tutor models and policies paired with a fixed simulated student, analyzing 1,179 confirmatory answer-phase turns. Using Claude Opus 4.8 and GPT-5.6 Sol as condition-blind judges alongside deterministic detectors, the study evaluated conversational versus pedagogical policies. The findings reveal that while policies do not differ significantly in general helpfulness, they are perfectly rank-separated under specialized pedagogy rubrics (Cliff's delta of 1.0 vs. 0.10). Furthermore, helpfulness rankings proved heavily judge-contingent—reversing directions across judges on two of three bases—while answer-revealing turns consistently resulted in less independent student work across all tested models.
Strategic Takeaway
For digital leaders and educational technologists, this research demonstrates the severe limitations of employing generic consumer metrics (like helpfulness or user satisfaction) to govern enterprise or educational AI deployments. Organizations must decouple user preference from competency development by implementing multi-layered evaluation frameworks. This means pairing targeted domain-specific rubrics with deterministic process measures rather than trusting off-the-shelf general-purpose LLM judges. In socio-technical system design, optimizing for immediate interactional friction reduction often undermines long-term capability building, necessitating rigorous, theory-driven audit pipelines before operationalizing generative AI agents.