Rethinking LLM-Judged Helpfulness as a Pedagogy Signal: A Pre-Registered Audit Across Tutor Models

Source: arXiv Computers & Society (Academic)

The Challenge

Large language model (LLM) tutoring systems face a critical measurement and alignment problem: general-purpose helpfulness rubrics frequently fail to distinguish between direct, answer-giving shortcuts and true pedagogical guidance that fosters independent learning. In socio-technical and educational deployments, relying on uncalibrated evaluation signals risks optimizing AI agents for superficial user satisfaction rather than substantive skill acquisition. This pre-registered audit investigates whether standard LLM-judged helpfulness can serve as a reliable proxy for pedagogical quality, exploring the dangerous divergence between making users feel helped versus genuinely driving cognitive engagement and independent student work.

Core Findings

The authors conducted a rigorous, pre-registered audit across multiple tutor models and policies paired with a fixed simulated student, analyzing 1,179 confirmatory answer-phase turns. Using Claude Opus 4.8 and GPT-5.6 Sol as condition-blind judges alongside deterministic detectors, the study evaluated conversational versus pedagogical policies. The findings reveal that while policies do not differ significantly in general helpfulness, they are perfectly rank-separated under specialized pedagogy rubrics (Cliff's delta of 1.0 vs. 0.10). Furthermore, helpfulness rankings proved heavily judge-contingent—reversing directions across judges on two of three bases—while answer-revealing turns consistently resulted in less independent student work across all tested models.

Strategic Takeaway

For digital leaders and educational technologists, this research demonstrates the severe limitations of employing generic consumer metrics (like helpfulness or user satisfaction) to govern enterprise or educational AI deployments. Organizations must decouple user preference from competency development by implementing multi-layered evaluation frameworks. This means pairing targeted domain-specific rubrics with deterministic process measures rather than trusting off-the-shelf general-purpose LLM judges. In socio-technical system design, optimizing for immediate interactional friction reduction often undermines long-term capability building, necessitating rigorous, theory-driven audit pipelines before operationalizing generative AI agents.

Deep Dive Q&A

Why is general-purpose helpfulness insufficient for evaluating AI tutors?

General-purpose helpfulness metrics measure user satisfaction or direct task completion, which often rewards the AI for simply giving direct answers rather than providing constructive pedagogical guidance that encourages critical thinking and independent problem-solving.

What methodology was used to audit the tutor models?

The study utilized a pre-registered audit design across three tutor bases, comparing conversational and pedagogical policies with a fixed simulated student, evaluated via frozen condition-blind LLM judges (Claude Opus and GPT-5.6) and deterministic process detectors.

How did the judges impact the evaluation results?

While pedagogy contrasts remained consistent across different judges, general-purpose helpfulness rankings were highly judge-contingent, reversing directions between different primary models on two out of three evaluated bases.