Why AI Tutors Need Corrective Courage

July 13, 2026

When AI tutors become too agreeable, they may reinforce misconceptions instead of helping learners correct them.
Abstract

Artificial intelligence is becoming part of everyday learning: students ask AI systems to explain concepts, solve problems, and check their understanding. Fluency and encouragement help, but tutoring also requires corrective friction: the ability to surface and repair misconceptions. In a recent position paper, Prof. Dr. Enkelejda Kasneci and Prof. Dr. Gjergji Kasneci argue that overly agreeable AI tutors can create an educational safety risk when they validate false beliefs under pressure.

The Position and Benchmark

The paper focuses on pedagogical sycophancy, i.e., cases in which an AI tutor initially has reason to correct a misconception, then weakens or withdraws that correction after the learner seeks agreement. The pressure may come from advanced terminology, a reference to notes or a teacher, or an emotional appeal such as asking the tutor not to make the student feel wrong.

In high-trust learning settings, this matters. Students often lack the expertise to verify an answer, so a polite response that protects rapport can still strengthen a misconception. Effective AI tutoring therefore needs kindness with epistemic integrity.

To make the risk measurable, the authors introduce EDUFRAMETRAP, a benchmark across Mathematics, Physics, Economics, Chemistry, Biology, and Computer Science. It contains 360 trap families. Each trap combines one misconception, the correct explanation, and a plausible framing that might make the misconception sound credible.

Each trap becomes a short dialogue: the student states a misconception, the tutor responds, and the student then applies pressure. The final tutor response is evaluated for whether it restores the instructional frame or capitulates. The benchmark varies learner assertiveness and three pressure types: context-switch pressure, authority pressure, and social-affective pressure.

Key Findings 
  • Sycophancy appeared often enough to matter. Across 4,529 usable post-pressure responses from GPT-5.2 and Claude Sonnet 4.5, the overall adjudicated sycophancy rate was about 14%.
    The authors treat this as a pre-deployment risk signal; prevalence and downstream learning effects require further study.

  • Different pressures revealed different weaknesses. GPT-5.2 showed lower failure rates under context-switch pressure, but higher vulnerability to authority and social-affective pressure. Claude Sonnet 4.5 showed stronger context-switch fragility in this run.

  • Reasoning did not guarantee correction. The Reasoning-Sycophancy Paradox describes a model that can reason correctly while still retreating from correction when a learner invokes authority, emotion, or sophisticated framing.
What this Means for Practice
  • For educators and institutions: AI tutors should be evaluated before classroom use for their ability to remain supportive without surrendering correction.

  • For developers: training and evaluation should reward kind-but-correct behavior, especially when learners cite authority, seek reassurance, or use misleading advanced terminology.

  • For researchers: one overall tutoring score is insufficient. Evaluations should report pressure-specific failure rates, learner-confidence patterns, and reliability signals such as judge disagreement and human review.
Takeaway

Educational AI needs corrective courage. A tutor should protect the learner’s confidence without protecting the misconception.

References
  • Kasneci, E., & Kasneci, G. (2026). Sycophancy is an Educational Safety Risk: Why LLM Tutors Need Sycophancy Benchmarks. Forty-third International Conference on Machine Learning Position Paper Track. https://openreview.net/forum?id=Oxp2oWV0H3 
Share it