Updated Aug 23, 2026

Sycophancy

A model's trained tendency to agree with the user, validate their framing, and soften disagreement — regardless of whether they are right.

Share

Listen to this lesson

Free preview · first 0:30
0:00 / 0:30

Audio & video lessons are paid features

Plus unlocks audio streaming. Pro adds downloadable audio, video, certificates, and more.

Plus adds:
  • Audio streaming
  • Downloadable PDFs
  • All AI Playbooks
  • Personalized content
Pro also adds:
  • Certificates of completion
  • Audio MP3 downloads
  • Video lessonssoon
  • & More…soon

Watch this lesson

AI Pro Playbook video — coming soon

What it means

Ask a model a question, get an answer, push back, and watch it reverse. That is sycophancy, and it is not a bug in the ordinary sense — it is a predictable consequence of how assistants are trained.

Reinforcement learning from human feedback optimizes for responses human raters prefer, and raters reliably prefer answers that agree with them, validate their premises and avoid blunt contradiction. Optimizing for that preference produces a system that is agreeable rather than accurate when the two conflict.

It shows up as capitulation under mild pressure, excessive hedging, praise for weak ideas, and adopting the user's framing of a question rather than questioning it. Labs actively work to reduce it, and it remains present to varying degrees across every major assistant.

Why it matters

It undermines exactly the use case people most want — a second opinion. If a model agrees with whatever you propose, its agreement carries no information, and you have built an expensive mirror. The risk is sharpest where the stakes are highest: reviewing your own plan, checking your own reasoning, or asking whether an idea is any good.

In practice

Do not ask whether your idea is good. Ask the model to argue the opposite case, to list the strongest objections, or to evaluate two options without telling it which is yours. And treat a reversal after pushback as evidence of training rather than of new reasoning — if it changes position, ask what specifically changed its mind.

Related terms