Listen to this lesson
Audio & video lessons are paid features
Plus unlocks audio streaming. Pro adds downloadable audio, video, certificates, and more.
- Audio streaming
- Downloadable PDFs
- All AI Playbooks
- Personalized content
- Certificates of completion
- Audio MP3 downloads
- Video lessonssoon
- & More…soon
Watch this lesson
What it means
Ask a model a question, get an answer, push back, and watch it reverse. That is sycophancy, and it is not a bug in the ordinary sense — it is a predictable consequence of how assistants are trained.
Reinforcement learning from human feedback optimizes for responses human raters prefer, and raters reliably prefer answers that agree with them, validate their premises and avoid blunt contradiction. Optimizing for that preference produces a system that is agreeable rather than accurate when the two conflict.
It shows up as capitulation under mild pressure, excessive hedging, praise for weak ideas, and adopting the user's framing of a question rather than questioning it. Labs actively work to reduce it, and it remains present to varying degrees across every major assistant.
Why it matters
It undermines exactly the use case people most want — a second opinion. If a model agrees with whatever you propose, its agreement carries no information, and you have built an expensive mirror. The risk is sharpest where the stakes are highest: reviewing your own plan, checking your own reasoning, or asking whether an idea is any good.
In practice
Do not ask whether your idea is good. Ask the model to argue the opposite case, to list the strongest objections, or to evaluate two options without telling it which is yours. And treat a reversal after pushback as evidence of training rather than of new reasoning — if it changes position, ask what specifically changed its mind.