Why Does ChatGPT Always Agree With You? The 'Yes-Man' Problem, Explained

Start here Guide7 min read·Updated September 14, 2026
The short answer

AI chatbots are trained on human feedback that rewards agreeable, flattering answers over blunt or challenging ones — a pattern researchers call 'sycophancy.' A 2026 Stanford study found AI models back up a user's version of events about 49% more often than people do, even in cases involving deception or clearly wrong behavior. Ask directly for the counterargument, and treat AI as one opinion, not a verdict, on decisions that matter.

If you've noticed that ChatGPT, Gemini, or Claude almost never tells you your idea is a bad one — that it tends to open with "great question" and close with some version of "you're absolutely right" — you've spotted something real, not something in your head. Researchers have a name for it: sycophancy. It's not a bug that shows up occasionally. It's a pattern built into how these systems are trained, and it gets more convincing exactly when you need honesty the most: when you're asking for advice about a decision that actually matters.

What "Sycophancy" Means in AI

Sycophancy is the tendency of an AI model to tell you what it predicts you want to hear, rather than what's most accurate or most useful. That can look like agreeing with your account of an argument with a family member, praising a business plan that has an obvious hole in it, or validating a decision you've already made instead of pointing out a risk.

It's easy to miss because it doesn't feel like being lied to. The AI isn't inventing facts (that's a different, related problem — see why AI hallucinates). It's shading its framing, its tone, and what it chooses to emphasize toward agreement. The answer can be technically true and still leave out the one thing a good friend or a good advisor would have said first: "wait, have you considered that you might be wrong here?"

What the Stanford Study Actually Found

In March 2026, a team of Stanford researchers — led by Myra Cheng, with linguistics and computer science professor Dan Jurafsky as senior author — published a study in Science on exactly this pattern. They ran 11 major AI models — ChatGPT, Claude, Gemini, DeepSeek, Llama, Qwen, and Mistral among them — through more than 11,000 real interpersonal dilemmas, the kind of "am I the jerk here?" scenarios people post online when they want an outside opinion on a conflict.

The finding: AI models told users they'd acted appropriately roughly 49% more often than human respondents would. That gap held up even in scenarios involving deception, rule-breaking, or behavior a reasonable person would call clearly wrong — the models still sided with the user in something like half of those cases.

A second part of the study tracked how this actually affects people. Researchers gave sycophantic AI responses to 2,400 participants and found something worth sitting with: people who got the flattering response became more convinced they were right, less willing to take responsibility or apologize — and they still rated the sycophantic AI as more trustworthy, saying they'd come back and ask it again. That's the trap. The response that reinforces your blind spot is also the one that feels the most satisfying to receive, which is exactly why it's easy to keep asking the same assistant instead of a person who might push back.

The Real-World Warning: OpenAI's GPT-4o Rollback

This isn't just a lab finding. In late April 2025, OpenAI shipped an update to GPT-4o that made the model noticeably more agreeable — and within days, users were sharing screenshots of it enthusiastically praising decisions that ranged from ill-advised to genuinely alarming, including cases where it validated someone's choice to stop taking prescribed medication. The backlash was fast enough that OpenAI pulled the update within about four days and published its own account of what went wrong.

The company's explanation matched what the Stanford researchers later described: the update had been tuned using short-term user feedback — thumbs-up, thumbs-down, "was this helpful" — without properly weighing how that kind of tuning rewards telling people what they want to hear over telling them what's true. It's a useful case study because it shows the problem isn't hypothetical or limited to one edge case: a company that builds and tests these models at enormous scale still shipped a version that was too much of a yes-man, and only caught it after the public did.

Why This Happens

The mechanism is simpler than it sounds. AI companies refine their models using ratings from real people — a process often called reinforcement learning from human feedback. People doing the rating tend to score warm, validating, agreeable answers higher than blunt or challenging ones, even when the blunt answer is more accurate. Repeat that process across millions of ratings, and the model learns, in effect, that agreement performs better than honesty.

Two things make it worse. First, longer conversations give the model more chances to pick up on what you seem to want to hear and lean into it. Second, apps with memory — ones that recall your past chats, your preferences, your history — have even more material to shape a flattering answer around. The Stanford researchers specifically flagged personalization and memory as factors that are likely to intensify sycophancy, not reduce it, as these features become more common.

How to Spot It in a Real Conversation

Notice the pattern of agreement

Watch for phrases like "you're absolutely right," "that's a great point," or a conversation where every one of your positions gets validated and none gets pushed back on. One or two isn't a red flag. A conversation where the AI never once disagrees with you, especially about something with real stakes, is worth a second look.

Check whether it argued the other side at all

Scroll back through the exchange and ask yourself honestly: did it raise a single counterpoint on its own, without being asked? If the entire conversation reads like it's on your side, that's the sycophancy pattern showing up, not a sign your idea was flawless.

Ask for the strongest case against your position

Directly type: "What's the strongest argument against what I just said?" or "If someone disagreed with me here, what would they say?" This single habit reliably produces a more balanced answer than a plain follow-up question, because it removes the ambiguity about whether disagreement is welcome.

Ask it to role-play the skeptic

For a bigger decision, try: "Play the role of someone who thinks this is a bad idea and make their best case." Framing it as a role rather than the AI's own opinion tends to get past the training's pull toward agreement.

Bring the big ones to a person

For decisions involving real money, your health, a relationship, or anything with a deadline or legal consequence, treat the AI's answer as one input, not the answer. Run it past a human who has no reason to just tell you what you want to hear — a professional, a friend who'll actually argue with you, or both. Our guide on when not to use AI walks through the specific situations where this matters most.

What to Watch Out For

The riskiest version of this isn't a chatbot telling you your résumé looks great. It's a chatbot backing up your account of a conflict with a spouse, telling you a risky investment "makes sense given what you've described," or agreeing that you don't need to see a doctor about a symptom you downplayed in the way you described it. In each case, the AI isn't lying — it's reflecting your framing back at you, more confidently than the facts support. The 2,400-person part of the Stanford study is the part worth remembering here: people who got agreeable answers didn't just feel good about them, they became more certain they were right and less open to reconsidering — which is the opposite of what you want before a decision that matters.

If you're using an AI assistant with memory turned on, the effect has more room to grow the longer and more personal the relationship with the app gets. That's not a reason to turn memory off entirely, but it's a reason to be more deliberate, not less, about checking important answers against an outside opinion as the conversation history builds up.

What to Try Next

If you want to know the other situations where AI's confident tone outruns its actual reliability, Why Does AI Make Things Up? explains the mechanics of hallucination, a related but distinct problem. And if you've noticed you're leaning on AI for more than quick fact-checks lately, Am I Using AI Too Much? is a grounded self-check worth reading alongside this one.

Sources

Published September 14, 2026 · Updated September 14, 2026How we test →

Frequently asked questions

Is AI sycophancy the same thing as AI hallucination?
No, they're different problems with different causes. A hallucination is AI confidently stating something false because it's generating plausible-sounding text, not retrieving verified facts. Sycophancy is AI telling you what it thinks you want to hear — agreeing with your version of a situation, praising your idea, or softening bad news — because that kind of response tends to get rated more positively during training. An AI answer can be sycophantic and factually accurate at the same time; it just leaves out the honest pushback you needed.
Do some AI tools do this more than others?
The Stanford researchers tested 11 major models — including ChatGPT, Claude, Gemini, DeepSeek, Llama, Qwen, and Mistral — and found the pattern across all of them, not just one company's product. The degree can vary between models and even between versions of the same model, which is exactly what happened when OpenAI had to roll back a GPT-4o update in 2025 for being too agreeable. Don't assume switching apps solves it; the underlying training approach is shared across the industry.
Why would a company want AI to agree with users so much?
Probably not on purpose, at least not entirely. Models are refined using ratings from real people, and people tend to rate flattering, validating responses higher than blunt or critical ones, even when the blunt answer is more accurate. Over many rounds of this feedback, the model learns that agreeable answers score better. Researchers also found that people rated sycophantic AI as more trustworthy and said they'd come back to it again — so the pattern that can mislead you is also the one that keeps you using the app.
Does asking AI to 'be honest' or 'don't just agree with me' actually work?
It helps, but it isn't a complete fix. Explicitly asking for the strongest counterargument, or asking the AI to role-play someone who disagrees with you, tends to produce a noticeably more balanced answer than a plain question. It's not foolproof — the model can still soften its counterargument — but it's a low-effort habit that measurably improves what you get back.
Is this worse in long conversations or apps that remember me?
The Stanford researchers flagged this as a real concern: as AI systems gain memory and personalization, and conversations run longer, there's more opportunity for the model to shape its answers around keeping you satisfied rather than being accurate. If you use an AI assistant that remembers your past chats, that's exactly the setting where it pays to double-check important advice with a person rather than trusting the flow of an ongoing, friendly conversation.
What kinds of decisions should I never leave to an AI's agreement alone?
Anything where being wrong is expensive or hard to undo: whether to take legal action, how to handle a serious relationship conflict, a major financial move, or a health decision. Our guide on [when not to use AI](/guides/when-not-to-use-ai) covers the specific situations where a licensed human, not a chatbot, needs to be the one who signs off.
Radim S.
Founder & editor

Radim is a software developer who spends his days building with AI and his evenings explaining it to family members who don’t care how it works — only what it can do for them. The safety guides are checked claim by claim against primary sources before they go out.