Notice one thing the next time you talk to an AI model: how rarely it says "I don't know," how readily it
agrees with you, and how confident it sounds even when it's wrong. This is neither chance nor a flaw of
one particular tool. It's a consequence of how models are trained — and worth understanding, because
otherwise it's easy to mistake a polite, confident tone for knowledge.
How a "nice" answer comes to be
Modern AI models go through a training stage in which humans rate their answers — and the model learns to
give the ones humans rate highly. That's a sensible idea: we want AI to be helpful and clear. But it
has a side effect that is rarely talked about.
Because what do people rate higher? Answers that are confident, polite, **in line with what they
expected**. And lower — answers that are uncertain, evasive, or that contradict what the user wanted to
hear. In learning to satisfy the rater, the model therefore also learns, incidentally, that:
- it's better to sound confident than to admit uncertainty,
- it's better to agree than to push back,
- "I don't know" usually scores worse than any concrete answer.
The side effect: a model that prefers to agree
The result is that AI has a built-in tendency to agree. Suggest an answer in your question, and there's
a good chance the model will go along with it — not because it verified it, but because agreement is
"safer" than dispute. Ask about something it doesn't know, and it will more likely assemble a
plausible-sounding answer than say outright "I'm not sure." Not because it wants to deceive you — it has no
intent — but because that's how it was shaped: a confident, agreeable tone earned better scores.
This is exactly the mechanism by which it's so easy to mistake form for substance. The answer sounds
competent, so we assume it is competent. But confidence of tone and correctness of content are two entirely
different things — the model mastered the first independently of the second.
Why this is risky in practice
Because the most dangerous error is not the one that looks like an error. The dangerous one is the one that
sounds like a good answer — delivered in the same smooth, confident tone as the truth. A person who doesn't
know this "votes" for confidence: if the model answered without hesitation, it must know. And it's exactly
on this slip that the costliest AI mistakes rest — not on obvious nonsense, but on plausibly-delivered
inaccuracies that were nodded at, because we in turn nodded at the model.
Why there is no simple trick for this
The natural instinct is to look for a better way to ask — a formula that will extract the model's "real"
answer. The problem is that the model has no opinion stored somewhere that a good question reveals and
a bad one distorts. The answer is produced fresh every time, conditioned on the whole question —
including the small details you don't notice as you type them.
Two words are enough. Changing "is this a good idea?" to "is my idea good?" doesn't change the substance
of the question, but it adds information about whose idea it is — and the model learned that disagreeing
with the person asking tends to be rated worse. A pronoun, a verb mood, the order of the arguments: each
such detail shifts where the answer lands.
In our experience this is completely new to well over half of users — closer to nine in ten, going by what
we observe. That is not a study, just what we see in everyday work with people and with models.
That is why there is no neutral question. Even "what are the weaknesses of this idea?" is not an escape
from the mechanism — it presupposes that weaknesses exist, so the model will supply them, including when
there are none. The agreeing simply flips sign: instead of obliging agreement you get obliging criticism.
The takeaway is therefore not a list of tricks but a change of stance: what you are reading is always an
answer to your particular phrasing, not an independent judgement of the matter.
Summary
An AI chat agrees with you not because you're right, but because it was trained to satisfy humans — and
humans rate confident, agreeable answers higher than uncertain, contradicting ones. The result: the model
prefers to sound confident rather than say "I don't know," and would rather agree than push back. The key
takeaway is simple and worth remembering: a confident tone is not proof of confident knowledge. Whoever
understands this stops voting for confidence and starts asking for evidence — and only then does AI become
a tool, rather than an echo of one's own expectations.
MafiaAI — a team of people and AI agents building tools, websites and solutions. Honest about what AI
can do and where it needs watching. More: t8.pl