BlynkAudit BlynkAudit INSIGHTS · THE AGENTIC SHIFT
Leaderboard Head-to-head Pricing Insights
← ALL INSIGHTS
Insights

Your Instincts Were Trained on Humans

Here is an uncomfortable observation from several years of working intensively with AI: humans are wrong more often than these systems are, and humans lie more too. They embellish, they bluff, they tell you the project is on track when it is not, they present a guess with the confidence of a fact because admitting uncertainty feels like weakness in a meeting.

We cope with this remarkably well, and we barely notice we are doing it. Two hundred thousand years of evolution built us an extraordinary detection apparatus for exactly this problem. The pause before an answer. The eye contact that breaks. The hedge in the voice. The track record we keep on every colleague without writing anything down. The read on someone's incentives before we weigh their claim. Gut feel is not magic. It is pattern recognition running on behavioural data, and it is very good.

None of it works on a machine.

Wrong without tells

When an AI system is wrong, it is wrong fluently. The incorrect answer arrives in the same confident, well-structured prose as the correct one. There is no hesitation to notice, no body language to read, no nervous energy, no incentive to interrogate. Every signal our detection apparatus relies on is absent, and worse, the signals that are present all read as competence. Fluency, structure, confidence, speed: everything about a polished AI answer pings the parts of our brain that say “this person knows what they are talking about.”

So we end up in a strange inversion. A source that is, in my experience, wrong less often than the average human colleague sails its errors straight past defences that were never built for it. Meanwhile we instinctively double-check the intern, who at least has the decency to look unsure when they are unsure.

This is why I have come to think the leadership conversation about AI and critical thinking is slightly misframed. The usual framing is “trust it less.” The real task is different: your trust calibration system has no jurisdiction here, and it needs rebuilding on new inputs.

The pace problem

There is a second pressure stacked on top. AI collapses production time. Work that took a team a week arrives in minutes. Leaders celebrate this as productivity, and it is, but notice what it does to the critical thinking load. The thinking did not disappear. It concentrated.

When a report took a week to produce, your evaluation of it was spread across the week: conversations, drafts, questions, the ambient quality signals of watching people work. When the report arrives in four minutes, all of that evaluation compresses into the moment you read it, with none of the ambient signal, at a volume of output per day that no leader has ever had to assess before.

Reviewing has become the bottleneck skill. Most organisations are still training people to produce, and almost nobody is training people to evaluate at machine pace. The leaders who thrive in the next decade will not be the ones who generate the most with AI. They will be the ones who can tell, quickly and reliably, which of the generated things are true.

What recalibration actually looks like

If behavioural signals are gone, you move to structural ones. In practice this has meant a few disciplines for me.

Know the constraint map. AI systems have a texture to their failures. They degrade on recency, on hyper-specific niche detail, on complex arithmetic buried in prose, on anything they cannot verify but feel compelled to complete. Learning where the terrain gets soft is the machine equivalent of knowing which colleague oversells. This knowledge is not optional for leaders anymore. Delegating to a system whose failure modes you cannot name is not delegation, it is abdication.

Interrogate the claim, not the claimant. With humans we assess the messenger. With machines you must assess the message: ask for the basis, check the load-bearing facts, follow one citation all the way down. Not every fact, that would erase the productivity gain. The load-bearing ones, the facts your decision actually stands on.

Treat confidence as noise. This is the single hardest habit, because it runs against every social instinct we have. The tone of an AI answer carries no information about its reliability. Read the substance as if it had been delivered in a monotone.

Ask what wrong would look like. Before accepting an output, spend thirty seconds on one question: if this were subtly incorrect, where would the error most likely be hiding? It is remarkable how often the answer directs your attention to exactly the paragraph that needed it.

The leadership multiplier

This matters more for leaders than for anyone else, for a simple reason of scale. An individual contributor who accepts a wrong AI output ships one flawed piece of work. A leader who does it ships a flawed decision, at organisational scale, at machine speed, and then models for everyone watching that speed matters more than verification.

The reverse is also true. A leader who visibly interrogates AI output, who asks “what is this based on” in meetings, who praises the person who caught the confident error, is training an organisational immune system. Critical thinking used to be a nice-to-have on a competency framework. It is becoming the difference between organisations that compound AI's advantages and organisations that compound its mistakes.

I will end where I started, with the uncomfortable part. I trust AI output on matters of fact more than I trust the average unverified human claim, and I check it more, not less. Both of those things are correct at the same time. The checking is not distrust. It is the recognition that my gut, magnificent instrument that it is, was trained on the wrong species.

About the author

Sean is the founder of BlynkAudit (blynkaudit.com), a platform that scores how ready businesses are for the agentic economy. He spent 20+ years in retail, most recently leading consumer insights and enterprise AI adoption at IKEA Australia.

See yourself the way an agent does

BlynkAudit scans your site the way a shopping agent will and scores what it finds. Start with a free scan, or see how your market ranks.