Earlier this year I watched a senior leader lose a room in about ten minutes.

They had brought work to present to a group whose opinion mattered. The work was AI generated. They never said so, and nobody needed to run it through a detector, because the room did what rooms do: it asked questions. Why this conclusion? What sits underneath this number? What changes if the budget halves? The presenter could not answer. Not would not. Could not. They did not understand the work they were standing in front of, and everyone present worked that out at roughly the same moment. The feedback afterwards was unanimously negative, and months later that meeting is still the one people reach for when they want an example of what poor looks like.

It would be comforting to leave the story there, because that version has a moral we already like: AI-generated work is hollow, rooms can tell, standards win. And in this case the room could tell, because the work was visibly weak and the person presenting it had nothing behind the slides.

The version that should worry you is the one where the work is good.

The failure you can see and the one you can’t

Every strong team I have been part of, and every strong team I have worked with since, has had the same underlying property. A range of viewpoints, and the conditions for those viewpoints to collide safely. People disagreed about the work, out loud, without it costing them anything socially, and the output was better for it. This is one of the better evidenced ideas in management research. Amy Edmondson has spent nearly three decades showing that teams where people can take interpersonal risks learn faster and perform better, and Google’s much-quoted Project Aristotle study found psychological safety to be the strongest differentiator among its own teams. Aristotle was internal research and its methods were never published for review, so treat it as illustration rather than proof, but the illustration points the same way as the peer-reviewed work.

All of that machinery assumes something so obvious nobody thought to state it: that the viewpoints in the room arrived independently.

That assumption is now quietly false in a growing number of teams. When everyone prepares for the same meeting by consulting the same model, the room still looks diverse. Different people, different phrasing, different slides. But the positions have a common ancestor. Anil Doshi and Oliver Hauser ran an experiment, published in Science Advances in 2024, in which writers given ideas from a large language model produced individually better work while the group’s output as a whole became measurably more alike. Read that again, because it is the whole problem in one finding. Each person improves. The collective narrows. And nobody in the room can feel it happening, because from the inside, everyone believes they did their own thinking.

A related body of work, summarised this month in MIT Sloan Management Review by Léonard Boussioux and colleagues, explains part of the mechanism: AI suggestions arrive early, before the discussion starts, and people anchor on output that seems good enough. The search space closes before the meeting opens. By the time your team is debating, it is debating variations of the machine’s answer, and calling it challenge.

The leader who lost the room was the visible failure. Convergence is the invisible one. The first announces itself through ordinary questioning. The second survives ordinary questioning, because the answers are fluent and every position being questioned grew from the same root.

The fix that doesn’t work

There is an appealing response to all this, and I held it myself until I pulled on it. It goes: stop treating AI output as true. Treat it as an opinion in the room, like any other, open to professional challenge and required to justify itself through debate.

The instinct is right. The mechanism fails, for a reason that becomes obvious the moment you try it. Debate finds truth because arguing carries stakes for the arguer. Challenge a colleague and one of two useful things happens: they defend the position well and the room learns something, or they fold and the idea dies, with a small reputational cost attached that keeps the next idea honest. Challenge a language model and you get a third outcome that is useless. It agrees with you. Push from the other direction and it agrees with that instead. And when it does hold a position, its weak arguments are delivered with exactly the same fluency as its strong ones, which is precisely the tell that lets you spot thin reasoning in a human. You cannot cross-examine something that concedes on contact and blushes at nothing.

The AI cannot hold up its end of a debate. Someone else will have to.

Three rules, adoptable on a Monday

This is the point in this sort of article where the advice usually turns to mist. I have spent three articles arguing that management judgement is being quietly de-skilled, and the honest continuation is not a mindset. It is a small number of rules specific enough to be disagreed with.

First: no AI output enters a decision without a person who can hold it under questioning. Note the wording. Not a person whose name is on it. The leader in my opening story had their name on the work; attribution is cheap. Ownership means being able to explain why the work says what it says, defend it when the room pushes, and extend it when the discussion moves somewhere the model never went. If you bring machine-generated analysis to the table, it is your position now, with your stakes attached. This restores everything the debate model needed and could not supply: accountability, track record, a human who can be embarrassed. It also relocates the checking work to where it belongs, before the meeting, done by a named person, rather than diffused across a room that assumes someone else has verified it.

Second: positions before prompts. Because the convergence happens upstream, the countermeasure has to sit upstream too. For decisions that matter, have people commit an independent view, even two rough sentences, before anyone opens a model. This is not a novel invention; it is the same logic that makes good chairs collect opinions before revealing their own, applied to a louder and more patient voice than any chair. The AI’s contribution is then one input into a genuinely diverse set, rather than the terrain everything else grows on.

Third: stop paying people to skip the checking. Randy Bean, Erik Strauss and Randeep Singh made this argument in Harvard Business Review this month: on conventional productivity metrics, the employee who accepts AI output at speed looks excellent, and the one who slows down to verify assumptions and catch errors looks inefficient at exactly the moment they are adding the most value. If your measurement system rewards throughput and is silent on verification, you have built an incentive to stop exercising judgement, and you will get what you built.

None of this is anti-AI. The tools are staying, the productivity is real, and a team that refuses them will be beaten by one that does not. The argument is narrower and older than the technology. In 1951, Eric Trist and Ken Bamforth studied coal mines that had all received the same mechanisation and found the results depended on whether work was redesigned around the people or the people rearranged around the machinery. The pits that treated the technology’s default working pattern as inevitable got the worst of both. Deliberate redesign was a choice then. It is a choice now, and the window in which teams set their norms around these tools is open at the moment and will not stay open indefinitely.

The leader in my opening story lost the room in ten minutes, and in one sense that was the system working. The norms were healthy, the questions were ordinary, and the hollow work was found. Convergence gives you no such moment. There is no session that gets referenced months later as the weak one, because every session felt fine. The room was diverse, the debate was lively, and everyone agreed, in the end, with the same quiet voice they had each consulted alone the night before.

You will not notice that happening. Which is why it has to be designed against, rather than watched for.

Sources

  • Bean, R., Strauss, E. and Singh, R. (2026) ‘Performance management needs new metrics in the AI era’, Harvard Business Review, 6 July.
  • Boussioux, L., Doshi, A., Hauser, O. and Hosanagar, K. (2026) ‘The hidden cost of AI-assisted creativity’, MIT Sloan Management Review, 9 July.
  • Doshi, A.R. and Hauser, O.P. (2024) ‘Generative AI enhances individual creativity but reduces the collective diversity of novel content’, Science Advances, 10(28).
  • Edmondson, A. (1999) ‘Psychological safety and learning behavior in work teams’, Administrative Science Quarterly, 44(2), pp. 350-383.
  • Trist, E.L. and Bamforth, K.W. (1951) ‘Some social and psychological consequences of the longwall method of coal-getting’, Human Relations, 4(1), pp. 3-38.
Share This