Plain-English primer · Part 1 of 2
What the two working styles are, where they came from, and what they are worth to a person, a team and a business
Part 1 · Primer›Part 2 · Literature review and evidence
Two people can hold the same AI tool, face the same task, and get opposite results. The difference is not the tool and not their intelligence. It is how they divide the work. Researchers gave names to the patterns they observed, because naming them made the difference visible and teachable. Centaurs split the job cleanly between person and machine. Cyborgs blend the two into one continuous conversation. A third group hands the whole thing over. All three look identical in a usage report. Only one of them stops learning.
For three years the public argument ran on one question. Does AI make people more productive? That question turned out to be badly formed, because the answer moves more within one person than between two people.
A large experiment at Boston Consulting Group showed this clearly. Consultants using GPT-4 completed 12.2% more tasks and worked 25.1% faster, at measurably higher quality[1]. On one task, chosen because it sat just beyond what the model could handle, the same population using the same tool became 19 percentage points less likely to get the right answer[1].
The tool did not change. The task did. And nothing on the surface of the task told anyone which kind it was.
The researchers called that invisible boundary the jagged frontier[1]. Two jobs that look equally hard to a person can sit on opposite sides of it. You cannot see the edge, so you cannot decide in advance how much to trust the output.
That left one question worth asking. If people cannot see the boundary, what do the ones who cope well actually do differently? Answering that meant describing behaviour rather than counting tool access. Centaur and Cyborg are the names for the two behaviours that worked.
The metaphor arrived from chess, and the story is worth knowing because it contains the whole argument in miniature.
Garry Kasparov loses a match to IBM's Deep Blue. The popular reading is that the machine has replaced the human at the top of the game.
Kasparov tries the opposite arrangement. He plays Veselin Topalov in the first Advanced Chess match, where both players use chess engines and databases during play. The match finishes 3-3[28].
An open online event lets any human-plus-machine team enter. The winners are Steven Cramton and Zackary Stephen, two American amateurs rated 1685 and 1398, using three ordinary PCs. They beat teams led by grandmasters running far stronger hardware[29].
Kasparov drew the conclusion that still carries the framework. A weak human with a machine and a better process beat a strong computer alone, and beat a strong human with a machine and a worse process[30].
Skill mattered less than hardware. Hardware mattered less than the method for combining the two.
That is why the framework describes how people work rather than what they know or which tool they hold. It is also why it transfers out of chess so easily. The amateurs won by running several engines on each position, comparing what they said, and using their own judgement to settle disagreements. Those are ordinary knowledge-work habits.
One honest correction. Kasparov called it Advanced Chess. The word centaur was attached later by writers describing human-machine teams, and is often projected back onto the 1998 match as if he had used it at the time[28]. The practice is his. The label is not.
The original 2023 research described two styles. A December 2025 re-analysis of the conversation logs from 244 of those same consultants found three, and the third one is the reason this matters commercially[4].
Splits the job along a clean line. The person decides what to do and how to do it, then hands the machine specific, bounded pieces. Think of a cook who uses a food processor for the chopping but tastes and seasons personally.
Gets better at the subject. Highest accuracy in the study.
Blends the two into one flow. Drafts, challenges, feeds in more data, reruns, argues back when the answer conflicts with their own view. There is no clean line, only continuous exchange.
Gets better at working with AI. Roughly equal to Centaurs on persuasiveness.
Collapses the whole job into one or two prompts, pastes everything in, and ships what comes back. In the study, 44% of this group changed nothing at all, and the rest made only surface edits.
Gets better at nothing. Lowest on accuracy and persuasiveness.
The proportions are the finding. The most accurate mode is the rarest — 14% — while 27%, more than a quarter of a highly capable professional population, had quietly stopped doing the work[4].
None of this is about effort or character. The pull toward the third mode is structural. An earlier experiment gave recruiters an AI that was very accurate, and their own attention dropped in response[6]. The better the tool performs, the less reason your brain finds to stay engaged. That effect has been documented in industrial automation for forty years[21].
Your mode shapes which expertise compounds. Centaurs deepened what they already knew. Cyborgs built a new capability in directing AI while holding on to their subject knowledge. Self-Automators built neither[4].
Neither of the first two is the correct answer. They are different bets. Centaur suits work where your judgement is the product and being wrong is expensive. Cyborg suits work where volume, drafting and exploration dominate. The practical instruction is to choose deliberately and to notice when you have drifted into the third mode without deciding to.
An experiment with 791 professionals at Procter & Gamble found that one person working with AI matched the output quality of a two-person team working without it[5]. The more interesting result was the second one. Without AI, technical people proposed technical solutions and commercial people proposed commercial ones. With AI, both produced balanced answers[5]. The tool acted across the boundary between functions, which is normally the hardest boundary in a business to cross.
Participants also reported feeling better about the work than those going it alone[5]. Part of what a colleague provides is company, and the tool supplied some of it.
This is where the framework earns its place in a policy conversation. Most AI governance says some version of keep a human in the loop. All three modes satisfy that sentence. A human was in the loop in every case, including the 27% who were Self-Automators — nearly half of whom changed nothing at all[4].
A usage dashboard cannot tell the difference between a workforce getting sharper and one quietly hollowing out. Only looking at how people interact can.
The second business implication concerns measurement. A randomised trial of experienced open-source developers found they worked 19% slower with AI tools, while estimating afterwards that they had been about 20% faster[9]. Perception and reality pointed opposite ways. Any organisation running on internal satisfaction surveys is at risk of that inversion.
Six questions about how you hand work over, how you check what comes back, and what you want to get better at. Three of them are not about AI at all. There is no score to beat and no right answer — every route through gives you something to read, including the one where you say you don’t use these tools. Nothing is stored or sent anywhere.
Everything above is the constructive case. It is not the whole picture, and presenting it alone would be a sales pitch rather than an explanation.
The largest review of the field pooled 106 experiments and found that, on average, human-and-AI combinations performed worse than whichever of the two was stronger on its own[7]. That is a serious challenge to the whole idea, and it has a resolution, but the resolution needs the evidence set out properly.
Part 2 does that. It tests three propositions against 28 sources, states which survive and which do not, and explains why the two bodies of evidence only appear to disagree.
Where a source appears in both documents the reference numbers match, so the two can be read together — which is why the numbering below jumps. Part 2 carries 1 to 27 and 31; references 28 to 30 cover the chess history, which Part 2 does not, and appear here only.