Essays

Delegating to machines that talk back

Four field studies — a call center, a writers' desk, a consulting floor, and a software team — converge on the same finding: raw access to AI barely predicts who gets ahead. What predicts it is whether someone built the habit of checking the machine's work, and learned exactly where to stop trusting it.

Ask any owner of a five-person marketing shop, a two-partner law office, or a boutique architecture studio what changed this year, and the answer arrives fast: everyone delegates to something that talks back now. A draft brief goes out to a chatbot before it goes to a client. A first-pass contract review runs through a model before a paralegal sees it. A junior hire’s rough sketch gets a second opinion from software that has read more building code than any one person could in a lifetime. The tool itself is no longer the interesting part of the story. What is still being worked out, quietly, inside thousands of small practices, is the discipline of using it well.

The tool that made novices better

The clearest early evidence came out of a call center, not a boardroom. Economists Erik Brynjolfsson, Danielle Li, and Lindsey Raymond tracked 5,179 customer support agents at a software firm as a generative AI assistant was rolled out to some teams before others. The headline number was modest: issues resolved per hour rose 14 percent on average. The number underneath it was not modest at all. Novice and low-skilled agents improved by 34 percent, while the most experienced agents barely moved. The tool, the researchers found, was disseminating the best practices of the ablest workers, compressing years of on-the-job learning into weeks for everyone else.

A similar pattern turned up among white-collar professionals doing nothing more exotic than writing. Shakked Noy and Whitney Zhang gave 453 college-educated marketers, grant writers, and HR professionals real writing tasks drawn from their own occupations, then randomly gave half of them access to ChatGPT. Average time on task fell 40 percent; output quality, graded blind, rose 18 percent. The gain was not spread evenly. Weaker writers improved the most and the strongest gained almost nothing. But the more revealing finding was what the tool did to the shape of the work itself: professionals largely stopped producing rough drafts by hand and shifted their saved time toward idea generation and editing. The machine wrote. The human checked. That division of labor turns out to be most of the story.

Where the frontier gets jagged

Two years into the boom, a second wave of research stopped asking whether AI helps and started asking exactly where it stops helping. The starkest answer came from a study of 758 consultants at Boston Consulting Group, run by a team spanning Harvard Business School, Wharton, and MIT. Given eighteen realistic consulting tasks, consultants with GPT-4 access completed 12.2 percent more of them, worked 25.1 percent faster, and produced work rated more than 40 percent higher in quality — but only on tasks that sat inside what the researchers named AI’s “jagged frontier,” a boundary of competence that behaves less like a straight line and more like a coastline: generous in some coves, treacherous just around the point.

Human consultants got the problem right 84 percent of the time without AI help, but when consultants used the AI, they did worse — only getting it right 60 to 70 percent of the time.Ethan Mollick, on the jagged-frontier findings

That single result may be the most useful sentence in the whole literature for a small practice owner, because it inverts the intuition most people carry into these tools. Enthusiasm is not the risk. Trust in the wrong place is. The same researchers found two workable ways of managing that risk, which they nicknamed Centaurs and Cyborgs. Centaurs draw a clean line — this part is mine, that part is the machine’s — and hold it. Cyborgs blend continuously, letting the model finish their sentences and correcting it mid-stream. Neither posture is more virtuous than the other. What both share, and what the consultants who fared worse on out-of-frontier tasks lacked, is a working theory of where their own line actually sits.

Adoption is not the same as trust

It would be convenient if the fix were simply to roll the tool out faster. The evidence says otherwise. In three randomized trials spanning 4,867 software developers at Microsoft, Accenture, and a Fortune 100 electronics manufacturer, access to GitHub Copilot lifted completed weekly tasks by 26 percent on average — junior developers gained 27 to 39 percent, senior developers only 8 to 13. Yet a year after rollout, adoption across the three companies still sat at only about 60 percent, even with the tool free, fast, and demonstrably useful. The researchers were candid, too, about a gap in their own data: they could measure how much code got written, not whether it was good code, because they never had access to it. Even inside companies with a formal engineering review process, quality remained the harder thing to verify.

Study Setting Sample Headline result
Brynjolfsson, Li & Raymond, NBER (2023) Customer support agents 5,179 agents +14% avg. productivity; +34% for novices
Noy & Zhang, Science (2023) Professional writing tasks 453 professionals −40% time on task; +18% quality
Dell’Acqua et al., HBS/BCG (2023) Consulting tasks, inside the frontier 758 consultants +12.2% tasks done; +25.1% speed; >40% quality
Cui et al., Microsoft/Accenture/Fortune 100 (2025) Software development, Copilot 4,867 developers +26% completed tasks; ~60% adoption at one year

Small practices do not have a review department to fall back on. There is no compliance desk to catch a fabricated case citation before it reaches a filing, no second reader to flag a client email that quietly overstates a claim. The procedure has to be invented by the owner, deliberately, before the tool is trusted with anything client-facing — and the studies above suggest roughly what that procedure should contain. It should identify, task by task, which parts of the practice’s own work sit inside its frontier and which do not, because that boundary is domain-specific and nobody has mapped it in advance. It should preserve a human editing pass on anything that leaves the building, the way Noy and Zhang’s professionals shifted their saved time into review rather than banking it as leisure. And it should stay suspicious of fluency, since an answer that reads confidently is not the same as an answer that is correct, and the consultants who trusted a fluent wrong answer did worse than the ones working from nothing but their own judgment.

What actually compounds

None of this is an argument against delegation. The customer-support agents, the writers, the developers who used these tools well came out ahead by wide and repeated margins, and the practices that never touch the technology will simply do less, more slowly, at a higher cost per hour of expertise. The argument is narrower, and for a small operation, more actionable: the gains above did not come from installing a chatbot. They came from a worker, or a firm, that built a habit of checking — dividing labor like a Centaur, blending like a Cyborg, or simply learning by trial which categories of task the machine could be trusted with unsupervised. That habit takes longer to build than a subscription takes to buy, which is exactly why it compounds. A practice that spends its first months with a new model mapping its own jagged frontier will still be running on that map long after the model itself has been replaced by the next one.