Insights · April 26th, 2026
The most useful piece of AI research published this year isn’t about a new model. It’s about what happens when you put one of today’s models on a team and let it work alongside people.
A March 2025 MIT field experiment ran 1,258 human-AI teams through a real marketing task, then spent real money running the resulting ads on X. The findings reframe the question CEOs should be asking. It’s no longer “how much faster will my people be with AI?” It’s “what does the work itself look like when half the team is software?”
The study, Collaborating with AI Agents: A Field Experiment on Teamwork, Productivity, and Performance, was conducted by Harang Ju and Sinan Aral at the MIT Initiative on the Digital Economy.
The team built a custom collaboration platform called MindMeld — think Google Docs plus a chat panel, where two collaborators can edit copy, select images, and generate new images via DALL-E 3 in real time. They then recruited 2,310 U.S. participants from Prolific, a research platform, and randomly assigned them to one of two conditions: paired with another human, or paired with an AI agent built on GPT-4o. Participants didn’t know which.
The task was real. Teams had 40 minutes to produce display ads for a think tank’s annual report. Both partners — human or AI — could send messages, edit headlines and body copy, pick from a stock image library, or generate new images. Critically, the AI was an agent, not a chatbot. It could take independent actions on the shared workspace every ten seconds, not just respond to prompts.
The output volume is what makes this study unusual: 11,138 ads, 183,691 messages, nearly two million text edits, and 63,656 image edits, all logged with timestamps. The researchers then took 2,000 of those ads, ran them as paid campaigns on X over twenty days, and measured 4.9 million impressions worth of click-through rates, cost-per-click, and post-click engagement on the actual document. To my knowledge, no published study of generative AI productivity has combined this scale of granular collaboration data with this kind of real-money field validation.
A second layer of the experiment randomized the AI’s “personality” — using prompts to induce high or low scores on each of the Big Five traits (openness, conscientiousness, extraversion, agreeableness, neuroticism). This let the researchers test whether matching AI personality to human personality changes outcomes.
Why It Matters
Three findings stand out, and they cut against some popular narratives.
First, AI agents don’t just make individuals faster — they redistribute where the work happens. People paired with an AI agent sent 137% more messages than human-human teams overall, but spent 20% less time directly editing text. The AI did the typing; the human did the directing. Communication shifted from social and emotional (“nice catch,” “ha, true”) toward content and process (“try a shorter headline,” “use the second image”). Human-AI teams sent 23% fewer social messages. The mechanism behind the productivity gain isn’t speed. It’s the elimination of what the authors call social coordination cost.
Second, the productivity gain is real but lopsided in a specific way. Each person on a human-AI team produced 60% more output than each person on a human-human team. Two human-AI teams of one person each match the output of one two-person human team — with comparable quality on the metric that mattered most in the field experiment. Click-through rates and cost-per-click on the live ad campaigns were statistically indistinguishable between the two team types. The headcount math is striking, but only if you also accept the second part of the next finding.
Third, AI agents are uneven across modalities, and the unevenness matters. Human evaluators rated the text quality of human-AI ads higher than human-human ads. They rated the image quality lower. GPT-4o is excellent at writing copy and mediocre at judging whether an image will land. In the field experiment, text quality drove higher click-through rates; image quality drove lower cost per click. So the human-AI teams won on engagement but paid more per click than they should have. The ads performed equivalently overall, but for different reasons on each side.
The personality-matching finding is genuinely novel and genuinely messy. Some pairings helped (a conscientious human plus an open-minded AI produced better images). Some hurt (an extroverted human plus a conscientious AI degraded text, image, and click quality). The signal is real but the rule isn’t simple, and anyone selling you a “personality optimizer” today is ahead of the science.
What This Means for CEOs
The temptation is to read “60% more output per worker” and start modeling headcount reductions. That’s the wrong first move. The right first move is to recognize that this study shows AI agents changing the shape of work, not just its volume. Five concrete actions follow.
1. Pilot agents in workflows, not as chatbots. Most enterprise AI deployments today bolt a chat window next to existing tools. The MIT result depends on the AI being inside the workspace — editing the same document the human edits, taking actions, adapting to what the human just did. If your current AI rollout is a sidebar, you are measuring a different (and weaker) intervention than this study did. Within 90 days, identify one workflow where an agent could share an editable artifact with a human and start there.
2. Audit your AI for the multimodal gap. The single most important practical finding is that GPT-class models excel at text and underperform on images. If your team’s output is multimodal — marketing, product design, sales collateral, internal comms with visuals — pairing GPT with a specialized image-evaluation or generation model is not optional. Ask your CTO: “where in our stack are we using a text-first model to evaluate non-text output, and what’s the second-tool plan?”
3. Resist the headcount math, at least for one cycle. The productivity finding is per-task and per-individual, in a 40-minute lab task with no career consequences. Real teams have onboarding, institutional knowledge, client relationships, and the kind of tacit coordination that doesn’t show up in a Prolific study. Treat the 60% number as evidence that capacity is opening up, not as evidence that 40% of your headcount is redundant. Use the slack for higher-quality output before you use it for lower payroll.
4. Rewrite the questions you ask vendors. Vendor pitches will increasingly cite this paper or studies like it. The red flag is anyone claiming AI agents work “regardless of context” or pitching personality matching as a solved feature. Ask: “Show me the modalities where your agent underperforms. Show me the field-experiment evidence, not just lab benchmarks. What happens when our specific user type interacts with your specific configuration?” If the answer is hand-waving, you’re being sold.
5. Watch the social fabric. Human-human teams in the study did something the human-AI teams didn’t: they built rapport, made jokes, expressed concern, apologized. That work isn’t measured in click-through rates, but it’s how teams stay viable across years, not minutes. As agents absorb more of the task-coordination layer, the social and emotional layer becomes scarcer and more valuable. The CEOs who will run the best teams in 2027 are the ones thinking now about what humans should be doing more of as agents take on the mechanical work, not just less.
The one question to put to your senior AI leader this week: “If we paired one of our people with a capable agent on our most ad-creation-like task, what would the workflow look like, and what would we be set up to measure?” If they can’t answer concretely, you don’t yet have a strategy — you have a budget.
Bottom Line
AI agents don’t just speed up workers; they change which parts of the work humans do, and the gain comes from eliminating coordination overhead, not from raw machine speed. The CEOs who win the next two years will be those who redesign workflows around that fact, while staying clear-eyed about where today’s agents — particularly on visual judgment — still need a human in the seat.
Read ‘Collaborating with AI Agents: AField Experiment on Teamwork, Productivity, and Performance’ – here
Other articles in ‘The CEO’s guide to AI’ series
Nikolas carefully scans and curates worthwhile research for the ‘The CEO’s guide to AI’ series – read more below:
About Nikolas Badminton
Nikolas Badminton is the Chief Futurist & Hope Engineer at futurist.com. He’s a world-renowned futurist keynote speaker, consultant, author, media producer, and executive advisor that has spoken to, and worked with, over 500 of the world’s most impactful organizations and governments.
Nikolas is an artificial intelligence expert and his 2026 keynote ‘The AI Leader: Create Incredible Productivity, Profit & Growth’ is the level up for the modern CEO and executive leader.
Please contact futurist speaker and consultant Nikolas Badminton to discuss your engagement.