Collective intelligence is real, and it is contested
Everyone in the room was chosen carefully. The meeting still didn’t move. What fifteen years of measurement on group performance actually establishes — and what it doesn’t.

Look at your leadership team on paper.
Every person on it was chosen. Someone read the CV, took the references, ran four rounds of interview and decided this one was better than the others. Most have been promoted at least once by people who watched them work up close. If you were asked to defend the quality of the individuals, you could do it for an hour without repeating yourself.
Now think about Thursday’s meeting.
It started on time. Everyone had read the pre-read. Nobody was difficult, nobody was underprepared, nobody dominated. And what the room produced in the hour it had was not what you would have predicted from the CVs.
Sometimes the shape is a sync that ends on schedule with nothing actually settled. Sometimes it is an offsite that runs three hours, examines the problem from every angle and closes without a direction chosen. Different shapes, one condition — in the weekly rhythms, and in the rooms where the decisions are largest. Same people either way.
When it happens often enough, the reflex is to look at the people. It is the most available explanation and the only variable a leader has real control over. Someone is not senior enough for this room. Someone is not being straight about what they think. We need a different Chief Product Officer.
There is a body of measurement on this, running since 2010. It does not support treating that reflex as the whole explanation. It is also more contested than most of the people citing it let on, and the contest is worth having in hand before you lean on the finding.
What was actually measured
In 2010, a group of researchers led by Anita Woolley published a result in Science that started the line. Across two studies, 699 people worked in groups of two to five — not intact teams from an organisation, but recruited participants, assembled for the study, meeting each other for the first time.
Each group spent several hours on a deliberately varied battery of tasks: brainstorming, visual puzzles, moral reasoning over a case, planning under constraints, negotiating over limited resources. The tasks had little in common with each other.
Performance on them did. Groups that did well on one kind of problem tended to do well on the others, and the ones that struggled struggled across the board. The researchers extracted a common factor from that pattern and called it c, for collective intelligence, deliberately echoing the g factor from a century of individual intelligence research. In those original studies, the factor accounted for roughly forty per cent of the variation across task scores.
Then they tested whether it travelled. They measured c on the battery and used it to predict how the same groups would do on a task they had not already completed — a complex architectural design problem in one study, a game of checkers against a computer program in the other.
It predicted — and it predicted better than the average intelligence of the members or the intelligence of the single most able person among them. Neither of those individual measures was strongly related to the group factor, and neither relationship reached statistical significance.
That is the finding that made the paper matter. What the group produced was not simply falling out of the most obvious measures of the people inside it.
The paper also reported correlates — groups did better when conversational turns were distributed more evenly, and when members scored higher on a test of reading emotional state from photographs of eyes. Whether those are things you could change in order to change the result, or things that merely travel alongside it, is a separate question with its own long and instructive history. It is not this article’s.
What fifteen years added
The 2010 result could have been a one-lab finding. It was not.
In 2021, Christoph Riedl, Young Ji Kim, Pranav Gupta, Thomas Malone and Anita Woolley published a meta-analysis in PNAS pooling 22 studies: 1,356 groups, 5,279 individuals. The population stayed broadly the same — groups working standardised tasks in lab and online settings, rather than intact leadership teams working their own problems.
The analysis sorted the predictors into two families. Composition: who is in the group, their skills, their measured ability. Collaboration process: what happens as the group works.
Process came out ahead of composition. Not as a footnote — as the headline.
Individual ability still matters in that result. Who is present still matters. What the finding says is that those variables do not exhaust the explanation once people start working on something together.
One number from that paper needs care, because it is the most-quoted figure in the field. The paper initially reported an average variance extracted of about 44 per cent for its collective-intelligence factor; a 2022 correction put it at 19.6 per cent. Both numbers get misused the same way. Average variance extracted is a statistic about how well a measurement model holds together — not an estimate of how much real-world team performance collective intelligence explains. The one-factor conclusion survived the correction. The headline number should never have been a headline.
The case against, at full strength
Two objections are serious enough that this literature should not be used without them.
The factor may partly be an artefact of the battery. In 2017, Marcus Credé and Garett Howardson reanalysed the original data and challenged whether a single general factor was the best description of what those groups were doing. Their argument is methodological: a common factor tends to emerge from any battery of tasks that shares setting and method, so the burden is on showing it beats the alternatives — and on their reading, it did not clearly do so. If they are right, part of what gets called c reflects how the tasks were assembled rather than a stable ability of the group.
The factor may be member intelligence wearing a different name. Also in 2017, Timothy Bates and Shivani Gupta ran their own studies with several hundred participants in small groups. They found a group-level factor too. But in their data the measured intelligence of the members accounted for that factor almost in full, and the additional role of social sensitivity did not reproduce. Their conclusion is in the title: smart groups are groups of smart people.
Take that at full weight rather than waving past it. If Bates and Gupta are right, the reflex at the top of this article is not a bias to be corrected. Composition is not one contributor among several. It is the dominant explanation, and the sensible thing to do about a room that does not move is to change who is in it.
Fifteen years in, this is not settled — and the sequence matters. Both 2017 objections predate the 22-study meta-analysis, which pooled far more groups than any single experiment and still put collaboration process ahead of composition. It did not settle the question their way. Whether c is a distinct group-level ability, how far it overlaps with member intelligence, and what actually produces it are all live. Anyone who tells you the literature has closed those questions has picked a side of an argument the literature is still having.
What survives the dispute
Notice what the two sides share.
Both treat what the group produces as the thing to be explained, and both had to measure it at group level to say anything about it at all. Woolley and colleagues ran groups. Credé and Howardson reanalysed groups. Bates and Gupta put people into groups and measured what came out. Even the strongest composition result is a claim about a group-level quantity — that it is predicted well by who is present, not that it stops existing.
The unit of analysis survives the argument intact. What is contested is what fills it.
That is not how the question gets handled in practice, though. Organisations do not treat composition as contested. They do not treat it as a question. Hiring is measured. Promotion is measured. Individual performance is reviewed twice a year against written criteria by someone trained to do it. In every one of those systems the unit is the person — while every result in this literature, on both sides of the argument, is a result about groups.
Thursday
Here is what the measurement supports, at the strength it carries.
There is repeatable structure in how groups perform across unrelated tasks, it predicts performance on tasks the group has not yet done, and the broad phenomenon has been reproduced across more than a thousand groups. There is real evidence that what happens during collaboration contributes to the result. There is also a serious, unresolved dispute about how much of it comes down to the abilities of the people involved.
And almost all of it was measured on assembled groups doing standardised tasks — a population with very little in common with a leadership team that has history, hierarchy and a Monday to get to. That limit is real, and it deserves an article of its own.
What survives is narrower than the usual headline and more useful. The question is not only who is in the room. It is what the room is doing — and fifteen years of argument have not produced an answer that can be reached without looking there.
Which is the harder question, because by the time everyone sits down, the roster is settled and the result is not. You chose those people months or years ago.
Thursday is on the calendar again this week. Same roster.
Watch what the room does.
References
Woolley, A. W., Chabris, C. F., Pentland, A., Hashmi, N., & Malone, T. W. (2010). Evidence for a collective intelligence factor in the performance of human groups. Science, 330(6004), 686–688. 10.1126/science.1193147
Riedl, C., Kim, Y. J., Gupta, P., Malone, T. W., & Woolley, A. W. (2021). Quantifying collective intelligence in human groups. PNAS, 118(21), e2005737118. 10.1073/pnas.2005737118. Correction: PNAS (2022), 119(19), e2204380119. 10.1073/pnas.2204380119
Credé, M., & Howardson, G. (2017). The structure of group task performance — A second look at “collective intelligence”: Comment on Woolley et al. (2010). Journal of Applied Psychology, 102(10), 1483–1492. 10.1037/apl0000176
Bates, T. C., & Gupta, S. (2017). Smart groups of smart people: Evidence for IQ as the origin of collective intelligence in the performance of human groups. Intelligence, 60, 46–56. 10.1016/j.intell.2016.11.004