Ask an AI tool to analyse a thousand open-text responses and it will give you themes. That is the pitch, and it is usually accurate. The question worth asking is what happened in between.
Four different things happen between a submission arriving and a finding appearing in a report. Someone reads it. Someone interprets it. Someone codes it. And someone analyses the result. Those four steps get collapsed into a single word — analysis — and the collapse is where the trouble starts.
Reading establishes what was said
Reading, in this sense, means establishing what is actually on the page. Every response, not the first two hundred. The ones written at length and the ones written in six words. The submission that arrived as a scanned letter and the comment left at eleven at night on a phone.
A model can hold all of it at once and notice that two hundred responses are circling the same concern in different words. This is the part of the work human attention handles worst. Fatigue is real: the four hundredth submission does not get the reading the fourth did. There is no good reason to keep doing this entirely by hand, and a reasonable argument that doing it by hand is less consistent, not more.
What a practitioner is actually doing while reading
Calling it "reading" undersells what an experienced practitioner does with 500 responses. Simultaneously, they are:
working out what each response means, including what it implies but does not say
holding the project context alongside it — the history, the last consultation, what the organisation has already committed to
recognising ambiguity rather than resolving it prematurely
deciding whether something is a new idea or another instance of an existing theme
judging whether a response belongs in one category or several
staying consistent with a decision made three hundred responses ago
thinking ahead to how all of this will need to be reported and defended
None of that is passive. They are analytical decisions, made continuously, and mostly never written down.
Coding is not the goal
Coding is where those decisions become explicit. It means choosing a frame — a set of categories that will structure everything downstream — and judging each response against it.
But nobody codes for the pleasure of attaching labels. Coding exists so that what comes after it is possible: quantifying patterns, comparing themes across groups or across time, explaining why a conclusion was reached, and producing something that survives being challenged.
That is the sequence. Reading establishes what was said. Interpretation works out what it means. Coding makes that interpretation consistent. Analysis turns it into evidence. Skip the third step and the fourth has nothing to stand on.
The real test is whether someone else would do it the same way
The question experienced qualitative researchers ask first is not whether these are the right codes. It is whether another competent person, working separately, would code the same material the same way.
That is what coding is really for. Less about labelling than about building a repeatable analytical framework, with categories defined well enough that two analysts would land in roughly the same place. Where they would not, the framework is doing less work than it appears to, and the findings are more fragile than the report suggests.
This matters because the frame encodes a view. Is a comment about construction noise about amenity, about trust in the delivery agency, or about a project that went badly three years ago? Two competent analysts can answer differently and both be defensible. What is not defensible is a frame nobody can articulate.
Where the line actually falls
So it is too simple to say that AI can read but cannot code. It can code, and it can do it well. Given a defined framework, a model will assign codes across ten thousand responses faster and more consistently than a tired human will — and consistency over volume is precisely what humans are worst at.
What does not transfer is the work around the framework: designing it, deciding what the categories are and why, refining them as the material pushes back, and governing the whole thing — recognising when the frame has stopped fitting and needs to change.
The risk was never that a model assigns labels. It is that a model asked to "find the themes" designs the framework as well, invisibly, without stating its assumptions and with nobody accountable for them. The output looks like a finding. It is a judgement nobody made on the record.
What this means in practice
This is the distinction Explore is built around. It groups comments by meaning before you have created a single theme, so you see the shape of the material before imposing a structure on it. Every verbatim stays attached to the cluster that produced it. The framework stays yours: nothing becomes a finding until a practitioner decides it is one.
It is worth being equally clear about what it cannot do.
It cannot tell you whether a cluster of angry comments about parking is a design problem, a communications failure, or the residue of a project that went badly three years ago.
It cannot weigh a small number of well-informed responses against a large number of superficial ones. Volume is not significance.
It cannot tell you which silence matters. The group that did not respond is often the most important thing about a consultation, and it is invisible in the data by definition.
It cannot be accountable. Someone has to stand behind a conclusion when a councillor or a community member challenges it, and that someone is a practitioner.
Why this matters now
Engagement is moving towards having to demonstrate influence rather than activity. The recommended evolution of the IAP2 Spectrum would oblige decision makers to communicate how input was used. Australia's draft National Environmental Standard asks for a summary showing how feedback informed the proposal. We have written separately about why that is a records problem before it is a practice problem.
Both assume the same capability: that you can trace a conclusion back to what people actually said, and explain the reasoning in between. A synthesis nobody can retrace is not evidence. It is an assertion with a chart attached.
Reading tells you what one person said. Coding is how you understand what a thousand people said together.
That is the difference between collecting feedback and producing evidence, and it is where professional practice begins. Anyone can read responses. The work starts when you have to analyse them consistently, defend the interpretation, and show how it changed a decision.
Related reading: AI vs Human Judgment, the risks of generic AI in community engagement, and what responsible AI looks like in engagement practice.
