Good Discovery When AI Does the Talking
The teams winning at AI-moderated discovery pair automation with real judgment

Fifty customer interviews used to take three weeks: recruiting participants, scheduling calls, then waiting on a research team to synthesize the transcripts. That timeline shaped how product teams approached discovery for most of my career.
An AI moderator can now run high volumes of structured interviews without the scheduling bottleneck that used to cap sample size. Nielsen Norman Group tested two AI interview tools in January 2026 and found they capably handle structured, known-question interviews. That’s an independent study, not a vendor describing its own product. Vendors report even bigger numbers. Perspective AI’s review of roughly 300 product teams found a three-week research cycle compressing to about three days. Worth knowing, but a vendor’s account of its own customers is a different kind of evidence than NN/g’s controlled test.
The speed is real either way. Where the good teams and the bad teams stop looking the same is what happens after the interview ends.
The Bottleneck That Disappeared
For most of my career, qualitative research meant choosing.
Five interviews or fifteen, never fifty, because a human researcher can only schedule and moderate so many conversations before the roadmap decision they were meant to inform has already been made without them. That constraint shaped how product teams approached discovery. Teams picked their highest-risk assumption and tested that one, because testing all of them wasn’t affordable (budget or time).
AI-moderated interviews remove much of the scheduling and synthesis cost, not all of it and not the interview itself. An agent opens the conversation, asks the seed questions, and follows up on vague answers the way a trained researcher would. If someone says pricing felt confusing, it asks what specifically felt confusing instead of moving on. Recruiting participants, managing incentives, and designing the research protocol in the first place still take real work and real human expertise.
What’s gone is the multi-week wait between “we have a question” and “we have transcripts.”
Strella, the research platform used by Duolingo, Amazon, and Daily Harvest, ran a video concept test for Duolingo that took two days instead of the more than six weeks a comparable study had taken before, according to Bessemer’s case study on the company. That’s Strella’s own reported result, not an independent benchmark, but it’s a useful marker for how fast discovery can move once the scheduling bottleneck is gone.
Fast discovery isn’t automatically good discovery, though. It’s just fast.
I wrote in December that primary research stays human, no AI shortcuts. That wasn’t wrong, but it answered a narrower question than it read like it was answering. What I’d tested were general-purpose assistants drafting questions from outside the room, not purpose-built tools running the interview itself. A lot’s changed in 8 months with the technology (that’s a good thing, but also keeps us on our toes and open to change).
What Changed Between December and July
The tools tested changed. The line between where AI helps and where it doesn't held.
What Transcripts Miss
The teams getting the least out of this shift are the ones treating transcript volume as a proxy for understanding. Fifty AI-moderated interviews produce fifty structured summaries, and it’s tempting to read the summaries, count the recurring themes, and call that discovery. That’s synthesis, not insight. It tells you what customers said. It doesn’t tell you what they meant.
Imagine a typical tax software customer. Ask what they want and they’ll say simplicity. Dig one layer past that and what they’re actually describing is relief from the fear of an audit, a need “simplify the interface” doesn’t capture and “add more tooltips” won’t fix.
That technique has a name. Laddering means asking why enough times to climb from a stated preference to the actual value or fear driving it, the correlation-versus-causation trap applied to interviews. No amount of theme-counting across fifty transcripts surfaces that gap if the interview protocol never asked a question built to find it.
I expected AI moderators to flatten emotional nuance, the inarticulate frustration that shows up in a pause before someone answers; however, what I’ve seen complicates that. Bessemer reports that Strella’s customers find participants more candid with an AI moderator than a human one, because the social pressure to please an interviewer, or worry about sounding foolish, mostly disappears. That’s a vendor’s reported experience, not settled research, but it undercuts the easy version of this argument, the one where humans get the real answer and AI gets the polite one.
The actual gap isn’t between human and AI moderators. It’s between protocols built to chase the “why” behind an answer and protocols that stop at the first coherent-sounding response.
Three Ways to Run Discovery With AI in the Room
Based on NN/g's 2026 study and the Perspective AI / Strella case studies cited above
Choosing The Right Option
The product teams doing this well aren’t choosing between AI-moderated interviews and human judgment. They’re deciding where each one belongs. Recurring, high-volume questions, like churn diagnostics after every cancellation and win-loss debriefs after every closed deal, are exactly what AI moderation is built for. The questions are known in advance, the volume matters more than any single conversation, and consistency across fifty sessions beats the deeper read one researcher might bring to five.
Novel domains work differently. When a team doesn’t yet know what question to ask, or the topic carries real emotional weight, a protocol written last quarter can’t adapt fast enough. That’s still where a human researcher earns their seat in the room, reading the pause and changing the next question based on a shift in tone. Sometimes that means deciding mid-conversation that the whole plan was wrong.
A researcher who used to spend a quarter scheduling and moderating fifty conversations can spend that same quarter reading fifty transcripts side by side and connecting a pattern from the churn interviews to something a sales rep mentioned in a win-loss debrief, building the case for a decision nobody asked the AI to make. The judgment didn’t get smaller. The raw material it has to work with got bigger than one person used to have time to gather.
Getting this split wrong in either direction costs something. Send a novel, ambiguous question to an AI moderator and the result is fifty confident-sounding summaries of a question nobody actually answered well. Send a known, recurring question to a scarce human researcher, and that’s premium craft spent on work that didn’t need it, while the questions that did need it sit in the backlog.
AI can run the interview now. It can ask the follow-up, catch the vague answer, and hand back a transcript before lunch. Deciding which two sentences in that transcript matter more than the other four hundred is a call we haven’t figured out how to hand off, and probably won’t for a while.
What’s the last product decision your team made off a research summary nobody went back and reread against the actual transcripts?







