Digital & Tech Innovation

Building directly on my previous article, I’m thrilled to share that I’ve launched a brand-new AI analysis, this time leveraging Claude! I was eager to see how its insights would stack up against my previous evaluations (Google NotebookLM and our 100% manual approach).
To keep the comparison perfectly consistent, I used the exact same dataset from 6 user tests (hands-on exploration of a clickable prototype) designed to craft a highly specialized, expert-level business tool. I repurposed the transcriptions from NotebookLM, as the web version of Claude required adapting video files into text format.

→ Analysis Duration: AI beautifully outpaces manual processing every time
When it comes to speed, AI remains unmatched—processing data 10 times faster than a human. Claude matched NotebookLM’s impressive speed of about 1 hour, though with a uniquely different distribution of effort. The initial setup took a bit longer due to the external transcription step, but the prompt delivery and final presentation generation were incredibly swift. In contrast, the manual approach required a dedicated 10 hours (including 6 hours of deep-dive video analysis).
→ Quantity: AI leaves half of the manual insights undiscovered
Claude successfully flagged 68 elements—a performance that, on the surface, looks nearly identical to the 73 elements uncovered by manual analysis. However, looking closer at Claude’s output, over 50% of the results weren’t actionable (36 elements): 23 were unclear due to awkward phrasing, 7 were repetitive duplicates, and 6 were outright errors. This left us with 32 highly valuable, actionable insights identified by Claude.
For instance, the value of this specific Claude insight remains a bit elusive: “Drawing ergonomics insufficient: shortcuts, double-click, drawing mode navigation.”
In the end, Claude performs quite similarly to NotebookLM, which also had a few discrepancies among its 40 identified elements.
→ Quality: AI provides a highly fragmented view of the data
Much like NotebookLM, Claude showed a lower detection rate for usability criteria (41%) than for utility (65%) compared to our comprehensive manual analysis. Claude actually struggled to differentiate between these two dimensions, requiring me to input custom definitions to refine its focus. Qualitatively, the insights delivered by Claude rarely hit the sweet spot—swinging from overly microscopic details to broad generalizations. The results lack clarity due to minimalist, verb-free phrasing and occasionally awkward word choices, including some English jargon.
For example, this particular output from Claude was difficult to leverage: “✔ Spatial notes anchored for traceability.”
Ultimately, Claude’s analysis feels less ready-to-use than NotebookLM’s because it doesn't prioritize the findings, opting instead to present everything in a flat bulleted list. This leaves the crucial task of prioritization entirely to the human reader.
→ Reliability: The AI is becoming delightfully subtle
Claude shines with a beautifully balanced neutrality, avoiding the slightly over-the-top, sales-like phrasing we saw with NotebookLM (which tended to use terms like “massive time savings”). After I requested a clearer view on cross-user patterns, Claude brilliantly scored each result (e.g., 4/6 users), which instantly elevates the validity of its analysis. Interestingly, I hadn't highlighted this exact score in my manual results (leaving it in the back-office analysis).
→ Synthesis: Extracting key takeaways remains an AI challenge
Claude highlighted 6 synthesis points (3 product strengths and 3 limitations), but just like the granular details, these key takeaways require extra refinement to be truly actionable.
For instance, Claude presented this key result: “Product direction and philosophy validated by all testers.” Ideally, I wanted it to paint a picture of that direction and philosophy; as a standalone summary, this phrase lacks analytical depth.
On the synthesis front, Claude favors neutral category labels, whereas NotebookLM excelled at creating an engaging narrative with compelling headlines like: “AI Assistance: Hover suggestions triumph over 100% automation.” This dynamic approach of scanning insights through engaging headings feels much more intuitive, mirroring how I structured the manual analysis.
→ Indirect Benefits: AI tools are wonderful facilitators in key areas
Claude delivered an exceptionally elegant visual layout that beautifully aligns with my organization’s branding, perfectly meeting my desire for a clean, sophisticated aesthetic—something NotebookLM struggled to achieve despite multiple prompts. Visually, I would be delighted to share Claude's generated output directly with a client, whereas NotebookLM only allowed for minor image integration. The manual method elegantly combines the best of both AI worlds, provided you invest the extra hours. Furthermore, the workflow is wonderfully flexible with AI, as it handles tasks in the background. I can easily advance the analysis in short, energized 10-minute bursts throughout the day, whereas manual analysis demands deep, uninterrupted 3-to-4-hour blocks of intense concentration.

→ The Verdict: AI tools bring brilliant strengths to the table, yet manual analysis remains the gold standard for depth and precision
For this exercise, NotebookLM alignes more closely with my manual approach to user test analysis—offering a precise, narrative-driven vision built on deep, structured insights. The main hurdles with AI tools remain the partial capture of insights compared to manual analysis, along with the occasional misinterpretations or misdirections that can crop up. As for Claude, it avoids over-interpreting, keeping things factual and flat, which requires the human eye to bring out the hierarchy. While Claude’s visual presentation is wonderfully polished, the depth of the writing still has room to grow.
As it stands today, I would find myself spending more time refining a deliverables document generated by Claude than one from NotebookLM. However, Claude is remarkably responsive to real-time prompt adjustments, whereas NotebookLM sometimes felt like it was “looping” without evolving the output. Of course, a single study is just a snapshot: as humans and machines continuously learn from each other through daily collaboration, the future of these insights looks incredibly promising.

Bastien
Sennegon PhD
Director of UX Research
Related Articles

How AI is Revolutionizing User Test Analysis (Part 1)
NotebookLM versus manual analysis across 6 user tests: duration, quantity, quality, and reliability. While AI saves valuable time, manual analysis remains key to mitigating risks. Read this insightful comparative study by Bastien Sennegon (Aktan) so you can discover how to elevate your design differentiation your design effectively.

User Testing: Reassuring Ourselves, or Truly Learning?
User testing observes real people to pinpoint exactly where they get stuck (Nielsen's rule: five users are more than enough!). Its true and exciting challenge: are we testing simply to reassure ourselves, or to pioneer new learning?

Ergonomics: tailoring the tool to the human, not the other way around!
Ergonomics shapes the tool and the task to fit human capabilities, never the other way around (human factors, Wisner). To simply blame 'human error' often overlooks and pardons a flawed design.
Our Associated Offers

User and client testing
Evaluate a service, product, or device in real-world conditions to measure its effectiveness and desirability.
Discover →

AI Empowering Operational Excellence
Design services where humans and AI collaborate effectively to maximize value.
Discover →

STARIA: The Framework for an AI That Truly Adapts
STARIA transforms AI into lasting value by revealing frictions and aligning solutions with actual work.
Discover →