Digital & Tech Innovation

A Concrete Comparison: Performance & Risks. What AI does beautifully and where it still falls short…
The Context
The playground for this head-to-head evaluation was the analysis of 6 user tests aimed at designing a highly professional digital tool for an expert audience. This complex tool was at an intermediate level of design maturity, converging on an exploratory clickable prototype.
The tests consisted of 1-hour remote sessions per user (recorded on video), conducted in either French or English. The users were autonomous with the digital prototype (via screen sharing) and were asked to perform various tasks based on a predefined usage scenario (7 functionalities to test).
The AI analysis was conducted after the manual one using Google Notebook LM—relying on a textual analysis of transcriptions alongside contextual documents.

→ Analysis Speed: AI Simply Outpaces the Manual Process
There was really no contest on this time-based dimension! The AI took just 1 hour to finalize the process, whereas the manual approach required 10 times more effort—clocking in at 10 hours (including 6 hours of video review).
→ Quantity: AI Delivers Half the Results of Manual Analysis
On this criterion, manual analysis identified 73 distinct elements compared to 40 elements identified by the AI (11 of which were only partial)—representing a 55% identification rate (excluding partial matches). Remarkably, no new elements overlooked by the manual analysis were proposed by the AI.
→ Quality: AI Offers an Uneven Analysis of Your Data
Compared to manual analysis, the AI shows a lower detection rate on usability criteria (40%) than on utility criteria (66%). This is likely due to the element of interpretation that usability criteria demand compared to what users state explicitly. Additionally, the final minutes of the recordings seemed to go completely unanalyzed (across all 6 tests). This had a major impact because the key feature of the tested tool was missed entirely simply because it was placed at the end of the scenario. On a bright note, two critical alerts were successfully identified by the AI.
→ Reliability: AI Lacks the Touch of Subtlety
Overall, the AI analysis does not grasp the subtlety of verbal feedback and tends to exaggerate points, whereas manual analysis more faithfully translates the nuances of user perception. Crucially, two of the AI's findings on major functions were highly misleading due to oversimplified shortcuts and the use of inappropriate terminology:
The AI highlighted a "massive time-saver" for the tested tool, yet this benefit remains highly hypothetical. It depends on a set of complex conditions that are difficult to implement and were not present during the test.
The AI concluded that the "automatic assistance feature is validated" following the tests. In reality, users rejected the automatic assistance feature in our mockup. While they expressed converging needs, they did so without having tested them, so we absolutely cannot call this a validation.
On a more specific note, we observed the inclusion of a single isolated case in the AI's results, which lacked inter-individual recurrence (contrary to the prompt instructions) and therefore should not have been reported.
→ Synthesis Synthesis: Highly High-Performing, Albeit Not Exhaustive
The manual analysis based its global summary on four key elements (2 validated features + 2 limitations). The AI captured three of these elements in common (2 validated features + 1 limitation), leaving one limitation unidentified. The alert missed by the AI (even though mentioned repeatedly by users) was critical to the validation process: the demo data used during the test was completely disconnected from the users' real-world data. Again, no new findings missed by the manual review were uncovered by the AI.
→ Indirect Benefits: To Each Their Own Strengths
The AI shines in its ability to automatically generate imagery to illustrate findings (with zero confidentiality or anonymity constraints), providing a fantastic boost in speed and overall appeal. Furthermore, the AI spontaneously initiates complementary research to verify concepts mentioned during tests (like alternative software, data formats, and mathematical models)—saving valuable time compared to manual analysis. On the other hand, the major strength of manual analysis lies in the deep ownership of data. This inevitably translates into an effortless ease in speaking to results and a natural ability to bring findings to life.

→ The Verdict: AI Wins on Speed, Manual Analysis Minimizes Risk
Do we want to sacrifice analytical precision to reduce timeline duration? This is a highly valuable strategic move for fast-paced projects (market windows, team availability, tight deadlines) and/or limited budgets. For other high-stakes projects, it is far better to validate and secure every element progressively throughout the design phase (where governance, complex deployment, or fierce competition are factors).
→ Rethinking Our Design Process and Mindset
By embracing AI in our workflows, we can unleash more frequent prototype/test iterations. This is exceptionally viable as AI tools excel beautifully at design prototyping, even if it requires higher user engagement (realizing some expert profiles have limited availability). Ultimately, this represents a cultural shift where we accept a slightly lower level of initial precision and embrace more immediate risk to gain spectacular design velocity on a global scale.
With AI tools evolving at lightning speed, this article will likely become obsolete very quickly…

Bastien
Sennegon PhD
Director of UX Research
Related Articles

AI Tools for Analyzing User Tests (Part 2)
I set up a head-to-head comparison between Claude, Google NotebookLM, and manual analysis for processing user tests. From time-saving efficiency to reliability and quality... what is the ultimate verdict?

User Testing: Reassuring Ourselves, or Truly Learning?
User testing observes real people to pinpoint exactly where they get stuck (Nielsen's rule: five users are more than enough!). Its true and exciting challenge: are we testing simply to reassure ourselves, or to pioneer new learning?

Ergonomics: tailoring the tool to the human, not the other way around!
Ergonomics shapes the tool and the task to fit human capabilities, never the other way around (human factors, Wisner). To simply blame 'human error' often overlooks and pardons a flawed design.
Our Associated Offers

User and client testing
Evaluate a service, product, or device in real-world conditions to measure its effectiveness and desirability.
Discover →

AI Empowering Operational Excellence
Design services where humans and AI collaborate effectively to maximize value.
Discover →

STARIA: The Framework for an AI That Truly Adapts
STARIA transforms AI into lasting value by revealing frictions and aligning solutions with actual work.
Discover →