A UX-led AI experiment
AI ships interfaces faster than we can review them. This is the question our UX team set out to answer, with a method that has proven itself for three decades.
The trusted measure
A heuristic evaluation is a fast, low-cost inspection method. An expert reviews an interface against ten established principles and ranks every issue by severity. No recruiting. No scheduling. Hours, not weeks.
Published in 1994 and refined since, these heuristics are the industry's most widely used standard for expert UX review.
Findings are ranked on a 0 to 4 severity scale. Leadership gets a number to act on.
Real problems surface before they reach engineering, at a fraction of the cost of a moderated study.
The honest tradeoffs
The challenge
Great method, real friction. So we wrote exactly what we needed into the skill:
Minutes, not days. Speed that keeps up with how fast we generate AI outputs.
A scale we can run with AI and by hand, then check the two against each other for interrater reliability.
Instructions that hold the AI back from hallucinating findings. Every output still gets a quick human review.
Findings anyone can open, understand, and act on. No researcher translation required.
The answer
We developed this skill in the first week of the working group. Designers, engineers, and product marketers were all generating AI outputs, and we needed a consistent way to measure our impact. Everyone uses the same skill, so we're all measuring usability with the same ruler.
It's not a prompt. It's a structured knowledge asset, now at version 20, that teaches Claude how to evaluate an interface: which heuristics to apply, how to score severity with fixed Frequency, Impact, and Persistence anchors, and how to deliver the findings. The result is a consistent usability score for everything we make.
Claude receives the evaluation skill plus the screens and flow context to review.
A benchmark score, then prioritized improvements ranked on the shared 0 to 4 severity scale.
The skill always prints an HTML report marked as a draft. We make it our culture that a human runs the same review to compare against the AI, or other human reviewers.
The receipts
Every experiment evaluated the exact same task with the same UX context (target user, constraints, and a workflow map): reschedule a Coterie diaper subscription delivery to an earlier date. I ran the human evaluation with the Studio, then compared my score against a different AI setup each time.
The setup. I evaluated the live Coterie website by hand in the Evaluation Studio. Claude for Chrome crawled the same live site with the skill.
The setup. Same task, but from 6 numbered screenshots instead of the live site. I evaluated them in the Studio. Claude Sonnet 5 got the same screenshots, the skill, and a detailed UX-context prompt.
The setup. My finished evaluation from Experiment 2 stayed as the human baseline. Claude Opus 4.8 got the exact same screenshots, skill, and prompt as Sonnet 5.
The setup. Same human baseline. I built a custom GPT with the same instructions as the Claude skill and gave ChatGPT 5.5 the same screenshots and prompt.
Can you pretty please conduct a heuristic evaluation of this workflow of the Coterie home page? /heuristic-evaluation Coterie is a brand creating ultra-soft, high-performing baby care products thoughtfully engineered to support baby, parent and caregivers. The site is catered to parents and caregivers, or gifters ordering diapers and managing auto-renew orders for their baby(ies). Task to review: Home page of managing / ordering diapers (and auto-renew), specifically the steps to change shipping date of coterie diapers auto-renew subscription to move to an earlier shipping date Goal: change shipping date from July 5, 2026 to June 29, 2026 Constraints: testing only the screenshots and workflow attached, not the nav, etc. I'm attaching the workflow in a PDF and the individual screenshots numbered in order 1-6. Please evaluate the workflow in this order: Screen 1 / Coterie HE 1 (dashboard) shows Ships on July 1, 2026. Screen 2 / Coterie HE 2 (dashboard) shows Ships on July 1, 2026 with hover over on "Change date" Screen 3 / Coterie HE 3 (side sheet overlay) opens a date picker on current ship date July 1, 2026 Screen 4 / Coterie HE 4 (selects earlier shipping date June 29, 2026, "Confirm button turns inactive gray to clear blue ready state" Screen 5 / Coterie HE 5 (confirmation overlay) reads "You've changed your shipping date for Amaya … June 29, 2026." Screen 6 / Coterie HE 6 has updated in real time to the correct new shipping date June 29, 2026
ChatGPT received the identical prompt without the /heuristic-evaluation skill trigger.
The AI-assisted heuristic evaluation, end to end.
The complete HTML output from an example evaluation of a single-task workflow, all ten heuristics and every finding.
Open the example AI-generated HTML report →If the AI-powered heuristic evaluation skill alone isn't enough structure, fear not. I built a Claude Artifact that guides you through everything you need to run the evaluation. Provide the UX context and click run.
Jump to the AI-powered Heuristic Evaluation Artifact →Human in the loop
The Heuristic Evaluation Studio guides any team member through all ten heuristics with the same scoring anchors the AI uses. Human judgment and AI speed, on one rubric, producing scores that can finally be compared.
Open the Human Evaluation StudioThe complete HTML output from an example evaluation of a single-task workflow, all ten heuristics and every finding, completed by a human reviewer (me).
Open the example human-evaluated HTML report →One irresistible outcome
UX quality becomes trackable across sprints. Rerun the evaluation after a change and see exactly what moved. That's the stewardship: we give AI all the context it needs, so it produces production-ready, user-friendly outputs informed by UX evidence and design principles.