AI-UX Score: questions & answers
Everything about the AI-UX Self-Assessment Questionnaire: what it measures, when to run it, how many responses you need, and how to read the numbers responsibly.
About the questionnaire
What is the AI-UX Self-Assessment Questionnaire?
It’s a diagnostic tool that helps organizations measure the perceived qualityof AI-powered user experiences. It captures what technical metrics can’t: how users feel about interacting with an AI system: whether they trust it, understand it, feel empowered by it, and believe it treats them fairly.
As AI becomes embedded in everything from healthcare to banking, user trust isn’t a luxury, it’s a requirement.
How does it work?
Users rate their experience with an AI system across seven key dimensions (such as trust, transparency, and fairness) using simple, statement-based prompts. Each dimension includes 3–5 Likert-scale items (e.g. “I trust the recommendations made by the system”) that capture how people perceive and interact with the AI in real-world contexts.
Once responses are collected, this tool processes and analyzes the scores to identify strengths, surface experience gaps, and guide future improvements.
What does it measure?
How users perceive the quality, reliability, and integrity of their experience with an AI system, across seven dimensions: Trust, Transparency, Agency & Control, Accountability, Fairness, Privacy & Data Practices, and Perceived Quality & Ease of Use. Together they give a holistic view of the system’s impact on user confidence, satisfaction, and ethical acceptance.
When should I run a study?
At key moments when you want to evaluate or improve the experience of an AI-powered system:
- Before or after launching a new AI feature, to gauge perceptions and readiness.
- During usability testing, to supplement qualitative insights with structured perception data.
- At regular intervals, to monitor how trust, control, and satisfaction evolve.
- After major updates, especially ones affecting automation, personalization, or decision-making.
- When adoption is low or feedback is unclear, to find hidden friction or trust barriers.
What are the pros and cons?
Pros
- Structured insight across 7 dimensions: a holistic view of user perception.
- User-centered: captures how users feel, not just what they do.
- Actionable and benchmarkable: scores you can track over time or compare across products.
- Flexible format and quick: typically under 5 minutes to complete.
Cons
- Self-reported: perceptions may not match actual behavior or system performance.
- Needs context: there are no universal “good” scores (yet).
- Not a substitute for usability testing: best used alongside interviews or behavior tracking.
- Benchmarking is still in progress: industry-wide norms are being developed.
Can other factors affect the score?
Yes, several external and contextual factors can influence a score beyond the system’s design or performance:
- User expectations and prior experience with AI.
- System maturity: early-stage tools often score lower.
- Use-case and domain sensitivity: high-stakes domains are judged more strictly.
- Cultural and demographic differences in how trust, fairness, and privacy are perceived.
- Communication and onboarding: poor introductions drag scores down.
- Device and interface constraints across platforms.
- Recent incidents or media coverage about AI in general.
Can I edit the wording of the questionnaire?
Best practice is to not change the wording of validated items: it can affect reliability, validity, and comparability. But some customization is expected: insert your product-specific reference in place of the [AI system] placeholder (e.g. “the recommendation engine” or “AutoBudget”). This helps participants relate to the questions, as long as the core intent of each item stays the same.
What are the pitfalls to watch for?
- Not diagnostic: it tells you how the experience is perceived, not exactly what to fix. Pair it with usability tests or interviews.
- Ceiling effect: high-performing systems score consistently high, making small improvements hard to detect.
- Response biases: halo effect and social desirability. Collect responses anonymously and consider randomizing question order.
- Context matters: a low score may reflect a mismatch, not a flawed system.
Sample size & reliability
How many users do I need?
There’s no single correct number: it depends on your goals, user diversity, and desired precision. As rules of thumb:
- 20–30 responses: exploratory research and early feedback; interpret cautiously.
- 50–100 responses: descriptive statistics and benchmarking across time or systems.
- 200+ responses: regression, subgroup comparisons, and psychometric validation (factor analysis, Cronbach’s alpha).
| Purpose | Minimum N |
|---|---|
| Exploratory insights | 20–30 |
| Benchmarking / internal comparison | 50–100 |
| Subgroup comparisons | 30–50 per group |
| Regression / reliability analysis | 150–200+ |
| Factor analysis / scale validation | 5–10 per item (≈115–230) |
What is a power analysis?
A statistical method for determining the sample size a study needs. It ensures you have enough statistical power, the likelihood of detecting a real effect if one exists (commonly set at 80%), without over-sampling and wasting resources. It helps you avoid underpowered studies (which miss real effects) and overpowered ones (which are needlessly costly).
How do I know my sample reflects my target users?
Define your target base (role, AI experience, industry, region) and compare your sample’s distribution against it. If key groups are over- or under-represented, consider stratified sampling during recruitment, or weighting after collection, used sparingly, as weighting adds variability.
How many responses per item or dimension?
Aim for 5–10 responses per item (≈115–230 total for the 23-item instrument) for full quantitative use. For dimension-level analysis, 30–50 responses per dimension is recommended, so composite scores reflect consistent perceptions rather than noise.
Is there a threshold below which I should be cautious?
Yes. Below 20 users, results can be volatile, especially for individual dimensions or subgroup comparisons. Use small samples (10–15) only for qualitative insight or early validation, not quantitative claims, and always report the sample size alongside your findings.
Weighting scores
Should I weight the scores? (Most of the time, no)
The questionnaire is designed to give simple average (mean) scores. Unweighted averages are usually good enough, especially if your sample roughly reflects your real user base. Weight only when your sample doesn’t represent your users well and you have reliable reference data on the real population.
Avoid weighting when:
- Your sample proportions already match your user base.
- You care more about group comparisons than overall metrics.
- You lack reliable reference data, or key groups are very small (weighting amplifies their noise).
Trade-off: weighting improves representativeness but can reduce precision. Unless you’re making high-stakes or externally-published decisions, it’s better to get the sample right up front.
How do I calculate a weighted mean?
Group-level (simpler), if you know each group’s average and its share of the real population:
Weighted Mean = (Group 1 Mean × Weight 1) + (Group 2 Mean × Weight 2) …Weights must add up to 1. Example: novices (75%, avg 70.8) and experts (25%, avg 76.9) →
(0.75 × 70.8) + (0.25 × 76.9) = 72.3Case-level (more precise): weight each response by reference proportion ÷ sample proportion, then average. Same result, but more flexible for advanced analysis.
Should I weight percentages, and how?
Percentages (e.g. the share who “Agree”) can be weighted the same way as means, because a percentage is just the mean of 0s and 1s. Weight them only when your sample is unbalanced and you have reliable reference data. Group-level formula:
Weighted % = (Group 1 % × Weight 1) + (Group 2 % × Weight 2) …Example: novices agree 60% (70% of users), experts agree 80% (30%) →
(0.7 × 60) + (0.3 × 80) = 66%Most teams can work with unweighted percentages during early discovery and internal testing.
Still have a question?
These answers are drawn from the AI-UX Self-Assessment Toolkit. For anything else (including licensing or research collaboration) reach out via jcerejo.com/contact-me.