Home  /  Blog  /  Artificial Intelligence
Artificial Intelligence

Generative AI Tools: Experts View and Results (Part 4)

July 17, 2023Artificial Intelligence2 min read
Table of contents
  1. Key takeaways
  2. Three methods, three answers
  3. Side-by-side prompt tests

Part 4 of 4, the payoff. Three research methods, three different winners. That contradiction is the most useful finding in the whole study.

Key lesson (2023 snapshot)

The big takeaway is methodological: user research, expert heuristic analysis and subjective expert ratings each crowned a different tool, and that disagreement is valuable data, not noise. User research ranked Leonardo first; heuristic analysis (Shneiderman’s golden rules) rated Firefly and Bing best and Midjourney worst; yet on subjective overall experience the same experts named Midjourney best. No single method tells the truth alone, triangulation does. Tool specifics are a July 2023 snapshot and have since changed.

Key takeaways

  • Heuristic analysis finds what's wrong (not what's right) and is run by an outside expert panel with no stake.
  • Three methods, three winners: user research (Leonardo), heuristics (Firefly/Bing), subjective (Midjourney).
  • The contradiction between methods is itself a finding: one method is never the whole picture.
  • Side-by-side prompt tests: Midjourney had the best prompt adherence; Bing over-compresses images.
  • Ethical/cultural flags: Leonardo appeared to reproduce stock-photo logos; 'beauty' prompts showed cultural bias.

Three methods, three answers

After the user-research results (Part 3), the team added two more lenses: a heuristic analysis by a three-person expert panel using Shneiderman’s eight Golden Rules (scored 1-5 via the Heurio app), and side-by-side prompt comparisons. The results diverged sharply:

ToolUser research (P3)Heuristics (expert avg)Subjective (Likert, /5)Author's verdict (/10)
Midjourney2ndworst (~1.75)best (4.67)8
Leonardo.AI1st~3.43.677
Bing5th~3.5-4 (high)1.335
Kaiber3rd~2.252.334
Adobe Firefly4th~3.6 (high)2.333
The same five tools ranked very differently by each method (July 2023).

The same experts who gave Midjourney the worst heuristic scores subjectively ranked it the best tool. Even that contradiction is valuable data.

On why one method is never enough

Why the split? Midjourney’s Discord-only, command-memorization interface tanks its heuristic score (poor consistency, memory load, error reversal), but its model output is so strong that experts still rated the overall experience highest. Firefly and Bing, backed by big UX teams, score well on heuristics but underwhelm on results. Leonardo is the balanced free option.

Side-by-side prompt tests

Using identical prompts (Kaiber excluded, since it can’t make a single image), Midjourney consistently gave the closest prompt adherence and best quality; Bing adhered reasonably but over-compresses images to save bandwidth (cardboard-on-blurry-background look); Firefly produced the weirdest faces/anatomy; Leonardo sometimes ignored the prompt and, notably, reproduced a legible stock-photo company logo, raising AI-ethics questions about its training data.

Cultural UX observation

Across many generations, Midjourney tended to render ‘beauty’ as a Caucasian person while Leonardo tended toward Asian features. A reminder that training data and company origin can encode cultural bias into ‘neutral’ prompts, a Cultural UX concern worth auditing in any generative tool.

What is heuristic analysis in UX?

An expert-panel method that identifies what’s wrong with an interface against a rule set (here, Shneiderman’s Golden Rules), scored and averaged. It finds problems, not strengths.

Why did the three methods disagree?

Each measures something different: user research captures lived use, heuristics catch rule violations, subjective ratings capture overall satisfaction. Midjourney scored worst on heuristics but best subjectively because its output outweighs its clunky interface.

Which tool did the author rate highest?

Midjourney (8/10), then Leonardo (7/10), Bing (5/10), Kaiber (4/10) and Adobe Firefly (3/10), as of 2023.

Do AI image tools have cultural bias?

They can. In these tests, ‘beauty’ prompts skewed Caucasian in Midjourney and Asian in Leonardo, reflecting training data and a Cultural UX concern.

Should you trust a single UX research method?

No. The study’s main lesson is triangulation: combine user research, expert heuristics and subjective ratings, because each alone can mislead.

Want research that triangulates multiple methods, not just one? Talk to Dorve.

Start a project

Want a team that knows the theory?

We research, design and build with people who can tell when an answer is wrong. Tell us what you are working on.

Start a project