Home  /  Blog  /  Artificial Intelligence
Artificial Intelligence

Choosing the Best AI Image Generation apps (Part 2)

July 17, 2023Artificial Intelligence2 min read
Table of contents
  1. Key takeaways
  2. Real generative AI vs filters
  3. The five apps tested
  4. The tester panel
  5. The 11 scoring variables

Part 2 of 4. What counts as 'real' AI image generation, which five apps made the cut, and the exact 11-variable rubric used to score them.

Quick answer

To separate real AI image generation from filters, one criterion did the work: can the app create completely new visuals from a prompt? Filter apps only manipulate existing images. That narrowed hundreds of ‘AI’ apps to eight, then five to test: Adobe Firefly, Bing AI Image, Leonardo, Kaiber and Midjourney (Dall-E 2 was set aside after a testing error). The method was qualitative think-aloud with five demographically varied testers, scoring 11 variables from 1 to 5. The study is independent: no affiliate links, no incentives, and the team paid for paid apps.

Key takeaways

  • The defining test of real generative AI: can it create new visuals from a prompt, not just filter an existing image?
  • Field narrowed from hundreds of self-proclaimed 'AI' apps to 8, then 5 tested.
  • Final five: Adobe Firefly, Bing AI Image, Leonardo, Kaiber, Midjourney (Dall-E 2 set aside).
  • Method: qualitative think-aloud, in-person, five varied testers, 15-minute discovery then a short tutorial.
  • Eleven variables scored 1-5, with explicit rules for missing/free/trainable features.
  • Fully independent: no affiliate links, no incentives, paid apps paid for.

Real generative AI vs filters

Many app publishers claim ‘AI’ but only apply filters no deeper than Instagram effects. The refined criterion, can it generate something entirely new from a prompt, separated the wheat from the chaff. Examples like Loopsie (background swap plus a pseudo-3D effect) and Disflow are filters, not generators. That left eight apps, trimmed to five by dropping near-duplicates in favor of the most representative options.

The five apps tested

  • Adobe Firefly
  • Bing AI Image
  • Leonardo
  • Kaiber
  • Midjourney

Dall-E 2 was intended as a sixth but excluded from scoring after two testers missed it, to keep the data consistent (a re-test was planned). The team stresses independence: no affiliation, no affiliate links, no incentives, and they paid for paid apps.

The tester panel

TesterProfileAI/design experience
Tester 1Female, 25, professional designer, good English. New to gen-AI
Tester 2Female, 29, professional designer, good English. Some AI experience
Tester 3Male, 52, professional designer, excellent English. High AI expertise
Tester 434, no design knowledge, good English. No AI/design
Tester 5Female, 40, no design knowledge, excellent English. No gen-AI

Each tester got 15 minutes of free discovery (compared against pre-existing data), then a short tutorial and a command list for apps that needed one. Testing was in-person think-aloud, in a non-English-speaking country, with participants of good-to-excellent English.

The 11 scoring variables

#VariableWhat it measures
1Image qualityQuality of the final output.
2Prompt-to-image adherenceHow faithful the output is to the text prompt.
3Image-to-image adherenceHow faithful the output is to an uploaded image.
4TrainableWhether the model can be trained, and how easily (perceived).
5WeightingWhether prompt weighting is available.
6Negative weightingWhether elements can be down-weighted or excluded.
7Special featuresZoom, pan, color modes, upscaling, models, etc.
8User expectationHow close the result was to what the user expected.
9UsabilityOverall usability, judged subjectively.
10Absolute pricePerceived absolute value.
11Relative priceWillingness to pay given the results.
The 11 variables, each scored 1 (low) to 5 (high).

Scoring rules: a missing feature scored 0, with two exceptions. If the app was free, price (Q10) scored 5; and ‘Trainable’ (Q4) scored 0 if untrainable and 5 if trainable, regardless of training quality, since testers couldn’t fully evaluate training. Business models may have changed since testing.

What counts as real AI image generation?

An app that creates completely new visuals from a prompt. Apps that only manipulate an existing image are filters, even if they use some AI internally.

Which apps were tested?

Adobe Firefly, Bing AI Image, Leonardo, Kaiber and Midjourney. Dall-E 2 was intended as a sixth but set aside after a testing error.

What research method was used?

Qualitative, in-person think-aloud with five demographically varied testers: 15 minutes of discovery, then a short tutorial, scoring 11 variables from 1 to 5.

Was the study independent?

Yes. No affiliations, no affiliate links, no incentives, and the team paid for any apps that required payment.

How were missing or free features scored?

Missing features scored 0, except price scored 5 if the app was free, and ‘trainable’ scored 0 (untrainable) or 5 (trainable) regardless of training quality.

Want an unbiased, rubric-based tool evaluation? Talk to Dorve.

Start a project

Want a team that knows the theory?

We research, design and build with people who can tell when an answer is wrong. Tell us what you are working on.

Start a project