Part 2 of 4. What counts as 'real' AI image generation, which five apps made the cut, and the exact 11-variable rubric used to score them.
To separate real AI image generation from filters, one criterion did the work: can the app create completely new visuals from a prompt? Filter apps only manipulate existing images. That narrowed hundreds of ‘AI’ apps to eight, then five to test: Adobe Firefly, Bing AI Image, Leonardo, Kaiber and Midjourney (Dall-E 2 was set aside after a testing error). The method was qualitative think-aloud with five demographically varied testers, scoring 11 variables from 1 to 5. The study is independent: no affiliate links, no incentives, and the team paid for paid apps.
Key takeaways
- The defining test of real generative AI: can it create new visuals from a prompt, not just filter an existing image?
- Field narrowed from hundreds of self-proclaimed 'AI' apps to 8, then 5 tested.
- Final five: Adobe Firefly, Bing AI Image, Leonardo, Kaiber, Midjourney (Dall-E 2 set aside).
- Method: qualitative think-aloud, in-person, five varied testers, 15-minute discovery then a short tutorial.
- Eleven variables scored 1-5, with explicit rules for missing/free/trainable features.
- Fully independent: no affiliate links, no incentives, paid apps paid for.
Real generative AI vs filters
Many app publishers claim ‘AI’ but only apply filters no deeper than Instagram effects. The refined criterion, can it generate something entirely new from a prompt, separated the wheat from the chaff. Examples like Loopsie (background swap plus a pseudo-3D effect) and Disflow are filters, not generators. That left eight apps, trimmed to five by dropping near-duplicates in favor of the most representative options.
The five apps tested
- Adobe Firefly
- Bing AI Image
- Leonardo
- Kaiber
- Midjourney
Dall-E 2 was intended as a sixth but excluded from scoring after two testers missed it, to keep the data consistent (a re-test was planned). The team stresses independence: no affiliation, no affiliate links, no incentives, and they paid for paid apps.
The tester panel
| Tester | Profile | AI/design experience |
|---|---|---|
| Tester 1 | Female, 25, professional designer, good English. | New to gen-AI |
| Tester 2 | Female, 29, professional designer, good English. | Some AI experience |
| Tester 3 | Male, 52, professional designer, excellent English. | High AI expertise |
| Tester 4 | 34, no design knowledge, good English. | No AI/design |
| Tester 5 | Female, 40, no design knowledge, excellent English. | No gen-AI |
Each tester got 15 minutes of free discovery (compared against pre-existing data), then a short tutorial and a command list for apps that needed one. Testing was in-person think-aloud, in a non-English-speaking country, with participants of good-to-excellent English.
The 11 scoring variables
| # | Variable | What it measures |
|---|---|---|
| 1 | Image quality | Quality of the final output. |
| 2 | Prompt-to-image adherence | How faithful the output is to the text prompt. |
| 3 | Image-to-image adherence | How faithful the output is to an uploaded image. |
| 4 | Trainable | Whether the model can be trained, and how easily (perceived). |
| 5 | Weighting | Whether prompt weighting is available. |
| 6 | Negative weighting | Whether elements can be down-weighted or excluded. |
| 7 | Special features | Zoom, pan, color modes, upscaling, models, etc. |
| 8 | User expectation | How close the result was to what the user expected. |
| 9 | Usability | Overall usability, judged subjectively. |
| 10 | Absolute price | Perceived absolute value. |
| 11 | Relative price | Willingness to pay given the results. |
Scoring rules: a missing feature scored 0, with two exceptions. If the app was free, price (Q10) scored 5; and ‘Trainable’ (Q4) scored 0 if untrainable and 5 if trainable, regardless of training quality, since testers couldn’t fully evaluate training. Business models may have changed since testing.
An app that creates completely new visuals from a prompt. Apps that only manipulate an existing image are filters, even if they use some AI internally.
Adobe Firefly, Bing AI Image, Leonardo, Kaiber and Midjourney. Dall-E 2 was intended as a sixth but set aside after a testing error.
Qualitative, in-person think-aloud with five demographically varied testers: 15 minutes of discovery, then a short tutorial, scoring 11 variables from 1 to 5.
Yes. No affiliations, no affiliate links, no incentives, and the team paid for any apps that required payment.
Missing features scored 0, except price scored 5 if the app was free, and ‘trainable’ scored 0 (untrainable) or 5 (trainable) regardless of training quality.
Want an unbiased, rubric-based tool evaluation? Talk to Dorve.
Start a project