Part 3 of 4. Five AI image generators, five testers, 11 variables each. Here's how they actually scored, and where each one shines or stumbles.
Per-app user-research results across the five tested tools. By total score, Leonardo.AI (3.89) and Midjourney (3.75) led, ahead of Kaiber (2.64), Adobe Firefly (2.09) and Bing (1.84). Midjourney had the best raw image quality and prompt control; Leonardo offered the most features for free. These are July 2023 scores: all of these tools have changed enormously since, so treat this as a methodology demonstration and historical snapshot, not a current ranking.
Scores at a glance
| Tool | Total (of 5) | Standout | Main drawback |
|---|---|---|---|
| Leonardo.AI | 3.89 | Huge feature set, many models (incl. user-trained), free at the time | Overwhelming UI, weak cognitive/visual accessibility |
| Midjourney | 3.75 | Best image quality and prompt adherence, accurate weighting | Discord-only, learning curve, low usability score, hidden renewal |
| Kaiber | 2.64 | Generates video, music reactivity, fast render | Only morphs images (not true video), questionable business practices, support denied a real bug |
| Adobe Firefly | 2.09 | Clean Photoshop-like UI, strong special features | Below-average text-to-image, no image-to-image, 'Not for Commercial Use' watermark |
| Bing Image Creator | 1.84 | Completely free, custom Dall-E, decent prompt adherence | Confusing flow, no options, lowest overall score |
Per-app notes
Adobe Firefly (2.09): polished and Photoshop-familiar, but text-to-image results were below average (odd hands, feet and faces), photographic quality average, and there was no image-to-image, not even ‘in exploration’. Downloads carried a prominent ‘Not for Commercial Use’ logo (framed as responsible AI).
Bing Image Creator (1.84): free, credit-based, running a custom Dall-E model. Flow is confusing (the ‘Create’ button sits far from the search box), it offers no options, and outputs four 1024×1024 images. Fun and free, but the lowest scorer. Noted as strikingly similar to You.com’s earlier AI-search-plus-generation model.
Leonardo.AI (3.89): the top scorer. Born for game assets, it grew into a feature-rich tool with many models (Stable Diffusion, proprietary, and user-trained), an ‘Alchemy’ mode for realism, upscaling, background removal, 3D Textures and AI Canvas. Free at the time (150 tokens/day; plans from $10/mo for 8500 tokens), behind a Discord-expedited waitlist. Excellent UI for its complexity, though cognitive and visual accessibility could improve.
Kaiber (2.64): unique in generating video (from images, text, audio, or clips) with camera control and music reactivity, ~$12-15/mo, fast render (a 30s video in under 10 minutes), upscaling to 1080p/4K. But it only morphs 3D-generated images rather than creating true video, watermarks outputs, uses a hidden auto-renewal, locks unused tokens to a 30-day cycle, and its Discord support denied a real, widely reported outage, hurting the customer experience.
Midjourney (3.75): self-funded, team-owned, Discord-only, with an undisclosed model. Best-in-class image quality and, thanks to its text-prompt approach, the best prompt quality and weighting of the group (from ‘/imagine’ onward), with well-documented parameters (resolution, quality, adherence, zoom, pan and more). Downsides: it inherits Discord’s problems, needs familiarity with two apps, is relatively pricey ($8-10/mo), scored low on raw usability, and also uses a hidden renewal.
Leonardo.AI (3.89 of 5), narrowly ahead of Midjourney (3.75), in the July 2023 test. Both far outscored Adobe Firefly and Bing.
Midjourney, with a perfect 5 on image quality and the strongest prompt adherence and weighting, though its usability score was low due to the Discord-only workflow.
Not really. At the time it morphed 3D-generated images rather than creating true video (it couldn’t render, say, people actually walking), and it watermarked outputs.
Firefly had below-average text-to-image results and no image-to-image; Bing offered no options and a confusing flow, giving it the lowest total score.
No. This is a July 2023 snapshot; every tool has changed substantially since. Use it as a methodology example, not a current ranking.
Need an objective, scored tool evaluation for your team? Talk to Dorve.
Start a project