I just tried the prompts with Imagen and Parti. They are similar to Dalle-2, with a bit more "variety" but none reproducing the author's specific prompt the way the author wants. For the prompt "a graph with 3 lines" both produce graphs with 3 lines at least 1/6 of the time.
Curious. I've played with quite a few of these models and one of the very consistent "tells" is that they're extremely bad at counting things. A friend of mine likes tarot and I tried a few prompts... great results for the major arcana, but good luck with "ten of cups"... without capability to edit & re-prompt, the only viable strategy appears to be "ask it to draw a bunch of cups repeatedly until you've collected all the numbers."
Getting 3 of something 1/6 of the time doesn't really sound like it groks the request.
Sure, but it doesn't have to count to be useful. When I run this locally on desktop GPU, I get 16 results in a few seconds. I can visually select the 2-3 that match what I want and pick one of them. I can try again ten times for 20-30 options, and it still takes less time than Fiverr.
I would not use these models for graphs yet, but for cool look tarot inspired "clipart" or background images, I think they are already usable.
The inability to count has some interesting knock-on effects. My favorite being eyes and hands. Especially the hands. The closer you look, the creepier it gets. That thumb has fingers of its own?! Thus far, the greatest utility I've seen has been on par with B-movies.
I am somewhat surprised at how bad these tools are at generating hands and feet. Is it just a matter of not having enough images to ingest?
The faces often look very good and they also have symmetric complexity and individual elements that come in a specific quantity (2 eyes, etc). Lower quality models do generate fly-like multi eye faces, but newer ones are so much more precise!
Heh, remember when you signed up and still bought into Google's mission is to organize the world's information and make it universally accessible and useful.
It's kinda sad that Google desperately wants the cool points for having its own DL models but you can only see it in the form of a store window display.
> Heh, remember when you signed up and still bought into Google's mission is to organize the world's information and make it universally accessible and useful.
They probably also remember signing an NDA, and maybe taking some training about how not all of the world's information should be made universally accessible and useful to everyone. For instance, the contents of a user's inbox.