Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Why is it not surprising? I don't see any fundamental reason for it. I think these models will be able to produce sensible graphs fairly soon.

You could equally say "it's not surprising that DALL-E can't draw words"... except that Imagen seems to be pretty good at it.

I think the real reason it's not surprising to you is that you've already seen enough DALL-E results to understand its limitations. It's not surprising that DALL-E can't draw graphs.



I do see a fundamental reason: The current crop of AI tools are horrible at logic. It's a complete inversion of how we think of computers.

If I want to convey happy emotions in the style of Rembrandt, SD or DALL-E will do brilliantly. If I want an apple BELOW a table, or worse, a geometric shape like a triangle, they'll crash-and-burn.

GPT-3 is also really empathetic, but struggles simple logic (and especially mathematics).

Graphs are like the horror case for these.

I can think of ways to make them better at this, but it's not a weekend of hacking.


We don't know if it's surprising, the author never tells us their hypothesis. They don't state any particular reason for the prompt they used, they don't explore the contents or qualities of the prompt compared to other AI-generated art, and they don't run multiple trials. Because of that, we can't conclude anything useful from this article. There's no frame of reference or scientific inquiry involved. If you find it entertaining, that's fine. As a scientific comparison, this verges on parody.


This is not a scientific experiment. The author compared results of a specific prompt.

Why are so many people overthinking this?


> Why are so many people overthinking this?

From reading the comments here, they're overthinking it because they seem to be taking this as a pre-planned "attack" on AI art generation, rather than just an interesting anecdote on the limitations of these tools.

As someone who has not played with said tools, it was an outcome I found interesting to know: DALL-E et al. can't do specific graphs or even specific logical things very well a lot of the time. That's good to know, and I didn't previously!


I played with these tools a lot, so expected bad results as soon as I read the prompt.

Still found the post really interesting as it explores a very realistic use case. A client needs something simple designed for a blog post. Should they use AI or a human designer?

I read somewhere in the comments here that these tools are very bad at counting. Which is an interesting limitation with far reaching implications.


> I read somewhere in the comments here that these tools are very bad at counting. Which is an interesting limitation with far reaching implications.

It may not make much difference, but it's not so much that they're bad at counting, as that they don't even try. The way the prompt is parsed and diffused doesn't allow for that sort of logic.

All a "two people" prompt or some such provides, is a hint to push the AI towards that section of latent space where training-set images titled "two people" exists.

That's not "counting", and it would be truly amazing if any sense of math emerged de-novo from this training process. Doesn't mean it can't be done — it means we aren't trying.


> it means we aren't trying

It'll be pretty exciting times once we do!


I think it's just a bad experiment. The author must have been truly ignorant about the capabilities and utilities of DALL-E if he thought this experiment would yield interesting results.


I am personally annoyed because I expected something more sensible from an article called “DALL·E 2 vs. $10 Fiverr Commissions”.

The author could also compare how well DALL·E draws text, but what would be the point of that? Is not being a scientific article a good defense for posting nonsense?


The most obvious reason would be that they're probably largely trained on art-type images, not charts and graphs.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: