Drawing graphs is probably one of the worst comparisons one can do in terms of evaluating these models. They seem to be trained to generate either photorealistic or stylized images.
It's not about it being theoretical, it's moreso that the language model is still far more simplistic than our own, and struggles with anything but the most basic relations between nouns. The "horse riding an astronaut" post is a good example of this.[0]
Did you even read the linked article? They even cite the picture you link, it still proves their point though these models have no understanding of the language (like Google claimed).
I think the path forward for something like this is models that learn to execute python code and incorporate the results into their outputs. There are already projects that can generate correct matplotlib calls for prompts like yours, but I don't think we are to the point where those python outputs can be automatically combined with a diffusion model for style or whatever.
One thing they're not very good at is deducing spatial relationships. Concepts like "above", "inside" and "behind", or "after". I'd say the prompts you gave make sense to a human who is thinking of a visual progression from left to right.
I bet you could write a few copilot prompts to generate code which would draw a graph like this, though.
I was surprised while reading your article how good the AIs did. I think it's fascinating that your intuition was that these tools would be able to do a good job of drawing a graph based on a description of it...
After reading your comment I asked Stable Diffusion to create photorealistic images of a graph with three lines (similar to the smallest prompt in the article).
Here's the results of three attempts with slightly different prompts: