In April 2022, OpenAI introduced DALL-E 2, a neural network model capable of transforming textual descriptions into intricate visual art. This model astounded many, even the skeptics, with its ability to generate detailed images from simple sentences. The excitement around DALL-E 2 was soon followed by the release of Google's Imagen, a model that seemed to surpass DALL-E 2 in generating photorealistic images from textual prompts.
Despite the excitement, there were voices of caution and skepticism. Gary Marcus, a cognitive psychologist, drew parallels between the overhyped claims of AI and historical incidents like the "Clever Hans" phenomenon, where a horse was believed to perform arithmetic but was actually responding to subtle cues from its trainer. Marcus argued that DALL-E 2 and similar models might be overclaimed in their capabilities. He pointed out that while these models could generate impressive images, they often failed with more complex or nuanced prompts..
The AI community had mixed reactions to these models. Some researchers and enthusiasts defended the models, suggesting reasons for their failures, ranging from training data issues to common sense interpretations. However, these explanations were mostly speculative. The Real Issue: Language Understanding Behnam Neyshabur from Google explained that the key to getting Imagen to generate a "horse riding an astronaut" was to rephrase the prompt more specifically. This pointed to a deeper issue: the model's failure wasn't due to an inability to generate the image but rather a lack of genuine understanding of the language.
https://huggingface.co/spaces/stabilityai/stable-diffusion-3-medium
https://openart.ai/create
"Money plays with a puppy 🐶"