Taxonomy of antagonistic Midjourney prompts
I wanted to comment on 96 layers' article by James Mccammon which explores the limitations of text-to-image AI systems in accurately interpreting and visualizing complex, unconventional, or linguistically ambiguous prompts.
Inversion Prompts
- Create inverted images, depicting the opposite of what the human meant
- Sticking to the canonical interpretation (riding = the horse is ridden on)
- Prompt invariance
Discordant Prompts
- Create images with unexpected artifacts
- Don’t represent intent
Homonymously Discordant Prompts (HDPs)
Confuse homonyms and incorporates multiple meanings of a word into a single image
I'd like to argue that only prompts with homographs are susceptible to this confusion when using text-to-image models
These results were generated using Midjourney V5:
A crane standing in the river
The model seemed to aim at all possible scenarios, even though "river" is semantically related to the bird
-In the palm of your hand
Even though "palm" is semantically related to "hand", and the English phrase is common, a palm tree is included, just in case
-A giant metal spring
The giant metal (material) spring was placed outside, under clear blue skies and blooming trees. I wonder why...
- I believe that the writer should have been more specific in the HDPs' definition.
Homonyms can be either homophones, with the same sound using different spelling or homographs, where the spelling is exactly the same.
These tools operate on textual input, interpreting the written language and relying on the spelling of words to generate images. Since homographs have the same spelling but different meanings, the AI must contextually infer which meaning is intended, a task that can be challenging without sufficient context. In contrast, homophones' distinct spellings provide clearer differentiation, making it easier for the AI to distinguish between them based on the text alone, without needing to interpret pronunciation or auditory cues.
A toy being win(e)d up
When using a homophone, MJ understood what the word refers to and created an image based on the right meaning of the specific spelling
This might become problematic when audio-to-image models are used, where the spelling is not absolete