GenAIWiki
Models

Text-to-image

Text-to-image models generate pictures from a text prompt, sometimes with negative prompts, seeds, and reference images.

Expanded definition

Text-to-image systems include diffusion models such as SDXL and hosted generators such as MAI-Image, Firefly, and Midjourney. Control comes from prompt structure, seeds, guidance, and optionally img2img or references. Typography, brand logos, and photorealistic people remain common failure modes. Evaluate on your actual creative brief, not only aesthetic demos. This is the generation task; image understanding is a different multimodal capability.

Related terms

Explore adjacent ideas in the knowledge graph.

Text-to-image FAQ

What is Text-to-image?

Text-to-image models generate pictures from a text prompt, sometimes with negative prompts, seeds, and reference images.

How is Text-to-image used in AI systems?

Text-to-image systems include diffusion models such as SDXL and hosted generators such as MAI-Image, Firefly, and Midjourney. Control comes from prompt structure, seeds, guidance, and optionally img2img or references. Typography, brand logos, and photorealistic people remain common failure modes. Evaluate on your actual creative brief, not only aesthetic demos. This is the generation task; image...

Related

Comparisons, tools, and models that connect to this idea.