Updated Aug 20, 2026

Text-to-Image

Generating original images from a written description, typically using a diffusion model.

Share

What it means

Text-to-image systems turn a written prompt into an image that has never existed. Most use diffusion, starting from noise and refining toward something matching the description.

Capability has moved fast in specific, uneven ways. Text rendering inside images — long a reliable tell — now largely works. Compositional control has improved, as has editing an existing image by instruction rather than regenerating from scratch.

The unresolved issues are legal and economic rather than technical. Models trained on scraped images are the subject of active litigation; some vendors now differentiate on licensed-only training data and offer indemnification. Style imitation sits in a genuinely unsettled area — style is not copyrightable, but generating work in a living artist's manner raises questions the law has not answered.

Why it matters

For most organizations this is the most immediately usable generative capability — drafts, concepts, and internal visuals at effectively zero marginal cost. Whether output is safe for commercial use depends on the vendor's training data and terms, and that varies enormously.

In practice

For commercial work, check the training-data provenance and whether the vendor indemnifies you. Expect to generate several and select; output is stochastic, and a seed value is what makes a result reproducible while you iterate.

Where this shows up

Tools and models in our catalog.

Related terms