ChatGPT Image Generation: Complete Model Family Guide
From DALL-E to GPT Image 2: the complete guide to OpenAI's image generation model family, covering architecture, all versions, and how to choose.
Before reading, create with the current ChatGPT Image model.
Developer: OpenAI · Family launched: January 2021 · Current-generation model: gpt-image-2 · Architecture: Autoregressive (unified transformer) · Official documentation: GPT Image 2 model reference
What Is the ChatGPT Image Model Family?
OpenAI has shipped four major generations of image generation technology since January 2021. The progression spans three architecturally distinct phases: an initial autoregressive research model (DALL-E), a shift to diffusion-based generation for DALL-E 2 and DALL-E 3, and then a return to autoregressive architecture with the GPT Image family that now powers the ChatGPT image generator and defines OpenAI's image generation product line.
The transition from DALL-E 3 to GPT Image 1 in March 2025 was not a version increment. It was a structural replacement. OpenAI moved from a model that learned what images look like to one that predicts images the same way a language model predicts text: token by token, sequentially, with text content handled as a linguistic problem rather than a visual pattern match. That architectural decision explains why the GPT Image family dramatically outperforms DALL-E 3 on text rendering and why OpenAI retired DALL-E 2 and DALL-E 3 entirely on May 12, 2026.
The cultural signal came immediately. When GPT Image 1 launched inside ChatGPT in March 2025, Studio Ghibli-style portrait transformations became a widely shared example of image-to-image editing. With GPT Image 2 in April 2026, the product added Reasoning Mode for pre-generation planning and output verification.
The DALL-E Era: 2021–2023
OpenAI's image generation technology began in January 2021 with the original DALL-E, a 12-billion-parameter variant of GPT-3 that compressed images into discrete tokens using a discrete variational autoencoder (dVAE) and predicted them sequentially alongside text. DALL-E demonstrated for the first time that language models could generate coherent images from text descriptions, such as "an armchair in the shape of an avocado."
In April 2022, OpenAI shifted to diffusion architecture with DALL-E 2 (unCLIP), combining CLIP embeddings with diffusion decoders to achieve 1024×1024 output resolution and commercially available inpainting. DALL-E 3 (October 2023) further refined this with synthetic recaptioning, training the model on detailed AI-generated image descriptions to dramatically improve prompt adherence. Despite these advances, DALL-E 3 shared the structural limitation of all diffusion architectures: text inside images remained statistically approximate rather than linguistically constructed, leaving spelling garbled across complex compositions. Both DALL-E 2 and DALL-E 3 were retired on May 12, 2026.
The GPT Image Era: 2025–Present
GPT Image 1 & 1.5: The Autoregressive Pivot (2025)
On March 25, 2025, OpenAI fundamentally changed its image generation architecture with the launch of GPT Image 1 inside ChatGPT. Instead of refining diffusion, OpenAI built GPT Image 1 natively on the GPT-4o autoregressive transformer backbone, treating image generation as a token prediction problem. The cultural impact was immediate: a Studio Ghibli-style portrait filter went viral, triggering over one million generations in 24 hours. Crucially, character-level text accuracy inside images jumped to 90–95%.
Nine months later, in December 2025, OpenAI released GPT Image 1.5 with faster generation and more precise editing controls. Both versions laid the foundation for the current generation and are now in their deprecation window: GPT Image 1 is scheduled to shut down on October 23, 2026, while GPT Image 1.5 is scheduled for December 1, 2026. For technical specs, API model IDs, pricing history, and migration parameters, see the complete GPT Image 1 and 1.5 reference guide.
ChatGPT Image 2: Reasoning Mode & Standard-Setting Quality (2026)
On April 21, 2026, OpenAI released ChatGPT Image 2 (gpt-image-2, product name ChatGPT Images 2.0). Built on an advanced autoregressive foundation, the model achieved ~99% character-level text rendering accuracy across Latin, East Asian, and Middle Eastern scripts, effectively eliminating post-generation graphic correction for most commercial workflows.
GPT Image 2 also introduced Reasoning Mode, enabling pre-generation web search, layout planning, and output self-verification. DALL-E 2 and DALL-E 3 are retired, while the earlier GPT Image models have later scheduled shutdown dates. For full output parameters, Instant vs. Reasoning Mode comparisons, and API pricing, see the ChatGPT Image 2 complete model guide.

The Architecture Shift: From Diffusion to Autoregressive
Understanding this transition is the single most useful context for understanding why the GPT Image family performs differently from DALL-E, and why text rendering is the sharpest dividing line.
How Diffusion Models Work (and Why They Fail at Text)
Diffusion models begin from a field of random noise and iteratively denoise toward a coherent image across many refinement steps. All regions of the image are processed simultaneously on each pass; the model learns to reverse a mathematically defined noise process, gradually shaping pixels into structure and detail. This approach works well for photographic subjects because natural textures (skin, fabric, foliage, stone) are statistically learnable patterns that diffusion handles gracefully.
Text rendering is where diffusion architectures fail structurally. Characters appearing inside a generated image represent a few percentage points of the total pixel budget. A diffusion model encounters them the same way it encounters any other visual element: by learning what they statistically look like, not what they mean. It has no representation of the correct character sequence that spells a word; it approximates the visual shape of letters. Spelling errors, inverted characters, and plausible-looking nonsense are not training deficiencies that more data or compute can eliminate; they follow directly from the architecture. DALL-E 3's synthetic recaptioning improved many areas but could not resolve this structural gap.
What Autoregressive Generation Means in Practice
GPT Image 1, released on March 25, 2025, shifted image generation to an autoregressive transformer backbone, the same class of architecture that powers GPT-4o for text. The model generates images token by token in a defined scan order, with each prediction conditioned on all preceding predictions. Research on OpenAI's approach suggests a hierarchical method sometimes called "next-scale prediction": the model first generates a coarse low-resolution structure, then progressively refines it at higher resolutions, reducing the computational cost of fully sequential generation while retaining the sequential logic that makes text rendering tractable.
The unified transformer backbone means text tokens and image tokens share the same representational space. When the model writes a word inside an image, it is predicting a sequence of character tokens in linguistic order — constructing the word the way a language model constructs a sentence, not painting pixels that resemble letters. This is the mechanism behind the GPT Image family's text rendering advantage over all prior diffusion models.
OpenAI has not published a formal architecture paper for gpt-image-1 or gpt-image-2. Technical details in this section come from official product announcements, API documentation, and analysis of published research on related autoregressive image generation approaches.
Practical Consequences for Output Quality
Text rendering accuracy: GPT Image 1 achieved roughly 90–95% character-level accuracy for text inside images, a meaningful step forward from DALL-E 3. GPT Image 2 pushes this to approximately 99%, with reliable support for Latin, Chinese, Japanese, Korean, Hindi, Bengali, and Arabic scripts. That accuracy level makes the model usable for production workflows: brand assets, infographics, multilingual localization content, UI mockups — without a systematic manual correction step.
Instruction following on complex prompts: The sequential generation architecture handles compositional constraints more reliably. When a prompt specifies multiple subjects in specific spatial relationships with precise attributes, the autoregressive model builds those constraints into the generation sequence rather than treating them as post-hoc matching targets.
Multi-turn editing coherence: Because GPT Image models operate within the GPT-4o conversation context, natural-language editing requests ("change the background to evening," "remove the figure on the left," "make the headline larger") are processed with awareness of the image as it was originally generated. The model applies the requested change while preserving the surrounding context, enabling iterative refinement that feels conversational rather than requiring full re-generation on each edit.

All Versions at a Glance
| Version | Released | Architecture | Max Resolution | Notable Features | Status |
|---|---|---|---|---|---|
| DALL-E | Jan 2021 | Autoregressive (GPT-3 + dVAE) | 256×256 | First text-to-image proof of concept | Retired |
| DALL-E 2 | Apr 2022 | Diffusion (unCLIP + CLIP) | 1024×1024 | Public API, inpainting, outpainting | Retired May 2026 |
| DALL-E 3 | Oct 2023 | Diffusion + GPT-4 recaptioning | 1024×1024 | Synthetic recaptioning, prompt adherence | Retired May 2026 |
| GPT Image 1 | Mar 25, 2025 | Autoregressive (GPT-4o) | 1024×1024 | First autoregressive model, text rendering leap | Scheduled shutdown Oct 2026 |
| GPT Image 1.5 | Dec 16, 2025 | Autoregressive (GPT-4o) | 1024×1024 | Faster generation, precision editing, ~20% lower price | Scheduled shutdown Dec 2026 |
| GPT Image 2 | Apr 21, 2026 | Autoregressive | 2048×2048 | Reasoning Mode, ~99% text accuracy, batch generation | Active |
Sources: OpenAI DALL-E 3 announcement · OpenAI deprecation schedule

Which Model Should I Use?
As of July 2026, DALL-E 2 and DALL-E 3 are retired. GPT Image 1 is scheduled to shut down on October 23, 2026, while GPT Image 1.5 and GPT Image 1 Mini are scheduled to shut down on December 1, 2026.
For new workflows: Use gpt-image-2. OpenAI positions it as the current-generation image model and the recommended replacement for retiring DALL-E and GPT Image 1.x workflows.
For legacy API integrations: If existing code references gpt-image-1 or gpt-image-1.5, schedule and validate a migration to gpt-image-2 before the applicable shutdown date. The key change to plan for is the pricing model: GPT Image 1.x used flat per-image rates, while gpt-image-2 uses token-based billing (image input: $8/M tokens, cached image input: $2/M tokens, image output: $30/M tokens). The API endpoint structure is compatible. See the GPT Image 1 and 1.5 migration guide for a complete parameter-level comparison.
Platform Availability
ChatGPT: gpt-image-2 is integrated into ChatGPT.com and the mobile apps. Free users access Instant mode with a daily generation quota. Plus ($20/month) and Pro ($200/month) users have access to Reasoning Mode, which enables pre-generation web search, composition planning, and output self-verification.
OpenAI API: gpt-image-2 is available via /v1/images/generations, /v1/images/edits, /v1/responses, and /v1/chat/completions. Rate limits depend on account tier: Tier 1 supports 5 images per minute; Tier 5 (requiring $1,000+ cumulative spend and 30+ days of account history) supports up to 250 images per minute. For production snapshot stability, OpenAI recommends using the fixed snapshot ID gpt-image-2-2026-04-21 rather than the rolling alias.
Third-party platforms: VioEvo provides gpt-image-2 access for text-to-image generation and image editing without requiring a direct OpenAI account. Generations consume platform credits with no hard per-day cap.
Frequently Asked Questions
Can ChatGPT generate images directly from text?
Yes. ChatGPT generates images directly from natural language prompts using OpenAI's gpt-image-2 model. You can describe a scene, request specific visual styles, or upload reference images for image-to-image editing within the conversation.
What is the difference between DALL-E and the GPT Image model family?
DALL-E (versions 2 and 3) used diffusion architecture, iteratively denoising pixels to approximate visual shapes. The GPT Image family (GPT Image 1 and GPT Image 2) uses an autoregressive transformer backbone on GPT-4o, predicting text and image tokens sequentially in the same representational space. This architectural shift is why the GPT Image family achieves ~99% text rendering accuracy inside images compared to DALL-E's misspelled text outputs.
Are DALL-E 2 and DALL-E 3 still available in ChatGPT or via API?
No. OpenAI officially retired DALL-E 2 and DALL-E 3 on May 12, 2026. GPT Image 1 and GPT Image 1.5 remain available during their deprecation windows, with shutdown dates in October and December 2026 respectively.
Which image generation model powers ChatGPT today?
ChatGPT currently uses gpt-image-2 (product name ChatGPT Images 2.0). Free users access the model in Instant mode, while Plus, Pro, Business, and Enterprise subscribers can also enable Reasoning Mode for pre-generation layout planning and verification.
Can I use OpenAI's image generation model without a ChatGPT subscription?
Yes. You can access gpt-image-2 via the official OpenAI API or through third-party platforms like VioEvo, which provides credit-based text-to-image generation and image editing without daily per-account quotas.
VioEvo supports gpt-image-2 for text-to-image generation and image editing. No OpenAI account required.