# How does the GPT4o in chatgpt web generate images

**URL:** <https://community.openai.com/t/how-does-the-gpt4o-in-chatgpt-web-generate-images/1055124>\
**Category:** ChatGPT\
**Tags:** gpt-4, chatgpt\
**Created:** [December 14, 2024, 1:26am UTC](https://community.openai.com/t/how-does-the-gpt4o-in-chatgpt-web-generate-images/1055124 "2024-12-14T01:26:51Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Goks](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/goks/32/511941_2.png) [@Goks](https://community.openai.com/u/Goks)\
**Post date:** [December 14, 2024, 1:26am UTC](https://community.openai.com/t/how-does-the-gpt4o-in-chatgpt-web-generate-images/1055124/1 "2024-12-14T01:26:51Z")

</div>

Hello Everyone,  
I know that gpt4o and gpt4o mini don’t have capability to generate images, only Dall -E can do that. Then how come in the chatgpt web if I ask gpt4o models to generate images it does? If anyone knows please let me know.

Thanks

---

<div class="post-metadata">

**Author:** ![edwinarbus](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/edwinarbus/32/455108_2.png) [@edwinarbus](https://community.openai.com/u/edwinarbus)\
**Post date:** [December 14, 2024, 1:29am UTC](https://community.openai.com/t/how-does-the-gpt4o-in-chatgpt-web-generate-images/1055124/2 "2024-12-14T01:29:57Z")

</div>

ChatGPT understands you want to create an image, and automatically routes the request to DALL-E.

---

<div class="post-metadata">

**Author:** ![EricGT](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/ericgt/32/20571_2.png) [@EricGT](https://community.openai.com/u/EricGT)\
**Post date:** [December 14, 2024, 1:30am UTC](https://community.openai.com/t/how-does-the-gpt4o-in-chatgpt-web-generate-images/1055124/3 "2024-12-14T01:30:30Z")

</div>

> [@Goks](#):
>
> only Dall -E can do that

> **[Multimodal learning](https://en.wikipedia.org/wiki/Multimodal_learning)**
>
> Multimodal learning is a type of deep learning that integrates and processes multiple types of data, referred to as modalities, such as text, audio, images, or video. This integration allows for a more holistic understanding of complex data, improving model performance in tasks like visual question answering, cross-modal retrieval, text-to-image generation, aesthetic ranking, and image captioning.
> Large multimodal models, such as Google Gemini and GPT-4o, have become increasingly popular since ...
