Hello OpenAI team,
First of all, thank you for creating such an impressive image generation system. I have been using it extensively to create a full-color manga, and after generating dozens of sequential pages, I would like to share two challenges that consistently affect long-form visual storytelling.
My goal is not to criticize the model, but to provide feedback from the perspective of someone who uses it for large creative projects every day.
1. Visual Consistency Across Sequential Panels
Current image generation models produce excellent standalone illustrations.
However, creating a long-form manga presents a different challenge.
The biggest difficulty is maintaining visual consistency across multiple panels while preserving:
- character identity;
- facial features;
- clothing;
- lighting;
- environment;
- camera position;
- spatial relationships;
- continuity between consecutive panels.
Even when detailed prompts and multiple reference images are provided, achieving one successful manga page can require dozens of regeneration attempts.
As a result, a significant amount of production time is spent correcting inconsistencies rather than creating the story itself.
Suggestion
Improve long-term visual consistency by allowing the model to better preserve:
- characters,
- environments,
- lighting,
- camera position,
- spatial layout,
- and previously established visual details across sequential images.
For long-form storytelling, consistency is often more valuable than generating a completely new interpretation every time.
2. Limited Creative Control During Image Generation
The second challenge is the lack of author control during the generation process itself.
Imagine a traditional artist who says:
“I have an idea for this painting.”
Someone replies:
“Great. Leave the room and come back in five minutes.”
When the artist returns, they notice that the tree is in the wrong place, the character’s clothing has changed, and the composition is different.
The explanation is simply:
“I decided to draw it that way.”
This is similar to how image generation currently feels.
The generation process behaves like a black box.
As an author, I can only provide prompts and references.
I cannot observe the generation process.
I cannot intervene while the image is being created.
I cannot correct mistakes before they become part of the final result.
Even very detailed instructions cannot always prevent the model from creatively interpreting parts of the image that were already clearly specified.
For example, this is a simplified version of the prompt structure I regularly use:
REFERENCE PRIORITY
Images 1 = style, lighting, rendering, atmosphere.
PANEL 1 — CAMERA LOCKED TO IMAGE 2. DO NOT CREATE A NEW ANGLE.
PANEL 2 — CAMERA LOCKED TO IMAGE 3. DO NOT CREATE A NEW ANGLE.
Images 4 = character identity, face, body, clothes and canon details.
Image 5 = location, time of day, light direction, spatial logic and mood.
Image 6 = page layout reference only.
Written brief = story beat, action, panel count, acting and camera.
If anything conflicts, preserve canon and continuity first.
Even with explicit instructions like these, the model may still reinterpret elements that were intended to remain fixed.
This greatly increases the number of regeneration attempts required to produce a consistent manga page.
Suggestion
Instead of treating image generation as a completely opaque process, consider providing creators with more control during generation.
For example:
- preview intermediate generation stages;
- lock approved elements (characters, environment, lighting, composition);
- regenerate only selected parts of the image;
- preserve scene state between panels;
- allow authors to guide or correct the generation before the final render.
This would make AI feel less like a black box and more like a collaborative creative tool.
Why This Matters
The issue is not image quality.
The issue is production efficiency.
Creating a long-form manga is already a complex creative process.
Today, a significant amount of time is spent fighting inconsistencies rather than developing the story itself.
Giving authors greater control over the generation process—and improving consistency across sequential images—would dramatically reduce unnecessary iterations and make AI a much stronger tool for professional long-form storytelling.
Thank you for your work and for continuously improving these models.
I hope this feedback is helpful.