I’m thinking of a sprite that slams a hammer on a forge each time ‘copy prompt’ is clicked in my tool.
What have you folks determined to be the cleanest workflow to create said sprite? Are there any tips I should be aware of before going into this blindly?
It’s not imperative to my flow today, but it’s something of a thing to have a dwarf smash a hammer onto a forge each time someone clicks ‘copy prompt’ in my app…
I have a list of things I can work on before I get to it, but I really want to understand the best practice
I just wrapped up version 2 of my LLM-driven game agent. Took a while but a big step up from my previous version. Enjoy and let me know what you think!
Here is something I made in 10 minutes. Took a screenshot of Monkey Island and put it in Codex together with one liner stating I want a MI sandbox replica and to use gpt-image-2 to create smooth sprite animations and tilemaps.
What’s interesting is that I get MUCH better sprite results this way, then if I got explicitly to gpt-image-2 and prompt it. It feels like the additional context (e.g. the interpretation of the screenshots and the “pre-planning” done by Codex) gives much better results.
This is a bit disappointing tbh. Context: I’m experimenting with a “detective-noir” style game and trying to generate a femme fatale character. Every AI model I have used has consistently struggle with this walk cycle. But when I get this safety/moderation blocks (it happens ALL the time), it is insanely frustrating.
Since ChatGPT was first released, I have been tracking a small set of tasks that LLMs historically struggled with. Each time a new model becomes available, I revisit these tasks to see how capabilities have evolved. This appears to be a case worth adding to that set.
My primary interest is not in games, but in generating images and videos related to cellular biology. The difference between reading about these processes in a book or paper and actually seeing them represented visually is substantial—well-crafted graphics or animations can convey far more detail and intuition than text alone.
Yesterday, I came across the following resource: (released: Apr 21, 2026)
This guide highlights prompting patterns, best practices, and example prompts drawn from real production use cases for gpt-image-2 . It is our most capable image model, with stronger image quality, improved editing performance, and broader support for production workflows. The low quality setting is especially strong for latency-sensitive use cases, while medium and high remain good fits when maximum fidelity matters.
One of the generated images stood out due to its level of detail:
Super, thank you @EricGT , I will have a closer look at the guide and see if there is any way to steer it to get that other darn leg to come forward
Ignoring the unnecessary moderation errors I get, the walking cycle issue is not intrinsic to OpenAI models - I see the same issue with Gemini and Grok. If you have a look at ppl posting their Codex game demos on Twitter - every one of them that has a sidescroller walk cycle, shows the same “drag” effect. It’s almost like for this particular use case, there is a black hole in the training data. But maybe the guide you linked will help
For those interested, I do have a pipeline that works however, you may find it useful as well:
(1) pencil sketch character → (2) give to gpt-image-2 to generate a static sprite side-view → (3) give to Seedance 2.0 to generate a smooth action cycle → (4) ffmpeg extract frames → (5) sub-sample in-between frames to get either 12 or 24 max frames → (6) pass to gpt-image-2 to remove background → (7) ImageMagic re-size to e.g. 128x128 → (8) load in aseprite for manual touch up.
It seems like a lot, but for rapidly prototyping new characters and action sequences this is an absolute game changer.
This is very tricky. I’ve also been trying to use reference poses but at some point the model ends up producing repeated poses. Perhaps this subject would be an interesting thread on its own? I’ll come back to it if I make some progress.
I got it “sorted” but found it is more of a temperature issue of the image-2 being too creative and hallucinating.
Since this is a very specific issue, I will open a new thread so that we can keept track on better ways of improving the process.
I am currently heading towards a video call avatar and I thought some of the stuff from that research might fit well for gaming as well. I mean most of it was made for gaming anyways…
Here is a fresh one:
---
And you should also watch out for the work of Olga Sorkine-Hornung and her work on:
Laplacian Mesh Editing
and especially
ARAP: As-Rigid-As-Possible Surface Modeling
That stuff is patented by Disney afaik… so you may be careful how you use it - but on the other hand it is good to know none the less.
---
You can connect your codex to blender (3D modelling tool) and then work on a so called mesh*
Here are a few more words to feed chatgpt with:
landmark extraction, pose control, mimiks, phenoms and shape keys.
In that context adding that little background information about those techniques make alot of sense.