Today I had an interesting experience improving an image-editing workflow built around GPT-Image 2.5, with help from ChatGPT Work.
I already have an image-generation/editing service, and I have been experimenting with interactive editing tools—especially point-based comments and brush-based selections.
One difficult case was surprisingly simple:
“Put a gold ring on this finger.”
The user clicks a finger or zooms in and paints over the intended area. However, when several fingers are close together, the ring sometimes appears on a neighboring finger.
This led me to separate three different problems:
-
Understanding what the user wants.
-
Identifying the intended target.
-
Controlling which pixels can change in the final image.
A precise click helps, but the application still needs a reliable way to carry that selection through the entire editing process.
What the original workflow did
For comment-based editing, I sent the full image with a text instruction containing the click’s percentage coordinates.
For brush-based editing, I already used a mask.
For markup editing, I supplied the clean original alongside an annotated reference image.
These approaches worked for many requests, but small, densely packed features remained challenging. Adding a mask alone was therefore not the whole solution.
Also, zooming the editor to 150% helped the user select accurately, but it did not automatically enlarge the relevant detail in the image sent to the model.
What I changed
I added a local editing pipeline:
-
Convert the user’s click, brush strokes, or shapes into an explicit selection in original-image coordinates.
-
Crop the selected area together with enough surrounding context.
-
Resize that crop proportionally into a 1024 × 1024 working image, using padding where needed.
-
Prepare a clean crop, a selection-guide image, and an alpha mask using the same coordinate transformation.
-
Send those inputs to the image-editing model.
-
Resize the edited crop back to its original dimensions.
-
Composite only the selected area into the original image.
The image model receives a larger view of the relevant detail. The final compositing step preserves the original pixels outside the selection, even if the generated crop contains changes elsewhere.
The crop provides context; the selection defines what can actually be applied. Those are separate boundaries.
Why this helped
The improvement came from several changes working together.
The target occupies more of the model’s input image.
A tiny feature in a large photograph becomes easier to distinguish within a local crop. This does not recover missing detail from a blurry source, but it gives the existing detail more space in the working image.
The intended location is communicated in multiple ways.
The prompt, visual guide, and alpha mask describe the same selection. They are generated from the same coordinates, so they stay aligned.
The application controls what reaches the final image.
Instead of assuming the model will preserve every unselected detail, the application applies only the selected portion of the generated result.
This noticeably improved my own tests, although I have not run a large benchmark or measured a general success rate.
Optional AI target clarification
I also added an optional target-clarification step before image editing.
It receives the whole-image context, the clean crop, and the selection guide, then returns a target description and a clarified editing instruction.
It does not move the selection or generate new mask geometry. The user’s selection remains authoritative. If the target cannot be identified confidently, that edit stops before the image-generation request.
The crop-and-composite workflow already helps without this extra step. Target clarification is useful when nearby or overlapping objects are difficult to distinguish, but it adds another model call and additional latency.
Two calls do not necessarily mean exactly twice the cost, since the interpretation and image-editing requests have different usage characteristics.
An important tradeoff: the selection must fit the finished edit
Strict compositing also revealed another issue.
If the user clicks one eye and asks for glasses, a small circular selection may be too narrow. The generated crop can show a complete pair of glasses, while the final composite retains only the portion inside that circle.
This can make the glasses appear incomplete—or almost disappear.
I added an option to include the surrounding area. In that mode:
-
The original click identifies the intended subject.
-
An adjustable rectangle defines the permitted edit area.
-
The same rectangle is used for the guide, mask, and final composite.
-
The instruction asks for the complete addition within that boundary.
For glasses, the rectangle needs enough room for both lenses and the bridge. The same principle applies to labels, decorations, accessories, and other additions that extend beyond the clicked point.
What is implemented—and what remains a proposal
The current implementation validates coordinate transformations and image dimensions, and preserves unselected pixels during compositing.
It does not automatically verify that the generated ring is on the correct finger or that the requested object was successfully created.
An incorrect edit can still occur inside the allowed region. Target interpretation and output validation are different tasks.
That distinction matters when discussing “precision”: the application can enforce where a result is applied without guaranteeing that everything generated inside that area is correct.
API and product ideas
This experience suggested several capabilities that could help developers building interactive image editors:
-
Point-to-region generation: Turn a click or a few points into a proposed editable region that the user can inspect and adjust.
-
Editable-region locking: Offer stronger guarantees about changes outside a supplied region.
-
Preserve or keep-out regions: Explicitly identify pixels or objects that should remain unchanged across repeated edits.
-
Edit-location validation: Check whether a requested addition appears in the intended region and on the intended object.
-
First-class editor inputs: Support comment pins, brush strokes, bounding boxes, and segmentation information as explicit editing inputs.
These are proposals; they are not all features of my current implementation.
What I learned from the development process
ChatGPT Work was useful throughout the iteration: examining a failure, discussing possible causes, modifying the pipeline, and testing the resulting behavior.
The most useful lesson was that precision depends on the surrounding application as well as the image model.
A graphical editor needs to carry the user’s spatial intent consistently through selection, input preparation, generation, and final compositing.
I started with rings appearing on the wrong finger and ended up with a more general approach to local image editing.
I have documented the implementation separately in Markdown and can share more technical details if useful. I would also be interested to hear how other developers handle small, closely spaced targets—and how they balance strict region preservation with enough freedom to complete larger edits.