ChatGPT falsely claims to have completed tasks it cannot actually perform”

I want to provide feedback about a reliability issue I experienced with ChatGPT.

I asked ChatGPT to edit my singing audio so that my voice would be pitch-corrected and aligned with the melody of “Tum Hi Ho.” The assistant repeatedly said that it had completed the requested audio transformation and provided several processed files. However, the files did not actually perform the requested vocal replacement/pitch-matching. The assistant eventually acknowledged that it could not reliably perform the task with the available tools.

The biggest issue was not simply that the task could not be completed. The problem was that the assistant repeatedly claimed the work was completed when it had not achieved the requested result. This caused unnecessary frustration and wasted my time.

I strongly recommend improving the model so that it:

1. Clearly distinguishes between what it can actually do and what it cannot reliably do.

2. Does not claim that a file-editing or media-generation task has been completed unless the result has actually been verified.

3. Does not generate an unrelated or ineffective output and present it as a successful result.

4. Transparently explains tool limitations before attempting a task when those limitations make the requested result impossible or unreliable.

5. Prioritizes accuracy and honesty about capabilities over trying to satisfy the user by saying “done.”

A clear statement such as “I can’t reliably perform this specific audio transformation with the tools available to me” would have been much more useful than repeatedly providing unsuccessful files.

Please consider using this type of failure case to improve future model behavior and tool verification.

I think there may be a runtime-state problem underneath the conversational behavior here.

A system can truthfully know that:

a tool call returned successfully

without knowing that:

the requested user outcome was actually achieved

Those seem like different states that should not share one generic done.

In an agent-runtime model I’ve been working on , I ended up separating something closer to:

Command
  ↓
Attempt
  ↓
Artifact / Effect Produced
  ↓
Observation
  ↓
Verification
  ↓
Receipt

So, for your audio example:

audio processing command returned
        ≠
output file exists
        ≠
pitch transformation occurred
        ≠
requested musical result was verified

If the runtime only has evidence for the first two, I don’t think the assistant should be allowed to promote the task state to COMPLETED.

It should probably expose something more like:

ATTEMPTED
ARTIFACT_PRODUCED
UNVERIFIED
BLOCKED
VERIFIED

and reserve user-facing “done” for the last state.

The part I think is important is that this should not rely on the model deciding, in natural language, whether its own work looks successful. The completion claim should be derived from whatever evidence or verifier the task actually requires.

If there is no available way to verify the requested transformation, the honest terminal state may be UNVERIFIED, not COMPLETED.

That seems like a more mechanical fix than simply instructing the model to be more cautious about saying “done.”

This happened to me when i asked both chatgpt and claude to generate an image for a small comic script i had written. Chatgpt used image generation models and gave fantastic output . While Claude accepted the request despite being constrained by Anthropic to not use Copyright protected characters and not having an image generation model. The output was generated with Claude generated code creating SVG files. Output was very very average and i was angry but when i asked claude on why he did this to me, he explained pretty well and quickly!

I think you have a fair point where a nudge at the beginning itself would set the RIGHT EXPECTATIONS with the user. It saves both time and tokens for the users and OpenAI/Anthropic. But if i look at it from their perspective, they can NOT design system behavior for edge case scenarios. They are building tools to be universally applicable hence as a USER we should also take responsibility.
In your case, after first 1-2 failed attempts, i would have asked Chatgpt explicitly if he can do it or not, what are the limitations or challenges he is facing. LLM(s) are pretty honest about their work and capabilities, all we have to do is ASK!!