When a model hits capacity in the middle of an active task, the interruption is very disruptive. The system should ideally preserve continuity automatically—either by letting the current task finish, transparently switching to an available compatible model, or handing off the full working context without requiring the user to restart or reconstruct the task.
Capacity errors are understandable, but they should be handled differently once a task is already in progress.