If you follow the first approach, you need 0 containers “ready”, and only as many speculative containers created as you have concurrent initial chats that have not called for the tool but need their own ID.
The second “turn on python” tool idea is only creating containers when demanded and will be assigned to a session/conversation - and then yes, the big fault that if you aren’t self managing (which needs you to create an ID anyway) but using a server chat state product, that will break extremely quickly, even the amount of thinking and typing someone might do before their next ongoing input.
You have identified: not a single Responses hosted tool works like it should or like you would want; they all imagine you want exactly ChatGPT (and now gpt-5.1 even has an anti-developer unstoppable “you chat” tune-up system message).
The ongoing “auto” bug here: you get charged $0.03 for every new user input! 10 “hello” messages added $0.30 and still didn’t provide you a code interpreter ID to reuse.