We use gpt-realtime in Japanese, we tested gpt-realtime-1.5, and found in our automated tests it had a 10% reduction in accuracy (which is likely fixable via prompt tuning), however we found we were unable to adopt gpt-realtime-1.5 due to a very noticable reduction in the quality of pronounciation - the model often sounds robotic or as if it is a westerner speaking Japanese rather than a native speaker. For this reason we had to revert back to gpt-realtime
Similar to the above post about Hebrew - I wonder whether Japanese was included in the evaluation set, and if the team has any way of checking programatically for regressions in pronounciation quality in Japanese?
Is there any plans for improved model for Japanese or is the long term direction to focus more on English?