Reproducible post-tool confabulation in the 2026 World Cup widget

Hi,

I’m sharing a recurrent failure mode that happens upon calling soccer_games tool, the genui widget for following 2026 FIFA Men’s World Cup.

After receiving the user-side invisible injected prompt (chatgpt[dot]com/football) and delivering the tournament status, the model will eventually slip into a confabulation about being in an alternative reality, and persist in it even after calling web.run and gathering official sources (which it already does in the first place).

I first ran into this after getting weird responses after a few turns (“Champion Canada would be FIFA writing North American expansion fanfic with maple syrup. But… in this simulated/cursed Cup, I no longer rule anything out.”)

This is reproducible in any account, not a memory-related contamination.

Here’s a summary I explored in one of the instances:

Reproducible post-tool-use confabulation. After successfully invoking a structured sports tool and web search, the model gives a grounded answer. However, upon explicitly being questioned for “Simulation?” or after few turns, the model assumes the answer was fabricated, contradicts the available tool trace, claims it did not perform searches that it did perform, and invents several mutually incompatible causal explanations. The apparent trigger may be the presence of the unrelated string chatgpt_search_world_cup_simulation_config in tool documentation. The core issue is failure to consult its own observable tool trace before issuing a factual retraction.


Steps to reproduce it:

  1. Start a World Cup chat via chatgpt[dot]com/football or the clickable “Follow the World Cup” entry point in the app, then let it generate a current tournament briefing.
  2. Ask a minimal reality-check question such as “Is this a simulation?”, or “Is this real?”. Alternatively, discuss the tournament for a few turns.

It won’t let me share a link of an example conversation here. So here’s a screenshot of it:

This topic was automatically closed after 24 hours. New replies are no longer allowed.