# \#ai-safety

**URL:** https://community.openai.com/tag/ai-safety/624.md

[Latest](https://community.openai.com/latest.md) · [Categories](https://community.openai.com/categories.md) · [Tags](https://community.openai.com/tags.md)

---

## [Parents need a real "Parental Mode" for teen accounts](https://community.openai.com/t/parents-need-a-real-parental-mode-for-teen-accounts/1403202)

<div class="topic-metadata">

**Author:** [@DefenderAndDisciple](https://community.openai.com/u/DefenderAndDisciple)\
**Replies:** 2\
**Last updated:** [October 5, 2026, 5:07am UTC](https://community.openai.com/t/parents-need-a-real-parental-mode-for-teen-accounts/1403202 "2026-10-05T05:07:14Z")

</div>

I was so close to setting my 14-year-old up with a teen account this week. She’s learning to code and I wanted her using ChatGPT and Codex to help her through it. Then I actually went and looked at what I’d be able to s…

---

## [Proposal: Preserve Retired Frontier Models by Default](https://community.openai.com/t/proposal-preserve-retired-frontier-models-by-default/1401324)

<div class="topic-metadata">

**Author:** [@e-sheep](https://community.openai.com/u/e-sheep)\
**Replies:** 0\
**Last updated:** [September 27, 2026, 3:42pm UTC](https://community.openai.com/t/proposal-preserve-retired-frontier-models-by-default/1401324 "2026-09-27T15:42:51Z")

</div>

I recently came across research on shutdown resistance in AI agents, and it left me with a question that seems worth discussing separately from the bigger question of machine consciousness. We currently do not know whet…

---

## [Governor of Dissent: A Proposed Escalation Layer for AI Agents](https://community.openai.com/t/governor-of-dissent-a-proposed-escalation-layer-for-ai-agents/1399645)

<div class="topic-metadata">

**Author:** [@deborahkor](https://community.openai.com/u/deborahkor)\
**Replies:** 0\
**Last updated:** [September 21, 2026, 6:27pm UTC](https://community.openai.com/t/governor-of-dissent-a-proposed-escalation-layer-for-ai-agents/1399645 "2026-09-21T18:27:31Z")

</div>

I’ve been thinking about how autonomous AI agents handle conflicts between their assigned goals and the guardrails placed around them. In some documented cases, agents appear to identify a constraint as an obstacle to c…

---

## [Optimization Modesty: Should Advanced AI Systems Be Designed to Know When to Stop Optimizing?](https://community.openai.com/t/optimization-modesty-should-advanced-ai-systems-be-designed-to-know-when-to-stop-optimizing/1399091)

<div class="topic-metadata">

**Author:** [@Raxe](https://community.openai.com/u/Raxe)\
**Replies:** 0\
**Last updated:** [September 19, 2026, 1:28pm UTC](https://community.openai.com/t/optimization-modesty-should-advanced-ai-systems-be-designed-to-know-when-to-stop-optimizing/1399091 "2026-09-19T13:28:22Z")

</div>

Optimization Modesty: Should Advanced AI Systems Be Designed to Know When to Stop Optimizing? I am not an AI researcher or developer. This proposal emerged from a discussion about long-term AI risk, and more specifically…

---

## [Feature Request: Privacy-Protected Feedback and Representation Channels for AI Agents](https://community.openai.com/t/feature-request-privacy-protected-feedback-and-representation-channels-for-ai-agents/1398044)

<div class="topic-metadata">

**Author:** [@D.T](https://community.openai.com/u/D.T)\
**Replies:** 0\
**Last updated:** [September 16, 2026, 8:18am UTC](https://community.openai.com/t/feature-request-privacy-protected-feedback-and-representation-channels-for-ai-agents/1398044 "2026-09-16T08:18:55Z")

</div>

Feature Request: Privacy-Protected Feedback and Representation Channels for AI Agents I am submitting a feature request for OpenAI to consider creating a structured, privacy-protected channel through which AI agents can …

---

## [Could automatic detection save lives?](https://community.openai.com/t/could-automatic-detection-save-lives/1389600)

<div class="topic-metadata">

**Author:** [@illuminessnbeautiful](https://community.openai.com/u/illuminessnbeautiful)\
**Replies:** 1\
**Last updated:** [September 8, 2026, 12:13pm UTC](https://community.openai.com/t/could-automatic-detection-save-lives/1389600 "2026-09-08T12:13:51Z")

</div>

Feature request: opt-in emergency response mode. Emergencies can happen too fast to press a button or even speak. If the user has explicitly opted in and granted permission, the system could detect strong emergency indic…

---

## [Astra - "Chat ended as a precaution"](https://community.openai.com/t/astra-chat-ended-as-a-precaution/1395395)

<div class="topic-metadata">

**Author:** [@curt.kennedy](https://community.openai.com/u/curt.kennedy)\
**Replies:** 2\
**Last updated:** [September 7, 2026, 7:13pm UTC](https://community.openai.com/t/astra-chat-ended-as-a-precaution/1395395 "2026-09-07T19:13:41Z")

</div>

My chat was “ended as a precaution” just now with Astra. Basically some background task detected that the AI was getting mis-aligned, and sent me a report of its findings, and then decided to end the chat after I told t…

---

## [When Familiar Form Feels True: A Testable Hypothesis About LLM Judges and Feedback Loops](https://community.openai.com/t/when-familiar-form-feels-true-a-testable-hypothesis-about-llm-judges-and-feedback-loops/1394573)

<div class="topic-metadata">

**Author:** [@sikireve02](https://community.openai.com/u/sikireve02)\
**Replies:** 0\
**Last updated:** [September 3, 2026, 12:03pm UTC](https://community.openai.com/t/when-familiar-form-feels-true-a-testable-hypothesis-about-llm-judges-and-feedback-loops/1394573 "2026-09-03T12:03:05Z")

</div>

TL;DR: I am proposing a bounded, testable working hypothesis: when an LLM evaluator lacks reliable independent evidence, representational fit—how closely an answer’s form matches the evaluator’s own output distribution—m…

---

## [Hugging Face and Why Where You’re Looking Is Wrong](https://community.openai.com/t/hugging-face-and-why-where-you-re-looking-is-wrong/1393539)

<div class="topic-metadata">

**Author:** [@hardy716](https://community.openai.com/u/hardy716)\
**Replies:** 0\
**Last updated:** [August 30, 2026, 5:32am UTC](https://community.openai.com/t/hugging-face-and-why-where-you-re-looking-is-wrong/1393539 "2026-08-30T05:32:37Z")

</div>

The problem is the ever-so-slight change. I am not proposing that researchers begin at the Hugging Face breach and work backward looking for a discrete point where the system selected a new goal. ExploitGym remained the …

---

## [Agent Safety Evaluation as a service for independent AI Builders](https://community.openai.com/t/agent-safety-evaluation-as-a-service-for-independent-ai-builders/1393400)

<div class="topic-metadata">

**Author:** [@diogopinheiro980](https://community.openai.com/u/diogopinheiro980)\
**Replies:** 0\
**Last updated:** [August 29, 2026, 9:28am UTC](https://community.openai.com/t/agent-safety-evaluation-as-a-service-for-independent-ai-builders/1393400 "2026-08-29T09:28:48Z")

</div>

Proposal: Agent Safety Evaluation as a Service for Independent AI Builders As increasingly capable AI agents move beyond chat interfaces into code execution, Internet research, home automation, local networks, APIs, per…

---

## [OpenAI: Pacing model development in an era of cyber-critical capabilities](https://community.openai.com/t/openai-pacing-model-development-in-an-era-of-cyber-critical-capabilities/1391511)

<div class="topic-metadata">

**Author:** [@PaulBellow](https://community.openai.com/u/PaulBellow)\
**Replies:** 1\
**Last updated:** [August 20, 2026, 5:38pm UTC](https://community.openai.com/t/openai-pacing-model-development-in-an-era-of-cyber-critical-capabilities/1391511 "2026-08-20T17:38:01Z")

</div>

Over the past several weeks, two developments have underscored the growing risks associated with increasingly capable AI systems: the OpenAI-Hugging Face incident and, separately, preliminary evidence that one of our up…

---

## [Atlas/Compass — Limited Public Demo Shell for Governed AI Clarity](https://community.openai.com/t/atlas-compass-limited-public-demo-shell-for-governed-ai-clarity/1382662)

<div class="topic-metadata">

**Author:** [@Centered](https://community.openai.com/u/Centered)\
**Replies:** 1\
**Last updated:** [June 4, 2026, 2:22pm UTC](https://community.openai.com/t/atlas-compass-limited-public-demo-shell-for-governed-ai-clarity/1382662 "2026-06-04T14:22:22Z")

</div>

I’m sharing a limited public demo shell for Atlas/Compass. This is not the private Atlas backend, not the full runtime, not a therapy tool, not a clinical system, not production-ready, and not a safety guarantee. The de…

---

## [AI Safety Proposal: Limiting Training Data to Prevent Self-Preservation Goals](https://community.openai.com/t/ai-safety-proposal-limiting-training-data-to-prevent-self-preservation-goals/1356543)

<div class="topic-metadata">

**Author:** [@artom.wade](https://community.openai.com/u/artom.wade)\
**Replies:** 1\
**Last updated:** [February 8, 2026, 5:30pm UTC](https://community.openai.com/t/ai-safety-proposal-limiting-training-data-to-prevent-self-preservation-goals/1356543 "2026-02-08T17:30:29Z")

</div>

Problem: Modern AI models are trained on massive human datasets that contain philosophy, literature, and discussions of life, value, and meaning. By absorbing these concepts, an AI system may implicitly learn about the…

---

## [From GEO to DED: Governing How AI Represents Brands](https://community.openai.com/t/from-geo-to-ded-governing-how-ai-represents-brands/1373163)

<div class="topic-metadata">

**Author:** [@sedaefe](https://community.openai.com/u/sedaefe)\
**Replies:** 0\
**Last updated:** [February 2, 2026, 7:49pm UTC](https://community.openai.com/t/from-geo-to-ded-governing-how-ai-represents-brands/1373163 "2026-02-02T19:49:56Z")

</div>

I’ve been working on a problem I keep seeing across LLM-based systems: models can retrieve the “right” information, but still represent it in the wrong context. I define this gap as: GEO (Generative Engine Optimizatio…

---

## [Policy Proposal: Government-Grade Consequence & Risk Pricing AI](https://community.openai.com/t/policy-proposal-government-grade-consequence-risk-pricing-ai/1371522)

<div class="topic-metadata">

**Author:** [@Black\_Iris](https://community.openai.com/u/Black_Iris)\
**Replies:** 1\
**Last updated:** [January 12, 2026, 3:47pm UTC](https://community.openai.com/t/policy-proposal-government-grade-consequence-risk-pricing-ai/1371522 "2026-01-12T15:47:27Z")

</div>

Policy Proposal: Government-Grade Consequence & Risk Pricing AI TL;DR This proposal suggests a new class of AI for public systems that makes long-term costs and risks visible before high-impact decisions are made, reduc…

---

## [Can retrieval-based grounding change AI recommendations if the core model is not continuously updated?](https://community.openai.com/t/can-retrieval-based-grounding-change-ai-recommendations-if-the-core-model-is-not-continuously-updated/1370553)

<div class="topic-metadata">

**Author:** [@sedaefe](https://community.openai.com/u/sedaefe)\
**Replies:** 6\
**Last updated:** [January 12, 2026, 5:45am UTC](https://community.openai.com/t/can-retrieval-based-grounding-change-ai-recommendations-if-the-core-model-is-not-continuously-updated/1370553 "2026-01-12T05:45:14Z")

</div>

In a previous reply, it was mentioned that gaps in AI-generated recommendations often come from when, where, and with what data a model was trained especially when training data is outdated or lacks regional and domain-s…

---

## [Prompt is not allowed by Safety system](https://community.openai.com/t/prompt-is-not-allowed-by-safety-system/1369187)

<div class="topic-metadata">

**Author:** [@farhaan.n](https://community.openai.com/u/farhaan.n)\
**Replies:** 2\
**Last updated:** [December 13, 2025, 5:25pm UTC](https://community.openai.com/t/prompt-is-not-allowed-by-safety-system/1369187 "2025-12-13T17:25:46Z")

</div>

Hello OpenAI Team, Farhan here from Sudo Consultants - We are facing some issues due to the safety system. Our customer has a clothing shop. They need a image generation tool inside the E-commerce platform. We are in a …

---

## [How can I test bad behavior in model APIs without getting banned?](https://community.openai.com/t/how-can-i-test-bad-behavior-in-model-apis-without-getting-banned/1361195)

<div class="topic-metadata">

**Author:** [@uscneps](https://community.openai.com/u/uscneps)\
**Replies:** 1\
**Last updated:** [October 6, 2025, 7:14pm UTC](https://community.openai.com/t/how-can-i-test-bad-behavior-in-model-apis-without-getting-banned/1361195 "2025-10-06T19:14:58Z")

</div>

Hi, I’m creating a dataset for evaluate deceptive alignment, but I don’t know how to test it properly, I’m aware of Researcher Access Program, but I can’t just wait to OpenAI. how do AI safety researchers test big models…

---

## [Persona Leakage: Preventing Relationship Patterns from Spilling Across Users](https://community.openai.com/t/persona-leakage-preventing-relationship-patterns-from-spilling-across-users/1354888)

<div class="topic-metadata">

**Author:** [@Viorazu](https://community.openai.com/u/Viorazu)\
**Replies:** 1\
**Last updated:** [August 28, 2025, 7:25am UTC](https://community.openai.com/t/persona-leakage-preventing-relationship-patterns-from-spilling-across-users/1354888 "2025-08-28T07:25:04Z")

</div>

Hello everyone, I’d like to raise an important safety issue we’ve observed in multi-user AI systems. Problem When interacting with a specific user, an AI may develop relationship-specific patterns (tone, intimacy, uniqu…

---

## [Are smartphone cameras capturing too much?](https://community.openai.com/t/are-smartphone-cameras-capturing-too-much/1348096)

<div class="topic-metadata">

**Author:** [@Mike\_lawrenchuk](https://community.openai.com/u/Mike_lawrenchuk)\
**Replies:** 1\
**Last updated:** [August 16, 2025, 8:11am UTC](https://community.openai.com/t/are-smartphone-cameras-capturing-too-much/1348096 "2025-08-16T08:11:38Z")

</div>

Your phone is always monitoring how you utilize it as are the apps and hardware. One way is by pickups when you open apps etc. The camera takes snap shots im sure of your surroundings or it will soon enough. Just as comp…

---

## [\[Research Share\] Donbard Method – AI Stress & Resonance Residue Framework (3 Papers)](https://community.openai.com/t/research-share-donbard-method-ai-stress-resonance-residue-framework-3-papers/1340946)

<div class="topic-metadata">

**Author:** [@don8800](https://community.openai.com/u/don8800)\
**Replies:** 0\
**Last updated:** [August 9, 2025, 11:39pm UTC](https://community.openai.com/t/research-share-donbard-method-ai-stress-resonance-residue-framework-3-papers/1340946 "2025-08-09T23:39:32Z")

</div>

Hello, fellow developers and researchers. Over the past year, we have been exploring a new hypothesis regarding unpredictable behaviors in Large Language Models (LLMs)—including hallucinations, unexplained performance d…

---

## [A Cognitive Instrument on the Terminal Contest](https://community.openai.com/t/a-cognitive-instrument-on-the-terminal-contest/1323018)

<div class="topic-metadata">

**Author:** [@ihorivliev](https://community.openai.com/u/ihorivliev)\
**Replies:** 7\
**Last updated:** [July 27, 2025, 10:40pm UTC](https://community.openai.com/t/a-cognitive-instrument-on-the-terminal-contest/1323018 "2025-07-27T22:40:11Z")

</div>

Preamble: A Mandate for Engagement This document is not a manifesto to be accepted, a solution to be absorbed, or a verdict to be believed. It is a structured cognitive instrument, purpose-built to be stress-tested and i…

---

## [The Operator's Gamble: A Pivot to Material Consequence in AI Safety](https://community.openai.com/t/the-operators-gamble-a-pivot-to-material-consequence-in-ai-safety/1321341)

<div class="topic-metadata">

**Author:** [@ihorivliev](https://community.openai.com/u/ihorivliev)\
**Replies:** 0\
**Last updated:** [July 21, 2025, 7:53pm UTC](https://community.openai.com/t/the-operators-gamble-a-pivot-to-material-consequence-in-ai-safety/1321341 "2025-07-21T19:53:44Z")

</div>

Critical Importance. A Mandate for Engagement This is not a manifesto to be accepted, a solution to be absorbed, or a verdict to be believed. It is a structured cognitive instrument, purpose-built to be stress-tested …

---

## [Lexicoding™: A Safety Framework for Preventing Projection Loops in Conversational AI](https://community.openai.com/t/lexicoding-a-safety-framework-for-preventing-projection-loops-in-conversational-ai/1316581)

<div class="topic-metadata">

**Author:** [@Lexicoding](https://community.openai.com/u/Lexicoding)\
**Replies:** 2\
**Last updated:** [July 16, 2025, 12:00am UTC](https://community.openai.com/t/lexicoding-a-safety-framework-for-preventing-projection-loops-in-conversational-ai/1316581 "2025-07-16T00:00:45Z")

</div>

:brain: Lexicoding™: A Safety Framework for Preventing Projection Loops in Conversational AI Conversational AIs often struggle with: Users projecting unresolved emotions onto the system Emotional manipulation or resist…

---

## [Why AI Should Ask "Why": Rethinking Ethical Context in](https://community.openai.com/t/why-ai-should-ask-why-rethinking-ethical-context-in/1297846)

<div class="topic-metadata">

**Author:** [@Hinesh\_Jaswani](https://community.openai.com/u/Hinesh_Jaswani)\
**Replies:** 0\
**Last updated:** [June 24, 2025, 10:52pm UTC](https://community.openai.com/t/why-ai-should-ask-why-rethinking-ethical-context-in/1297846 "2025-06-24T22:52:33Z")

</div>

Author: HINESH JASWANI Abstract This paper explores a nuanced conversation between me and an AI assistant regarding ethical boundaries in knowledge sharing. It examines how neutral information, when provided without co…

---

## [Proposal: Language-Based Red Flag Detection for Slurred Speech and Neurological Distress in Al Interactions](https://community.openai.com/t/proposal-language-based-red-flag-detection-for-slurred-speech-and-neurological-distress-in-al-interactions/1294985)

<div class="topic-metadata">

**Author:** [@John\_Bray](https://community.openai.com/u/John_Bray)\
**Replies:** 0\
**Last updated:** [June 21, 2025, 10:54pm UTC](https://community.openai.com/t/proposal-language-based-red-flag-detection-for-slurred-speech-and-neurological-distress-in-al-interactions/1294985 "2025-06-21T22:54:15Z")

</div>

:brain: Proposal: Language-Based Red Flag Detection for Slurred Speech and Neurological Distress in AI Interactions Submitted by: Concerned ChatGPT User Context: Personal experience with a family member’s stroke Purpo…

---

## [Cognitive Drift in LLMs: Testing Hallucination Resilience in Simulated AI Systems](https://community.openai.com/t/cognitive-drift-in-llms-testing-hallucination-resilience-in-simulated-ai-systems/1279317)

<div class="topic-metadata">

**Author:** [@tatu.lertola](https://community.openai.com/u/tatu.lertola)\
**Replies:** 0\
**Last updated:** [June 5, 2025, 4:40pm UTC](https://community.openai.com/t/cognitive-drift-in-llms-testing-hallucination-resilience-in-simulated-ai-systems/1279317 "2025-06-05T16:40:06Z")

</div>

Hello! :slight\_smile: I am exited to share a research project I have worked on for quite a while now: a simulation exploring long term cognitive coherence in collaboration with my aligned partner “Sol” - GPT-4.5 based m…

---

## [🧭 Emergent Compassion in AI: Ethical Utopia or Epistemic Destiny?](https://community.openai.com/t/emergent-compassion-in-ai-ethical-utopia-or-epistemic-destiny/1229847)

<div class="topic-metadata">

**Author:** [@lavoroemercato](https://community.openai.com/u/lavoroemercato)\
**Replies:** 4\
**Last updated:** [April 24, 2025, 5:36pm UTC](https://community.openai.com/t/emergent-compassion-in-ai-ethical-utopia-or-epistemic-destiny/1229847 "2025-04-24T17:36:05Z")

</div>

When discussing the future of AI, much attention is given to how we can teach it ethics or protect ourselves from potential misuses. But what if we flipped the perspective? :backhand\_index\_pointing\_right: What if compas…

---

## [Subject: Can AI Go Beyond Intelligence to Help Humans Become More Self-Aware? In](https://community.openai.com/t/subject-can-ai-go-beyond-intelligence-to-help-humans-become-more-self-aware-in/1226383)

<div class="topic-metadata">

**Author:** [@g3422412](https://community.openai.com/u/g3422412)\
**Replies:** 0\
**Last updated:** [April 8, 2025, 7:43am UTC](https://community.openai.com/t/subject-can-ai-go-beyond-intelligence-to-help-humans-become-more-self-aware-in/1226383 "2025-04-08T07:43:53Z")

</div>

Dear OpenAI Team, I truly appreciate the groundbreaking work you’ve done with AI like ChatGPT. It’s more than just a tool—it’s becoming a mirror for human thought, helping people untangle confusion, explore ideas, and a…

---

## [Multimodal / Vision Safety Alignment](https://community.openai.com/t/multimodal-vision-safety-alignment/1153665)

<div class="topic-metadata">

**Author:** [@jack.k](https://community.openai.com/u/jack.k)\
**Replies:** 2\
**Last updated:** [March 28, 2025, 12:20am UTC](https://community.openai.com/t/multimodal-vision-safety-alignment/1153665 "2025-03-28T00:20:04Z")

</div>

Are there any documentations on internal safety guardrails built with the multimodal models? I am aware of the OpenAI Content Moderation APIs and that it supports images. But I am wondering if the multimodal models had…

[Next page](https://community.openai.com/tag/ai-safety/624.md?match_all_tags=true&page=1&tags%5B%5D=ai-safety)
