ChatGPT should offer a local preprocessing/anonymization mode for sensitive documents, especially medical, legal and forensic documents. The mode should detect and pseudonymize names, dates of birth, addresses, file numbers, hospital IDs, rare location/time combinations, image metadata and other re-identification risks before any content is uploaded to OpenAI servers. This would be highly relevant for physicians, legal professionals, researchers and public institutions working with confidential records.
Thanks for sharing this, @Schmallo. This is a thoughtful suggestion, especially for healthcare, legal, forensic, and public-sector use cases where re-identification risks matter.
Part of this is already supported in healthcare through ChatGPT Health, which includes additional privacy protections and keeps health conversations separate from model training. For organizations, ChatGPT Business and Enterprise also provide strong data privacy controls and don't train on business data by default.
That said, the specific idea of local anonymization and pseudonymization before any data is uploaded isn't broadly available today, particularly for legal and forensic workflows. That feels like a distinct gap and a useful feature request.
We're sending this to the team for logging so it can be considered alongside similar privacy-focused feedback.
-Mark G.