Hello OpenAI Product, Research, ChatGPT Work, Safety, Legal, and Policy Teams,
I would like to propose a strategic direction for the next generation of ChatGPT and ChatGPT Work.
The goal should be clear:
ChatGPT should not only be the best AI assistant. It should become the best system for completing professional work from idea to production-ready output.
I believe OpenAI has three major opportunities:
1. Regain clear leadership in AI video generation.
2. Surpass Claude in professional document creation and editing.
3. Adopt a precise provenance policy that complies with law without unnecessarily watermarking ordinary AI-assisted professional work.
Together, these could make ChatGPT Work a universal AI production operating system.
PART I — REGAIN LEADERSHIP IN AI VIDEO
AI video is moving beyond short prompt-generated clips.
The next category is multimodal video production using combinations of:
Text
Images
Multiple images
Video references
Audio
Characters
Products
Locations
Motion references
Camera instructions
Storyboards
OpenAI should build one unified system supporting:
Text → Video
Image → Video
Multiple Images + Text → Video
Multiple Images + Text + Audio → Video
Video → Video
Video Editing
Video Extension
Storyboard → Video
Character Reference → Video
Product Reference → Video
Environment Reference → Video
Motion Reference → Video
Camera Reference → Video
Users should not need different systems for each workflow.
A. Semantic reference control
Users should be able to assign roles to references:
Image 1 = Character face
Image 2 = Clothing
Image 3 = Product
Image 4 = Environment
Video 1 = Camera motion
Video 2 = Body motion
Then give instructions such as:
“Use Image 1 only for facial identity.”
“Preserve the product geometry exactly.”
“Use Video 1 only for camera movement.”
This would provide much stronger production control.
B. Persistent scene and identity memory
Characters, products, objects, and environments should remain consistent across shots.
The system should maintain persistent information about:
Character identity
Face
Body proportions
Clothing
Products
Objects
Environment layout
Lighting
Camera state
Narrative state
Every shot should reference the same Scene Memory or World State.
C. 3D-aware world understanding
Future video models should maintain stable representations of:
Depth
Geometry
Object permanence
Occlusion
Camera trajectory
Lighting
Physics
Motion
This should reduce:
Identity drift
Morphing objects
Changing environments
Broken anatomy
Impossible interactions
Camera discontinuities
D. Professional camera control
Users should be able to control:
Lens
Focal length
Aperture
Depth of field
Camera height
Camera angle
Pan
Tilt
Dolly
Crane
Orbit
Zoom
Handheld motion
Focus pull
ChatGPT should translate cinematographic language into precise generation controls.
E. Native audiovisual generation
Video and audio should share timing and semantic context.
The system should jointly support:
Dialogue
Lip synchronization
Ambient audio
Foley
Music
Sound effects
Spatial acoustics
F. From clip generation to full production
ChatGPT Work should support:
Project
→ Scene
→ Sequence
→ Shot
→ Frame
A request such as:
“Create a 90-second commercial.”
could produce:
Brief
Script
Storyboard
Shot list
Characters
Product references
Camera plan
Video shots
Dialogue
Sound
Music
Final edit
This would turn AI video generation into AI video production.
G. AI Video Studio inside ChatGPT Work
ChatGPT Work should eventually provide:
Assets
Characters
Products
Locations
Storyboard
Timeline
Shots
Audio
Versions
Exports
Users could edit conversationally:
“Use the same character as shot 2.”
“Change shot 7 to a 50mm lens.”
“Keep the product unchanged.”
“Regenerate only frames 40–80.”
“Replace the background.”
H. Professional outputs
Future exports should include:
Final video
Individual shots
Image sequences
Alpha channels
Audio stems
Subtitles
EDL
OpenTimelineIO
Project metadata
Where practical, integrate with workflows involving:
Premiere Pro
After Effects
DaVinci Resolve
Final Cut Pro
Blender
Unreal Engine
I. Automatic video verification
ChatGPT Work should use:
Generate
→ Inspect
→ Detect errors
→ Correct
→ Re-render
→ Validate
Evaluation should include:
Identity consistency
Reference fidelity
Temporal stability
Anatomy
Physics
Object permanence
Camera continuity
Lip sync
Audio sync
Product consistency
The user should not have to manually find every failure.
J. Benchmark goal
OpenAI should explicitly target:
#1 Text-to-Video
#1 Image-to-Video
and develop stronger internal tests for:
Multi-reference generation
Identity persistence
Product consistency
Camera control
Narrative continuity
Audiovisual synchronization
Long-form generation
The goal should not merely be impressive short clips.
It should be professional production reliability.
PART II — SURPASS CLAUDE IN PROFESSIONAL DOCUMENTS
Claude has developed a strong reputation for long-form document work.
OpenAI should treat professional document production as a major competitive category.
The key principle is:
A professional document is not a long chatbot response.
The document itself is the product.
A. Stronger long-context coherence
ChatGPT should maintain across long documents:
Instruction consistency
Entity consistency
Argument consistency
Terminology consistency
Evidence tracking
Citation consistency
Cross-references
Narrative continuity
Context-window size alone is not enough.
B. Document-native reasoning
ChatGPT should reason directly about:
Headings
Paragraphs
Tables
Figures
Captions
References
Footnotes
Styles
Page breaks
Headers
Footers
Numbering
Appendices
These should be structured objects, not formatting added after writing.
C. Internal Document Representation
A structured representation could include:
Document
├── Sections
├── Paragraphs
├── Tables
├── Figures
├── Citations
├── References
├── Styles
└── Layout Constraints
This could compile into:
DOCX
PPTX
HTML
Markdown
ODT
D. Multi-agent document production
Professional documents should use specialized roles such as:
Research Agent
Writer
Editor
Evidence Verifier
Citation Verifier
Layout Agent
Visual QA Agent
Final Reviewer
The workflow should automatically perform:
Draft
→ Critique
→ Fact verification
→ Citation verification
→ Editing
→ Formatting
→ Rendering
→ Visual QA
→ Final revision
E. Rendered-document QA
ChatGPT should inspect the actual final pages and automatically detect:
Blank pages
Broken tables
Text overflow
Tiny fonts
Large empty spaces
Bad page breaks
Misaligned figures
Caption separation
Header/footer problems
Poor visual balance
F. Surgical editing
When modifying an existing file, ChatGPT should preserve what the user did not ask to change.
This includes, where appropriate:
Images
Formatting
Styles
Tables
References
Captions
Manual edits
Document structure
G. Version control
Document projects should support:
Diff
Accept
Reject
Rollback
Compare Versions
Semantic Comparison
H. Citation integrity
Citations should be linked to:
Source
Claim
Evidence location
Publication date
Reference format
The system should detect:
Unsupported claims
Citation mismatch
Incorrect quotations
Broken references
Missing citations
Contradictory evidence
I. Define success by professional usability
The target should not be:
“ChatGPT can generate a DOCX.”
The target should be:
“A professional can send the document directly to a professor, client, executive, board, investor, government official, or customer without rebuilding it.”
OpenAI should aim to surpass Claude in both benchmarks and professional human preference.
PART III — DO NOT MAKE INVISIBLE AI WATERMARKING THE GLOBAL DEFAULT
I strongly encourage OpenAI not to make persistent invisible AI watermarking the global default for ordinary professional text and documents.
This does not mean ignoring law.
The recommendation is:
Comply fully where legally required, while avoiding unnecessary invisible watermarking where the law does not require it.
The long-term principle should be:
No unnecessary invisible watermark.
No unnecessary hidden AI-generated marker.
No unnecessary AI-generation metadata in ordinary professional files.
No permanent forensic signal merely because ChatGPT assisted with writing, editing, formatting, translation, research, or document production.
A. AI involvement is not AI authorship
There is a major difference between:
Human authored
Human authored + AI proofread
Human authored + AI edited
Human directed + AI drafting
Human-AI co-created
Primarily AI generated
Fully AI generated
A binary hidden watermark cannot accurately represent this continuum.
A watermark may show that AI influenced wording.
It does not establish:
Who created the idea
Who conducted the research
Who constructed the argument
Who verified the facts
Who edited the document
Who accepts responsibility
Who owns the intellectual contribution
Provenance should describe process, not redefine authorship.
B. AI assistance should not become a permanent forensic property
Nobody expects a Word document to permanently announce:
“This sentence was corrected by Microsoft Editor.”
A photograph does not normally announce:
“Photoshop adjusted this image.”
Source code does not permanently announce:
“Autocomplete helped write this.”
Generative AI is increasingly becoming another professional productivity tool.
If every artifact touched by AI receives a persistent hidden marker, eventually a large percentage of ordinary human digital work could become machine-marked.
That will become less informative, not more informative.
C. Maximum compliance, minimum unnecessary friction
AI laws and enforcement vary across jurisdictions.
OpenAI should avoid automatically applying the strictest implementation everywhere simply because it is technically easier.
Instead, provenance should be:
Jurisdiction-aware
Risk-aware
Content-aware
Legally scoped
If legally required:
Apply the required marking.
If not legally required:
Do not automatically apply persistent invisible marking to ordinary professional work.
If content creates serious deception or impersonation risk:
Apply stronger safeguards.
D. Prevent regulatory arbitrage
If responsible AI providers create significantly more friction than competing systems, users may not stop using AI.
They may instead:
Switch providers
Use local models
Use open-weight models
Use models hosted elsewhere
Move workflows into less accountable systems
This could reduce transparency rather than improve it.
The safest AI ecosystem should also be attractive enough that professionals want to remain inside it.
E. Separate high-risk synthetic media from normal professional work
A fabricated video of a political leader is fundamentally different from a business report whose grammar was corrected by AI.
Strong provenance makes particular sense for:
Deepfakes
Synthetic impersonation
Fraud
Fabricated evidence
Deceptive public-interest media
Ordinary work such as:
Reports
Research documents
Business proposals
Technical documentation
Emails
Translations
Proofreading
Formatting
should not automatically receive identical treatment.
F. Human responsibility should matter
If a professional:
Reviews the entire artifact
Checks claims
Verifies citations
Rewrites material
Approves the final content
Accepts responsibility
that workflow is meaningfully different from anonymous one-click AI generation.
ChatGPT Work could support:
Human Reviewed
or:
Professionally Reviewed
Where legally permissible, this could allow clean professional export.
The workflow would be:
AI assistance
→ human review
→ human revision
→ human responsibility
→ final professional artifact
G. Avoid unnecessary metadata
For formats such as:
DOCX
PPTX
XLSX
TXT
Markdown
OpenAI should avoid unnecessary metadata whose main purpose is permanently classifying ordinary professional work as AI-generated when such marking is not legally required.
If provenance is required, include it.
If it is optional, allow appropriate user control.
H. Invisible watermarking has a durability problem
If watermark signals can disappear after extensive rewriting or transformation, they cannot provide an absolute historical record.
This creates an imbalance:
Ordinary compliant users may remain detectable.
Deliberate evaders may transform content until the signal disappears.
That is not an ideal long-term trust architecture.
I. Prefer proportional provenance
A better framework is:
High provenance where authenticity risk is high.
Required provenance where law requires it.
Minimal provenance where risk is low.
Optional provenance where users benefit from it.
No unnecessary persistent marking where it is not required.
J. Build better alternatives
OpenAI should continue developing richer options such as:
Content credentials
Cryptographic provenance
Signed generation receipts
Enterprise audit logs
Publication-level disclosure
Privacy-preserving server records
Human-review attestations
These can provide more useful information than a binary hidden watermark.
K. Clean Professional Export
I recommend a ChatGPT Work policy called:
Clean Professional Export
This would not remove legally required watermarks.
It would define situations where unnecessary AI-generation markings are not inserted in the first place.
Possible conditions:
No legal requirement for persistent marking
Low deception risk
Human review completed where needed
User accepts responsibility
Artifact is ordinary professional work
The principle is:
Do not add unnecessary provenance
rather than:
Add provenance and later remove it.
FINAL WATERMARKING RECOMMENDATION
Do not make invisible AI watermarking the permanent global default for professional text and documents.
Where legally required, comply fully.
Where it is not required, avoid unnecessary permanent marking.
Develop:
Jurisdiction-aware compliance
Risk-aware provenance
AI-generated versus AI-assisted distinctions
Human-review recognition
Clean Professional Export
Strong safeguards for deceptive synthetic media
The principle should be:
Transparency where necessary.
Protection where authenticity is at risk.
Compliance where required.
No unnecessary invisible watermark merely because AI was used as a professional tool.
PART IV — CHATGPT WORK AS A UNIVERSAL PRODUCTION LAYER
These proposals converge on a larger opportunity.
Professionals currently move between:
ChatGPT
Claude
Gemini
AI video generators
Image generators
Word
PowerPoint
Excel
Photoshop
Premiere
Figma
Canva
Notion
Research tools
because different applications own different stages of production.
ChatGPT Work could coordinate the entire workflow:
Research
→ Think
→ Write
→ Design
→ Generate Images
→ Generate Video
→ Edit
→ Verify
→ Export
→ Publish
For example:
“Create a complete launch campaign for this product.”
ChatGPT Work could:
Research the market
Analyze competitors
Develop positioning
Write strategy
Create a proposal
Build a presentation
Generate imagery
Generate a commercial
Create social variants
Generate subtitles
Localize content
Verify branding
Package deliverables
Export editable files
That is not simply a chatbot.
It is an AI production operating system.
PART V — DEVELOPMENT ROADMAP
Phase 1 — Benchmark Leadership
Target leadership in:
Text-to-Video
Image-to-Video
Multi-reference video
Professional document benchmarks
Phase 2 — Document-Native ChatGPT Work
Build:
Document IR
Specialized document workflows
Multi-agent writing and verification
Citation validation
Rendered QA
Version control
Professional Office/PDF export
Goal:
Surpass Claude in professional document quality.
Phase 3 — Unified Multimodal Video
Support:
Text
Images
Multiple images
Video
Audio
Storyboards
Semantic references
in one architecture.
Phase 4 — Persistent World Memory
Maintain consistent:
Characters
Products
Locations
Objects
Lighting
Camera state
Narrative state
Phase 5 — ChatGPT Work Video Studio
Integrate:
Script
Storyboard
Assets
Timeline
Generation
Editing
Audio
Versioning
QC
Export
Phase 6 — Automatic Production Verification
Documents:
Facts
Citations
Formatting
Layout
Consistency
Videos:
Identity
Reference fidelity
Temporal coherence
Audio sync
Physics
Camera continuity
Phase 7 — Proportional Provenance
Implement:
Jurisdiction-aware compliance
Risk-tiered provenance
AI-generated versus AI-assisted distinctions
Human-review recognition
Clean Professional Export where legally permissible
Strong deepfake safeguards
Phase 8 — Universal Production Orchestration
Combine:
Research
Documents
Images
Video
Audio
Code
Data
Design
Automation
Agents
Verification
Publishing
into one coordinated environment.
PART VI — SUCCESS CRITERIA
Success should mean:
#1 Text-to-Video performance
#1 Image-to-Video performance
State-of-the-art multi-reference video
State-of-the-art identity and product consistency
Top professional document benchmark performance
Higher professional preference than Claude
Strong long-context coherence
High citation reliability
Low hallucination rates
Low revision burden
Professional editable outputs
Minimal unnecessary provenance friction
High user trust
FINAL RECOMMENDATION
The standard should no longer be:
“Can ChatGPT do this?”
It should be:
“Is ChatGPT the best system in the world for doing this?”
For video:
Outperform the strongest specialized video models and provide superior professional control.
For documents:
Surpass Claude through document-native reasoning, long-context coherence, specialized tooling, multi-agent verification, and better artifact QA.
For provenance:
Comply fully with applicable law without unnecessarily turning ordinary AI-assisted professional work into permanently machine-marked content.
The winning AI platform will be the system where a professional can begin with an idea and finish with a complete, editable, verified, production-ready result.
Research.
Documents.
Images.
Video.
Audio.
Code.
Data.
Design.
Automation.
All inside one intelligent production environment.
The evolution should be:
Chatbot
→ Assistant
→ Agent
→ Production Partner
→ AI Production Operating System
OpenAI should not build only the best chatbot.
It should build the best AI production system in the world.
The most capable system.
The highest-quality output.
The strongest professional workflow.
The greatest user control.
The lowest unnecessary friction.
The highest trust.
From idea to finished result.
Inside one system.
Thank you for considering this proposal.