Feature Request: Native Video Upload Support

Feature Request: Native Video Upload Support

I’d love to see ChatGPT support uploading video files (such as .mp4 and .mov) the same way it currently supports images and documents.

Many tasks require understanding what happens over time, not just in a single frame. Native video support would make ChatGPT much more useful for:

  • Analyzing movement or body language
  • Reviewing sports or exercise form
  • Diagnosing mechanical issues from sounds and motion
  • Understanding pet behavior
  • Summarizing recordings or events
  • Providing feedback on presentations, interviews, or speeches

Ideally, ChatGPT could:

  • Accept common video formats.
  • Automatically transcribe spoken audio.
  • Analyze both visual and audio information.
  • Allow questions about specific timestamps (for example, “What happened at 2:14?”).
  • Summarize entire videos or focus on selected segments.

This would be a natural extension of the existing image and document upload features and would greatly expand the kinds of real-world problems ChatGPT can help solve.

Hey @aethryxx, welcome to the community! This is a really useful suggestion, and the examples make the need pretty clear.

A lot of real-world questions depend on motion, sound, timing, or changes across a scene, so video upload support would open up use cases that single images can’t fully cover. Things like sports form, presentations, mechanical issues, pet behavior, or timestamp-specific questions would all benefit from being able to analyze the full clip.

I can’t share a timeline right now, but I’ll pass this feedback along internally.

- Sunny