The title explains it already.
The current implementation of “Computer Use” is incredibly primitive at best - more like a demo or proof of concept than any kind of functional feature.
Most actions and activities take place in a sequence of time, with all kinds of interactions and visuals that are not available from a static screenshot format. Furthermore its not really capable of “Using” anything in this way - except perhaps a simple webpage.
The model can recognize Video input natively - it already has this ability. It should not be very difficult to implement a “video capture/viewer” which runs a Timestamped chronological log, along with a Timestamped log of Keyboard Inputs and Cursor Positions with Mouse Clicks.
This would be an actual and true situation of “Computer Use” that would be actually valuable in terms of building and error checking.