Three to six seconds. No audio. Resolution below what the same platform's still images produce. Motion that is convincing for small movements β a turn of the head, hair, a shift in expression β and falls apart on anything involving hands, a second figure, or a change in camera position.
The failure mode is specific and consistent: the first second is good, the model loses track of the subject's identity somewhere in the middle, and the last second is a different person. This is the same identity-persistence problem that makes still-image consistency hard, except a video shows you the drift happening in real time instead of letting you discard the bad generations.
Text and hands, already the weakest part of still generation, are worse in motion because errors now flicker frame to frame.