MARATTO

article

An Object-Oriented Deep Learning Method for Video Frame Prediction

Abstract

Predicting the next frames in a video sequence is a fundamental task in computer vision and video analysis. A great deal of success has been achieved in this task thanks to recent advances in deep learning. However, most existing techniques are tested in constrained settings such as static camera, static scene, or few moving elements. In this paper, we develop a new method of predicting the next video frame based on the estimation of transformation parameters for each object in the scene. Our method disentangles camera intrinsic motion from objects motion. First, it estimates the projection parameters for each object within the scene to align with its view in the next frame. Then, each object’s motion is captured by estimating an affine transform that aligns the warped view of the object with the ground truth next frame. The sequence of projective and affine transforms is then fed to a deep neural network (a transformer network) to predict next frames. Experiments using the Indian Driving Dataset (IDD) demonstrate the merits of the proposed method.

Research topics

  • Advanced Image Processing Techniques
  • Video Analysis and Summarization
  • Advanced Vision and Imaging

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/icasspw62465.2024.10626448

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.