VideoPoet by Google
Transforming language models into video generators.
About VideoPoet by Google
"VideoPoet, by Google Research, is a groundbreaking innovation in video creation, focusing on generating dynamic, engaging, and realistic movements.This revolutionary tool converts autoregressive language models into a top-tier video generator. Featuring the MAGVIT V2 video tokenizer and SoundStream audio tokenizer, it converts images, videos, and audio clips of varying lengths into a series of distinct codes within a unified vocabulary.These codes are linked with text-based language models, enabling seamless integration with other forms of media like text. Embedded within this tool is an autoregressive language model that comprehensively learns from video, image, audio, and text inputs to predict the subsequent video or audio token in the sequence.Moreover, it incorporates diverse generative learning objectives into its training structure, including text-to-video, text-to-image, image-to-video, video frame continuation, video inpainting and outpainting, video stylization, and video-to-audio capabilities.VideoPoet can create videos in square or portrait orientations to accommodate short-form content needs. It also supports the generation of audio from video inputs.With its ability to handle a range of video-centric inputs and outputs simultaneously, VideoPoet showcases how language models can effectively synthesize and edit videos with consistent temporal flow."