nVidia’s new text-to-video AI shows an insane rate of progress

“Will Smith eats spaghetti” seemed like a bit of a joke when text-to-video generative AI was just a month or two ago, but now nVidia seems to be blowing its previous efforts out of the water. demonstrated a new system that looks like The pace of progress here is astonishing.

Presented at the IEEE Conference on Computer Vision and Pattern Recognition 2023, nVidia’s new video generator started as a latent diffusion model (LDM) trained to generate images from text, and then what was used to generate the images. Introduce an extra step that you want to animate. It learned from studying thousands of existing videos.

This adds time as a tracked dimension, and the LDM is responsible for estimating what might change in each region of the image over a given period of time. Create a number of keyframes throughout the sequence and use another LDM to interpolate the frames between the keyframes to produce similar quality images for all images in the sequence.

nVidia tested the system with low-quality dashcam-style footage and found that at a resolution of 512 x 1024 pixels, it was able to produce this kind of video for several minutes in a “temporally consistent” manner. . This fast-moving field.

However, it works at much higher resolutions and can also work with a very wide range of other visual styles. generated. Each of these videos contains 113 frames and is rendered at 24 fps, so they are approximately 4.7 seconds long. Pushing much longer than that in terms of total time makes things seem to break, leading to even more weirdness.

They were clearly AI-generated, and there are still many strange mistakes to be found. It’s also obvious where the keyframes are in a lot of the videos, and you’ll see strange speed and slow motion around them. It’s an incredible leap.

In this formative age when we’re just starting to understand how images and videos work, it’s pretty cool to see these amazing AI systems. Think of all the things they need to understand. One is in 3D space, and think about how the realistic parallax effect continues as you move the camera. Then there’s how the liquid behaves, from the splashing spectacle of waves crashing against rocks at sunset, to the gently spreading wake left by a swimming duck, to how steamed milk mixes when poured into coffee. Up to how it foams.

Then there are the subtly changing reflections on the spinning bowl of grapes. Or how the flower garden flutters in the wind. Or the flames spread along the campfire logs and lick the sky. Needless to say, a huge variety of human and animal behaviors need to be recreated.

In my eyes, it epitomizes rapid progress across the gamut of generative AI projects, from language models like ChatGPT to image, video, audio, and music generation systems. At first glance at these systems, they seem ridiculously impossible, but then you realize they’re amazingly good and very useful. We’re somewhere between hilariously bad and surprisingly good right now.

Due to the way this system is designed, nVidia seems to be trying to offer the world’s first ability to grab an image as well as a text prompt. This means you may be able to upload your own images, or images from a specific AI generator. They evolved into videos. For example, if I had a lot of pictures of Kermit the Frog, I could generate a video of him playing the guitar, singing, and typing on his laptop.

So it looks like we’ll be able to daisy chain these AIs together to create a ridiculously integrated form of entertainment relatively soon. A language model might write a children’s book and have an image generator explain it. Such models then take the text of each page and use it to animate illustrations, while other AIs provide realistic sound effects, voices, and fine-tuned musical soundtracks. A children’s book turned into a short film, fully retaining the visual feel of the illustrations.

From there, the entire environment for each scene is modeled in 3D to create an immersive VR experience or a story-driven video game. Once that happens, you can talk directly to any character about anything you want, because custom AI characters are already capable of incredibly complex and informative verbal conversations.

The craziest thing of all is creating prompts for the overarching AI to probably be a lot better than you or better results from other AIs in the chain, assessing the results and asking for corrections – So these entire projects are probably generated from a single prompt and a few iterative change requests. This is truly amazing. At some point closer than you can imagine, you’ll be able to make the leap from conceptual idea to fully fleshed out entertainment franchise in minutes.

nVidia is now treating this system as a research project rather than a consumer product. Perhaps the company has little interest in paying the processing costs of open systems, which can be substantial. You may also be trying to avoid copyright issues that may arise from your training dataset. And when these systems start churning out realistic videos that never happened, it’s clear there are other dangers to avoid.

But make no mistake, this stuff is coming, and it’s coming at a rate that feels thrilling or terrifying.

Source: nVidia



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *