Intro to AI Art — Lesson 5: Getting Familiar with GAN
Speaker 1 As we enter the second quarter of the twenty first century, human civilization is experiencing an accelerated pace of evolution. A technological and socioeconomic singularity is ahead of us, and it's driven by artificial intelligence and blockchain technology. NPEAK equips entrepreneurs and business professionals with practical knowledge to keep them ahead of the curve in the exponential age. Our live mentoring sessions, on demand training, and exclusive networking opportunities keep you at the cutting edge of Web three, NFTs, the metaverse, decentralized finance, automation, and so much more. Meet top industry leaders during our live mentoring sessions to ask your questions directly or simply follow the recorded sessions in your own time. In PEEK, inclusive, inspired, in the know, in this together.
Speaker 0 GM GM, welcome everyone. Speakers, artists, DJs, happy to see you here today. As more people are trickling in, please as always head over to the polls. We have some poll questions. And I am very excited for today's session because Benjamin is back for, I believe, it's his fifth session already. Yeah. Okay. Got that right. He has some very exciting AI news to share with us today that was not announced before from the past twenty four hours, and we're going through, yeah, really exciting session about GANs today. And to be honest, please enlighten me because I had no idea about GANs before. So I'm very excited to learn myself. And all your questions, please put them in the question tab, and please head over to the polls. And then I'm gonna give it over to Benjamin. And actually, I know here in the audience, Andrew, Nat, welcome.
You both know Ben already. But for anybody who might be watching on demand, Benjamin Meadows is an AI artist. He's been with AI Art for the past nine months. He's been selling his own NFTs, and also he created a special NFT artwork for the n for the Impic Goal Getters, a challenge we run-in our Discord, and he's been sharing this in his past four sessions, how he created this for some really cool hands on learning. So please also take a look at his previous sessions. And with that, I'm excited to learn. So, Benjamin, it's off to you. Well,
Ben good morning, everyone. Yep. So today, we're going over GANs. It's a little bit funny because GANs came out sort of before all the models we've been dealing with back in 2014. So we're kinda going back to the future because there's some new developments with GANs, and we'll talk about that in just a minute. You know, I always like to start out with an interesting AI photo. Today, we have Harry Potter. If Wes Anderson had filmed the movies, Pretty amazing. This is mid journey version five, but pretty crazy what you can do now. It's getting really realistic and very cohesive. You know? You can there's not a lot of artifacts. There's not a whole lot of weird things going on. This looks like one would imagine Harry Potter if Wes Anderson directed it. Another note for today's session before we dive into it, if you go to the references, bmeadows.xyz/npeak, it'll take you to a GitHub repo.
If you go to the a I r dash references file and you go down to lesson five, there are a lot of references for this lesson, and a lot of them are interactive. Last time when we talked about ControlNet, I included a couple of hands on demos so that you could actually use ControlNet. This time, there's probably a dozen or so. So I would encourage you, you know, as we're going through some of these things and looking at it, if you're really excited about something, hop in, look at one of the demos, and play around with it. And then at the end, hop on stage, and let's talk about it and see what you see what you did and see how it's working. So what are we doing today? So the main point of this lesson is to go over generative adversarial networks or GANs.
But I also wanna talk a little bit about some new developments in mid journey and stable diffusion just in the last two weeks. And then let's go through once we talk about what a GAN is, let's talk about how they're actually used nowadays, and then I've got a couple surprises at the end. So let's dive into it. We've had four lessons before this one. Lesson one, kind of a comprehensive what is AIR, where does it come from. Lesson two, we looked at the three major latent diffusion models or diffusion models. You've got DALL E two, mid journey stable diffusion, and then we dove into those in parts one and two, you know, kinda got our hands dirty with how they actually work. We didn't really get our hands dirty because I'm always worried that it'll be like the cooking show where you put the turkey in and, you know, in the real world, you wait forty minutes for the turkey to be ready.
So we kinda we kinda did the, you the old cooking trick, and I already did some of the some of the practical work on the front end. So this is our last lesson that's scheduled. I've got ideas for another five or six lessons, But I'm gonna take a break for the next month or two, and then we'll kinda come back after that probably and look at some more things. So alright. So we're gonna dive right into mid journey, stable diffusion news, but I wanted to stop for a second and sort of reflect on the left. This was Midjourney's version one model from last July. Here we are not even, what, nine, ten months later, April 2023, and this is mid journey's version five model on the right.
So sometimes it's just amazing to take a step back and look at how far AI generative artist come generated art has come in the last nine, ten months, from something that's sort of incoherent, doesn't look like a person, doesn't look like a picture, to something that you would have a very hard time telling on the right, those aren't actual photographs. Just pretty amazing, how fast we're moving. So probably be even more amazing to see how far we come in the next nine, ten months. So what's going on? What's new? Midjourney released three new styles for Neji version five, sort of their anime themed model version five model. A user actually came up with an idea of multi prompt image blending. We'll talk about that in just a minute. Just came out a couple days after our last class.
And then there's some mid journey news about real time drawing, their new version six and seven models, and some more things. And then real quickly, I'll just mention it now. Diffusion b, the stable diffusion UI for Mac, sort of the turnkey application is getting ready to add control net. They're already beta testing it, and that'll be built right in. I'm sure if they're doing that, other sort of turnkey stable diffusion, GUI installs are also gonna be including ControlNet and some of the newer features. So it'll be a little less, you know, hacker ish. You won't have to hop on to a website or install via the command line to use some of these new things. Alright. So let's look at Nidjie, these styles real quick. I I love kinda show and tell, seeing what it looks like.
So that first of all, a few days back, they added two new styles, a cute style and an expressive style. The expressive is slightly more three d and westernized. The cute style is very cute. We'll see that in just a second. And then recently, like, within the last couple days, they added the scenic style, which is more for beautiful backgrounds or characters in beautiful environment. So let's look at that real quickly. Here we have on the left mid journey's default, which is still version four as they're working out some kinks in version five, and I gave it the prompt, a lion looking out over the savannah. You can see these look very much like mid journey version four, sort of very stylized. And on the right, we have a lion looking out over the savannah in version five. Here, you can tell that this is definitely version five.
It's a little more photorealistic, more detail. And then let's start looking at these new Nidji styles real quick. So on the left is the basic Nidji version five style. Up on the top right and the left almost looks like a scene from a Disney movie or the Lion King. They still have a little more of that realistic aesthetic coming, you know, because of version five. If we look on the right, this cute style is indeed, as mid journey warned us, very cute. Almost looks like a children's cartoon or a children's anime on the top right there. So then if we look at the expressive style on the left, it is more toned down. It is more western cartoony. And then on the right, it does give us a very scenic and less western style.
So, anyway, lot of variation here just by adding this style parameter to Nidji, and it just shows you how fast mid journey is iterating. Now you've got version four, version five, Nidji version four, Nidji version five, these different styles. So it's almost experiencing sort of exponential growth in how fast and how far you can customize it. So just a really cool little thing. So multi prompt image blending. So this came out on April 15. This is just a couple days after I taught our last class, and this is what's cool. This isn't an official feature. This is a user that's been using some of the new features, like the blending and the describe command that we just looked at just just two weeks ago, and he found a new way to use this to produce new things. So let's just dive into what he figured out, and let's look at it real quickly.
So this guy said, hey. Sometimes when I wanna blend, you know, two photos, it doesn't do what I want. It comes out malformed or with objects or artifacts. So he takes these two images from Unsplash. And first one's an elegant hotel lobby, second one is a photo of the Royal Greenhouse in Brussels. And he says, I just wanna mix these together, blend them together. So mid journey does what it does, and it comes out with these two photos. And his complaint here is if you look at the ceiling, there's some weird artifacts, the walls, but they're also just very similar, very same, and very unimaginative if you look at the two images that were blended together. So he said, you know what? Let's use the describe command to describe the blended image. So he feeds the blended image that mid journey made into mid journey's describe command, which, by the way, uses a slightly different AI model.
So the model that Midjourney uses to produce images is actually a slightly different model than the model it uses to describe images. So he uses that to describe one of the blended images, and it gives him these four different descriptions. So he decided to go with the first option on the list. The lobby of a modern building has several glass ceilings in the style of tranquil garden scapes, neoclassical simplicity, 32 k UHD, Indonesian art, neoclassical influences, rounded green and bronze. And then it gave him these four images below that. And he said that's really cool, but is that the end of this? No. So now you could feed the original two images, the hotel lobby and the greenhouse, into the describe command, giving you two more sets of starter prompts each. Now this is where this gets really interesting. So it's one thing to blend images.
It's another to blend them and then take the description. But now we're gonna use multiprompting and image waiting like we talked about last week. So what's he do? He takes number two and three from each of these sets. He then combines them into a multiprompt, which if you wanna learn more about multi prompts, go back to lesson three where we go over multi prompts in mid journey. So he combines them into a multi prompt with each starter prompt as its own segment. So mid journey is interpreting each of these prompts as a separate item. So now it's deviated a bit from the original blend blended concept, but but look at what it outputs. Now we can also then weight these multi prompts. So now we've got three multi prompts. Right? The original prompt then the two from the blended images. But now we can weight them.
If we wanted to emphasize, for instance, the hotel lobby concept more than the greenhouse concept, we could then weight them. So, anyway, just really, really innovative, smart way to use something that just came out. So I guess all that to say play around with these new features, you know, and new use cases are emerging as we speak. Another interesting thing that's coming up. So Midjourney's got a whole lot in the pipeline. The last couple weeks, I really mind some of their sort of updates and office notes for news. Nick Saint Pierre did a great little recap on their real time drawing feature. So Midjourney is working pretty heavily right now on a real time drawing feature. It's almost like a control net or a sketch to draw, except it looks like it's gonna be real time. And as you draw, it'll be iterating and trying to make the image.
Additionally, they have their version six, which should be coming out early to mid June. And then they already are working on version seven, which should be coming out late summer, and they're saying it's gonna have a significantly better understanding of language, and they're using the words impressive to describe it. Additionally, some other Midjourney news that may or may not be, you know, public, they have a web beta coming out, so you will not need to use Discord to use it. I think a lot of people are gonna be excited about that. They're also working on outpainting. So just like stable diffusion and Dolly already have out painting, mid journey will also have that. No timeline for that. And they're working on repeatable elements. So, you know, how can you use the same characters or color palettes or styles and reuse them as you're making new images?
Let's see. They also are looking at and I think this is interesting. So they're seeing a trend from text only inputs to image prompts or image prompts with text, and I think that is something that we're gonna see as we move forward. And we'll talk about that a little bit more, but mid journey is seeing that coming down the pike, they're already working on that. Alright. So let's talk about GANs. So what are GANs? Very simply, GANs are comprised so it's a generative adversarial network,
Ben n. They're composed of two neural networks, a generator and a discriminator, and they work together to create new data. If you wanna dive into, you know, exactly how they work, go back to week one. And relatively early on, we looked at the history of AIR. We looked at the history of GANs, going all the way back to 2014. Little bit deeper, how do GANs actually work and train? They're a type of deep learning algorithm, two neural networks, the generator and the discriminator. This is the adversarial part of the discriminator, and they compete against each other to generate new data. The generator creates new data based on a given input, and it's trying to create data that's indistinguishable from real data, and the discriminator then evaluates that generated data, and tests it and kinda gives it a pass fail, and they go back and forth, back and forth a number of times, you know, in a number of steps or iterations until eventually the discriminator agrees that what the generator made was close enough, and it passes through the GAN, and you have the the output.
So here's a quick history and timeline of GANs. 2014, they're introduced by Ian Goodfellow in a paper. 2015, there were some early GANs. 2016, conditional GANs look a little more like what we're thinking. By 2017, progressive growing GANs come out, then these, from them comes style GANs a year later, then there are big GANs, and then there's some other things that we'll talk about. Now something interesting is these last three slides were made entirely with AI. I didn't pick the photos. I didn't pick the data. I think I changed the background color on one of these slides. So all three slides were made entirely with AI, which is incredibly impressive also. So let's look at some general purpose GAN models. We're just talking about style GANs here 2018. Let's go through these. What are these? These are GANs that are used to produce images.
Generally, you're not gonna hop on the demo online and and create images with these. They're a little bit larger models. You generally install them on a computer, run them locally. So first one we'll look at quickly is stat GAN, invented in 2017 by a group of Chinese researchers. Then in 2018, StyleGAN was introduced, a paper by NVIDIA. And then 2020, VQGAN, came out clip in 2021, and then they combined them and used them to create images. I'd say that's probably that and StyleGAN are probably the two sort of predominant general purpose GANs at this point. So let's dive in real quickly and look into what they are. So StackGAN, been in 2017. It's a type of progressive growing GAN. It's a two stage model. The first stage generates low resolution images, and and I'm talking very low resolution images, 64 by 64 pixels from text description.
The the second stage takes these low resolution images and then generates higher resolution images with them. Now this might be familiar when we talked about diffusion, you know, how it takes a fuzzy image and makes it less fuzzy. This is different, but it's not entirely different in a sense. So StackGain uses a technique called progressive growing to generate high resolution images, Starts with a small image, increases the size in each iteration. StatGAN itself is a two stage, you know, progressive growing GAN, but you could have more stages, and we'll actually see that a little bit later in the presentation. One of the other things to note is when we were talking back in week one, one of the questions was, why don't we still use GANs? Or why did we move to these diffusion models? What do GANs do well, and what do they not do well?
And one of the things that they don't do well is GANs tend to or have a tendency to collapse and continue to create the same image or just to make noise. They they've got some issues. By doing it this way, by by gradually increasing the resolution, this prevents some of the early issues that we saw with GANs. So here's some outputs from stat GAN, and and you'll notice that they're very, very descriptive. This flower has a long, yellow petals and a lot of yellow anthers in the set A small yellow bird with a black crown and a short black pointed beak, and you have your stage one and your stage two. The other thing to note about these GANs is they generally only work at generating images based on a pretrained dataset. So maybe this dataset has flowers, birds, people.
You know, there there are it is a large dataset, but it it only works on the items that have been classified in that dataset. So it has some constraints that the later diffusion models don't necessarily have, and and we'll talk about why that is a little bit. So StyleGAN comes out a year later. NVIDIA, it is also a type of progressive growing GAN, but it has some improvements that make it technologically different from stack GAN. And if you notice, NVIDIA NVIDIA continued to build on this, and they're all the way on StyleGAN three released just a couple of years ago. One of the biggest improvements is the use of a style transfer technique called adaptive instance normalization. This allows StyleGAN to learn the style of an existing image and apply it to a new image. So some of the constraints we were just talking about, you know, this gets around them.
So this makes it possible to create images with a wide variety of styles from realistic to cartoony. So we're not stuck with the bird looking like a bird. Now we could have a cartoon bird or an anime bird. Another improvement is the use of a noise vector to add variation to the generated images. This also prevents the images from becoming too repetitive or boring. So we talked about the last one used a couple of stages to scale up the resolution to do this. This is much more similar to a diffusion model. They're literally adding noise into the scan. So let's look at some results. So here we have some more impressive results, obviously, and you can also see some of the variation, right, in these different rows that they're similar, but there's variation. So huge jump from stat GAN twenty seventeen to style GAN twenty eighteen, and then as we move forward to VQGAN and CLIP.
So this is a combination of two different types of neural networks, and they're used to generate images from text descriptions. VQGAN stands for vector quantized GAN, while CLIP is a contrastive language image pretraining model. In other words, VQGAN is a generative adversarial neural network that's good at generating images that look similar to others but not from a prompt, and CLIP is another neural network that is able to determine how well a caption or prompt matches the image. So it's almost like having, you know, your generative and discriminator. It's it's like having two different neural networks that are doing the same
Ben Clip is is determining how well the caption matches the image that v q GAN created. So let's look at how that actually works. And CLIP doesn't have to be combined with VQGAN. Open AI uses CLIP, in their Dolly models. So here we have a text prompt, a beautiful painting of a dog riding a dolphin. And the first time, this is step one, v q glant GAN generates this sort of noisy image, and clip gives it a score of 5.3. It pretty much says this is this is not a dog riding a dolphin. Maybe it's a painting because there's some multiple colors on it. Here, we're at step 50, and clip says, this is 40.7% a dog riding a dolphin. There's a dolphin there. There's what looks like a dog. There's some painting artifacts here, but you're still not passing.
You're, you know, two fifths of a of a painting of a dog riding a dolphin. Now here we're at step 500, and clip gives us a score of 70.7, so just over two thirds. So it's still not agreeing that this is really a beautiful painting of a dog riding a dolphin, but we're at step 500. Maybe we're we're at our maximum number of steps, and so this is where it stops. But this is how clip works with the image generator to sort of grade it. Here are some VQ GAN plus clip generations. One of the things to note about these is they lack the I would say they lack the overall coherency of the newer diffusion models, and and I would say that's generally true of VQGAN plus clip. It creates some very interesting artistic things, but you're not gonna mistake it for a photograph, and there's always weird sort of anomalies going on with this particular model.
So how is VQGAN plus CLIP different from, like, style GAN that we just looked at? So VQGAN plus CLIP can generate images from open domain text prompts without any prior training, while StyleGAN requires pre training on a specific domain of images such as faces, cars, birds, flowers. So it's more versatile. VQGAN, it uses CLIP as guidance to match the image and the text, while StyleGAN instead uses a mapping network to map points in latent space to intermediate latent space that controls the style of the image. So they literally just work differently. Additionally, VQ plus clip can generate images of a variable size, while style gain can only generate images of a fixed size that depend on the pretrained model. And VQ plus clip can also edit existing images by adding or removing text prompts, while StyleGAN can only synthesize new images in a new style.
So so VQ can edit, StyleGAN can synthesize and apply a new style, but it it's not really editing in the same way. Now here's where I said you can also add CLIP to other GANs. So in 2021, a paper came out, StyleClip, text driven manipulation of StyleGAN imagery, and it presented a method to manipulate images using a driving text. Their method uses the generative power of a pretrained StyleGAN generator along with the visual language power of clip. And here we can see what they mean by that. So in the top left, there's a picture of Barack Obama, and then in the bottom, StyleGAN has applied a mohawk hairstyle to Barack Obama. We have a picture of a cute cat in the top I mean, a cat in the top, and in the bottom, we have a cuter cat, more like a Puss in Boots cat with the big eyes.
So or even more impressive is it took, I think, this picture of a tiger and turned it into a lion. So pretty impressive what you can do with a GAN when you add this clip to it. Now one of the other interesting things, if we go to the references let's see. So using stat GAN, style GAN, BQ GAN, like I said, is it plus clip, not necessarily something you're just gonna spin up and do right this second. However, I think let me see. Nope. I do not have a demo for, you know, StyleGAN plus CLIP. I was hoping I did. I do not think I do. But
Speaker 0 Ben, There's a question that just popped up in the question tab by Matt, and she's asking, have you explained about the job ID slash seat in mid journey before? If yes, can you quickly let me know in which session it was, and also if you have experience using it yourself?
Ben So I have played around with the job ID seed in Midjourney. I'm trying to remember. It was either in lesson three or four that I touched on. I think it was in lesson three. Let me see if I can find it real quick, man. But I think it was lesson I'm almost certain it was lesson three. And, yes, I do talk about it. You have to go into the you have to go into the web interface to find the job ID seed in Midjourney. You can't find it in Discord. And, yeah, it looks like it was lesson three. So it is addressed there.
Ben Absolutely. Alright. Let's see. Okay. Face restoration. So we've kinda gone through general GANs that are used for image generation, things like that. So let's look at, you know, actual sort of practical uses of GANs today. Because, realistically, if you're gonna generate images, unless you're looking, you know, for this very artistic, strange, sort of otherworldly look, you're not gonna use a GAN. You're probably gonna use one of the newer diffusion models. However, there are still modern uses of GANs, and one of them is face restoration. They're really good at this. I've used it in the past. You know, when mid journey or stable diffusion put out a photo that I I really liked, but maybe the face was a little bit weird. You can actually run the face through a GAN, which remember GANs have constraints. They're trained on a very specific set of data.
So there are multiple GANs. We're gonna look at three real quickly that are specifically made, to make faces. So what are they? GFP GAN came out in 2021. This is probably the most popular facial restoration GAN that's out there. It used StyleGAN version two, and it's often bundled with stable diffusion installs. There is a newer model called Codeformer, came out last year based on basic SR, which is, again, based image restoration toolbox along with three other facial data models. It's a little more advanced. I think it gives a little bit better results. And as a quick note, basic SR, ESR GAN, SFT GAN, we're not gonna deal with those today because they're really the building blocks for sort of more user facing GAN models, like real ESR GAN or GFP GAN. So if you wanna know more about those, shoot me a message.
We can dive into some of the technicals of how they work. But I tried to focus on once we kinda went over the general stuff, now we're gonna focus on what you can actually use. And then most recently, there's g pen. It came out in 2021, but they've been heavily updating it. They've already updated it this year. So this is a GAN prior embedded network for blind face restoration in the wild. Has some very promising results. I would say it may even, in some instances, do better than CodeFormer, but not as well known. Looks like it's mostly Chinese language, but also some really cool things. So let's look at these practically real quick and look at how they work. And then there are demos for all of these and the references where you can go try these right now while we're while we're talking. So GFPGA.
Here's some examples. On the left is the input. Here are some older models from and then on the right is the GFPGA model. A couple things to note. It can colorize an image. So if you look at the top row, the little girl is colorized, and it also does a better job of retaining sort of the original input. So if you look, for instance, at the older woman in the middle row, some of these GAN models have turned her into a bald man or on the bottom row, it de aged this fellow and took away his mustache. So GFP GAN does a pretty good job. Here, I actually used one of the demos, and we took this older photo of a blanket on the left and ran it through facial restoration, g f p gan. On the right is our output.
I upscaled it or blew it up a little bit so you could see it a little bit better. And you can see that Abe Lincoln's face looks a little less waxy, a little less strange. Maybe it took away some of the dust marks in the background. We'll talk about how it did that in a minute. It's actually using ESR GAN to do some of the upscaling and image restoration on the background, and it's using GFP GAN to do the face on the foreground. So pretty pretty impressive results. We're gonna look at some more models. So this is using using GFP
Speaker 2 GAN model number one.
Speaker 0 Okay. Sorry, everyone. It looks like we lost Ben for a moment. So yeah. If there's any questions before Ben is coming back, please let us know in the question tab so Ben can answer when he's back. And if anybody wants to come on stage with me, maybe to share how you used Sorry.
Ben They're doing work right by our house. So No worries, man.
Speaker 0 Thought that Glad glad you're back. I was struggling a bit. I'm like, how how am I gonna take over? I have no idea how to follow-up on this.
Ben This is the 1.4 model. If you if you look, here's some comparisons amongst different model versions. So it's not necessarily that the newer models are always better. Sometimes they chose to, you know, focus on different things. So the 1.3 model, for instance, is not as sharp as the 1.2 model, but sometimes the 1.2 model creates some more unnatural looks. And so if you look here in the picture, the version 1.2 makes her face much sharper, unnaturally sharp. Version three looks better, but it's a little bit soft. So something to be aware of, and this is true like we talked about with mid journey stable diffusion. You know, sometimes you have to choose the right model for, you know, the actual image that you're working on. So something to be aware of. I would say one of the cool things about these demos is you can definitely try this out.
But if you were gonna go try and use this on 50 photos, the demo is not gonna do it. You really need to, you know they really you really need to install it locally so that you could run this sort of en masse. Alright. So let's look at here's Codeformer. And like I said, I I I think Codeformer is a little more impressive than GFPGAN. GFPGAN is sort of the default that's generally bundled with stable diffusion, but you can also install Codeformer. Here's an example from their website, showing an old photo, upscaled, the face is restored. Here's an even more impressive photo from their GitHub repo that really takes a very, very pixelated low resolution photo and upscales it really impressively. You know, you would have a hard time telling that this had been restored using a GAN if you didn't know any better.
Speaker 1 Here's a Codeformer colorizing
Ben and upscaling a face, also very impressive. Now one thing to note is if you look at the hair on the right side of this photo, the GAN definitely made something new with it. While the hair looks a bit messier in the original black and white photo, it made up some details. So, you know, it is interpolating data that's not there. And so you know? And I've seen this myself when I take a photo of myself that's incredibly pixelated and use one of these GANs on it, And then my wife looks at the photo and says, that doesn't look like you. Well, that's true. It it it is making up data, essentially. Here, we also have essentially face in painting. So, you know, all this data is not there. It's erased, and then it made the face.
Now it made a convincing looking face, but, again, that may not actually do what that person looks like or what that person's teeth look like or nose looks like. So just something to be aware of, but still incredibly impressive. Here's g pen. There's also a online demo for this. You can take an extremely pixelated photo, like, on the left, and it both mildly colorized it and, you know, created an upscaled version of this face. Here's some other examples from g pen. Once again, really impressive what they can do with an incredible I mean, you look at the bottom left photo, it's hard to tell that's a face. Now is that what that person looks like? Who knows? But it did do a good job of of restoring face. Colorization, g pen's also pretty good at facial colorization. I don't know how it is at general colorization.
Remember, these scans are specifically trained on faces. So if you, you know, try and colorize a car, it may not give you good results. But also very impressive face colorization, face in painting. Lots of demos available for this as well. Now this is where g pen, I think, is pretty cool. So you can take a segment to face. So kind of what we looked at with ControlNet, you could do that with a face. You could have, you know, just a very segmented, you know, sort of sketched photo, and it created these faces off of that, which is really mind blowing. You know, what we were looking at couple weeks ago with ControlNet and diffusion models, you can do with this very specific GAN model here. So really cool some of the new things they're coming out with and working with on this particular GAN.
And, again, you know, one of the cool things about this is it's pretty lightweight in the sense that all it does is faces. You can go hop on a web demo, upload something, and in ten, twenty seconds, have results. You don't have to go install stable diffusion and control net and then wait several minutes. So pretty neat. Alright. So colorization. One of the things we noticed with GFP GAN, Codeformer is they're also and then GPEN is they're able to colorize photos. I also wanna look at a couple of other pieces of software, one that's a GAN, one that's proprietary, whose entire job is to colorize photos. Sometimes, you know, purpose built tools can do this even better. Like, we talked about GFP GAN, Codeformer, gPEN specifically made for faces. These other colorization models are made for, you know, a little broader dataset.
So the first one we'll look at and you can tell, you know, those were made in 2020, 2021, '20. These are 2022 last year. So DeOldify is the first one we'll look at. It uses a new type of GAN model that the author calls NoGAN. I tried to dive into it and understand it a little bit. I stopped at a certain point, but here are a couple examples. So here's a famous photo from the nineteen thirties during the Great Depression. The oldifies colorized it on the right, and I think it it looks pretty realistic. It's pretty convincing. And this used to take someone doing this manually, here and you can run it through this program. There's also some demos available for this, and it looks really convincing. Here's the building of the Golden Gate Bridge around 1937. Now some Thing to take note of in this photo is the water looks really good.
The sky looks really good. But if you notice the bridge is white, the author of this particular program, this GAN based program, went back and and looked at the historical record and determined that the bridge was probably red in 1937. It had already been painted with its red primer, so it was probably not white. So once again, something to notice is when we're saying colorization or restoration, these programs are guessing, and they're making their best guess, and sometimes they're making stuff up, but still very impressive. Next, we're gonna look at palette. So this is a commercial closed source software. A Google engineer went off on his own, created this. There is a three tier, but it's also commercial software used by Netflix and a bunch of other companies. And I would say that it is he calls it second generation colorization.
I don't know what machine learning, you know, sort of algorithm it uses. But like I said, we're here to talk about GANs, but we're also here to talk about actual useful tools. This is a very useful tool. You can go on their website, palette.fm. You can try it and use it right now. And I would say it does give you sort slightly more lifelike colors. If you look at the example on the right, it is more lifelike than, for instance, what we have here from the oldify. It looks more like a newer photo. This is another example from palette.fm, Audrey Hepburn. This is shocking. You you couldn't tell this was originally a black and white photo. This is so good. This looks like a period photo, but it looks like it was originally a color photo. So really impressive. You know?
Once again, if you wanna go upload a photo and try it out and hop on stage and show us, you know, go for it. It'd be it'd be fun to see. Alright. So what's another use? So we talked about face restoration. Hold on, Ned. It is there you go. Alright. So we talked about face restoration. We talked about colorization. Now we're gonna talk a little bit about upscaling and overall image restoration. And, you know, these are very similar but a little bit different. And here's on the right a post from ClownVamp, a famous well known AI artist talking about how their new piece is 10,000 pixels by 7,000 pixels. And, obviously, AI generators don't put out at that resolution. So they have a upscaling pipeline that combines gigapixel, ultimate SD upscale, and then they've hand added some texture overlays. So let's go through quickly how to upscale and restore anything, really.
It could be a photo, but in this case, we're focused primarily on AI output. So real ESR GAN, which came out last year, probably the most popular upscaling algorithm GAN, often bundled with stable diffusion installs. It's used actually by both GFP GAN and Codeformer to do the background upscaling. There's also a custom version of real ESR GAN, a super resolution model that comes with the Russian Dolly, RU Dolly, and I've included some links to that. If you wanna play around with that, we'll look at that quickly. There is also SwinIR, which uses a type of transformer model, computer vision transformer model to upscale. And then there are some diffusion upscalers, and then there's a piece of commercial upscaling software that we'll look at in just a minute that ClownVamp mentioned, this gigapixel. So let's dive into these real quickly. Alright. So real ESR GAN.
Here's some examples. I love the probably the two on the bottom left to the Cheesecake Factory sign and the tree. You can really see how how clear and how good of a job this GAN does. And remember, this GAN is trained. There's some other back end, you know, older GANs behind it, whether that's basic SR, ESR GAN itself, SFT GAN, which was originally made to recover textures from photos. So there's a lot going on when you use this, but I think it does a really impressive job, especially as sort of the default upscaling algorithm out there. And there's some demos of this. You can go on the references, and there's at least two or three different demos you can try this out with. Alright. So here is the version of real ESR again customized for r u Dolly, the Russian Dolly, which is not a copy of Dolly, but it's based on some of the same ideas.
Here, I actually they also have a demo you can hop on and use. Here, I actually used their 2.1 model and asked it to draw a line looking out over the Savannah. Impressive. Did not look into the technology behind it, but I did look into this real ESR GAN super resolution model, and it's really impressive because while the normal ESR GAN does two or four times resolution, this one had no problem with an eight times scale. So we took this little bitty photo of a computer chip and upscaled it eight times, and it did a pretty good job at it. I mean, there are some a couple odd little artifacts in there, but really impressive and, in some ways, more impressive than the standard real ESR GAN implementation. There's a demo for that, so you can hop in and use that.
Then I also looked at swim I r, which uses a transformer model. And here you can see it compared to ESR GAN, and it really does a great job on the textures and the details. So if you look on the left, we have this picture of this, you know, old classical, neoclassical building, and you can see that the brick looks pretty pixelated. Real ASR GAN does a pretty good job, but the SWIMIR does an even better job at retaining the brick detail on the texture. Here is an example of that. I tested it. I used it four x resolution. Did a great job on this computer chip. Did a little different job than the Russian super resolution. I would say it retained the text a little bit better, maybe the metallicness and the shadows on the pens, but also an impressive upscaling algorithm. Alright.
So so those are the well, if we looked at two GAN models, then we looked at this transformer model. Now we're gonna look at some diffusion models for upscaling. So here's sort of the vanilla latent diffusion super resolution model. There's a demo for this. That's what this comes from. And what they've noticed is the results so far showed to be good with textures and fine details on generated images, but poor on real world jobs. So it does a good job on these textures, but if you throw it a real world photo, it's not trained like the GANs are for real world photos, and so you'll get some very strange artifacts and things. Demo's included for that, but it's good to understand what it is, how it works. Here we have the stable diffusion two and two point one, two x, and four x upscaler.
Something to notice here, and and I think this is, you know, a little bit what we were just looking at, is it almost takes this photo and makes it look like a generated image. You can see some weird patterns going on in the background, a little bit more fuzziness in the text, but does a good job, and this is the standard upscaler that's built into stable diffusion. Now the artist ClownVamp mentioned this ultimate upscale for stable diffusion that they like to use, and this is what it is. Not something you can demo online, but something that you can install with automatic one one one one stable diffusion. And I think the most interesting thing about this is it does use a custom version of the ASR GAN, real ASR GAN with a whole bunch of options.
And, essentially, what it does is it takes the upscaler, in this case, say, a really a star again. It divides the image up into a bunch of little tiles and then upscales and then paints everything together and paints the scenes. And the results end up looking like this. On the left, we have a original image of five twelve by seven sixty eight generated image, and on the right, this has been four x upscaled with the seams fixed. And they also said how much to denoise it, how much to pad and blur in between the seams. There's a half tile offset path. Really impressive results with this, but also really technical in terms of all the options, how you have to install it, how you have to use it. Let's see. And then we're gonna look at a piece of commercial software.
So this is Topaz Gigapixel AI, probably the most common piece of commercial software that's currently used upscaling AIR. Not cheap. I think it's a $100 for a license, but it does a really good job. It can easily upscale four to six times, has a whole lot of different algorithms and options, closed source, so I don't know exactly how it works. I know it uses, you know, some sort of machine learning, interpolative algorithm to do what it does. Here, we can see on the left the amount of detail it recovered on these buildings. It's got a really good line model. And like I said, it's commercial software used by a lot of commercial firms. If you're gonna be doing this professionally, as much as GANs are really interesting and diffusion models, this is probably the most efficient thing to do from a cost benefit standpoint with your time.
You can go through 30 photos into this, and it will automatically pick which algorithm to use to clean up the artifacts and and upscale. And then you can let it go run-in the background. So pretty cool. Alright. So let's look at style transfer. So these are modern GANs. And if you notice, these three are from 2021. I didn't find a whole lot more modern than that because you can do this so easily with diffusion models nowadays. So wall face restoration, image upscaling, colorization, GANs are really good at these. They're also really good at this, but so is diffusion models, and they're a lot easier to work with. But we're gonna look at three examples real quickly just to understand what can be done. So Jojo GAN took the Mona Lisa. It applied this Jinx style to it. Here's some other examples. Pretty impressive.
It literally just takes a photograph, generally a face. If I remember right, these are all trained on faces, and it can apply a particular style to that face. Anime GAN, same thing. Here we've got Will Smith. Use version two, and it made an anime version of Will Smith. Here it made a creepy young version of Bill Gates. Here's ArcaneGAN. It's making it it once again, also does faces in an Arcane style. Here's another Bill Gates. You can see how the GANs are different in their outputs, and here are a few more outputs from that arcane GAN. And all three of these GANs also have online demos. They're in the references so you can play around with them, upload a photo of yourself, see what it does with it. So what's next for GANs? So we've looked at what are they good at.
They're good at facial recognition, image upscaling, colorization, maybe stylization. What's going on now? 2023. Right? So paper just came out last month, scaling up GANs for text to image synthesis. This came out in Cornell University. And like I said, this was just I don't see the date. March March ninth, so less than eight weeks ago. And here's the actual paper. This is also in the references. You can go in and read this paper. But this is what's interesting. Can GANs be trained on a large dataset for general text image synthesis? So they present this model that does it way, way faster than stable diffusion, Dolly. It generates 512 pixel outputs at point one three seconds. So orders of magnitude faster than diffusion models do. And it can also upscale really, really effectively. So this gig again, they they they're training it on upscaling photos.
So in this case, we have a photo that's only a 128 pixels upscaled to 4,000 pixels, and it does an amazing job. And it did it in three point six seconds. So GigaGAN, like I said, there's no demos. It just can't you know, the paper the research paper just came out last month, but there's gonna be some really cool things coming down the line where, you know, it looks like GANs, you know, still have their place to play when it comes to AI art. So look out for GigaGAN. I'll let you all know if I see any news about it. One more thing. So that's what I've got on GANs. If anybody has any questions or wants to say anything about GANs. Yes, Ned. Let me post it in chat real quick, and I will also post the link to all the references while I'm thinking about it.
So there's all the references, and let me find the one for gig again real quick and copy it. Alright. There is the paper for gig again. Alright. So one more thing. Unlike Steve Jobs, I'm not gonna unveil an iPhone or anything, but I am gonna unveil three more cool new things. So new models. So these are all diffusion models, so we're kinda done with the GAN portion. But these are so brand new that one of them, the last one, Deep Floyd I f, came out less than twenty four hours ago, and the other two both came out this year. And there's demos of two of these three, you can you can play around with these, and and they're pretty cool. So the first one is GlideGen, open set grounded text to image generation. And if you look, you can see there's these funny boxes, then there's some text, and there's images.
So let's dive in and see what it is and how it works real quickly. Really, really cool. So here's the caption. A teddy bear sitting next to a parrot, and here we have the grounded text is what these boxes are. So teddy bear is grounded to this box on the left in red, and parrot is grounded to this box on the right. And so if you look at the photo that's generated, the teddy bear is where the teddy bear box is, the parrot's where the parrot box is. Here's where it gets neat. Here's the exact same prompt but with different grounding boxes. So we see that the parrot moved from sort of at the teddy bear shoulder down to, you know, mid level with the teddy bear. And now we see something else interesting as we move to the next photo. So take a look.
The parrot box is wider than it is tall, and the parrot sort of got this wide stance in these photos. We go to the next one, and the parrot box has been squeezed in, and now the parrot's feathers aren't sticking out. Same text prompt, but these grounded boxes allow us to say where we want things in the image, and it gets even crazier. So we have a text prompt, a painting of a fox sitting in a field at sunrise in the style of Claude Monet. And here we see the you know, in the first set, the fox is on the bottom, the sunrise is on the top. On the second set, the sunrise is on the left, the fox is on the right. Compared with existing models such as DALL E, Glidegen enables the new capability to allow rounding instruction.
And we can see what Dolly one and two made, and you didn't have any control really over where the sun and the fox were. Here, you do. You can also and this is something that if you've tried to generate a lot of photos, you have trouble with counterfactual generation. And this is what I mean, and this is what we talked about a little bit earlier with mid journeys working on its language understanding for its model seven, and people are gonna be impressed. I think it's gonna be something along these lines. So by explicitly specifying object size and location, GlideGen can generate spatially counterfactual results, which are difficult to release through existing models. So we have a hen that's hatching a huge egg. Normally, stable diffusion is gonna make a hen sitting on a bunch of eggs or it's very hard to make a huge egg.
However, here, we've drawn the egg much larger than the hen, and that's exactly what we get. Or we have an apple and a same size dog, and we literally have a giant apple next to a dog where stable diffusion just makes a dog with a bunch of apples. It also allows us to do inpainting just because of this grounding model. So you it doesn't, you know, require a special inpainting feature. It can sort of automatically do it because you're telling it this is where this object belongs exactly, and it does a great job at coherently inpainting images. So in this case, it turned the flower into a lamp and did it very well. It also allows you to use mappings like candy maps and HDD maps, and, essentially, it allows you to to do control net, it's built into it.
So in this case, we have a chair and a table in different styles. So I think this GlideGen model is gonna be pretty revolutionary. Once again, there's a demo online. You can hop on in the references, and you can try it out right now. This was the one I tried this morning, a dog and an apple. And if you look, I made the apple, like, twice the size of the dog, and it made an apple twice the size of the dog. I didn't give it any style commands or other things. I was just playing around with this. But play around with this. It's pretty cool and has a ton of potential. The next thing we're gonna look at is this instruct picks to picks. So we were just looking at style GAN control net. We're talking about changing styles.
Here's a paper that just came out this year also that specifically deals with helping diffusion models follow image editing instructions better. So for instance, add fireworks to the sky or in Magritte's famous photo, make his jacket out of letter leather on the bottom right. And let's look at that a little bit. So this is a method for editing images from human instructions. And the way they do this is they take a language model, like a g p t three and LLM, and they take a text to image model, like Stable diffusion. And then they combine the two, essentially, short form. So here on the left, we have this beautiful picture of a mountain lake. They said add boats to the water, and in the middle, it adds boats. And here's what's cool. Note that the isolated changes also bring along accompanying contextual effects.
So the addition of boats also added wind ripples in the water. And when we asked for replacing the mountains with a city skyline, the city skyline's reflected in the lake. So this is sort of mind blowing how good this is. Here we have Vermeer's girl with a pearl earring, and I I I think just the cohesiveness of these photos is what's incredible. So we have our input, and then we have face paint. What would she look like as a bearded man or with a pair of sunglasses? Or turn her into Dwayne The Rock Johnson on the bottom right? And I would say this does a better job than existing diffusion models even within painting, even probably than ControlNet, we saw last week, the results can be a little bit unpredictable. Sometimes it sort of does what it wants. This is incredibly cohesive with the outputs, and there's also a demo of this.
So I only I only gave it one shot. I didn't take a lot of time, but I took our picture of Abe Lincoln from earlier, put it in here, I said wearing a clown hat. Now back to what we learned a minute ago, it didn't do a great job with the clown hat. But as it said here, sometimes it also brings along accompanying contextual effects. It figured if a Blankton's wearing a clown hat, he's a clown, and so it gave him clown face paint. And it did a really good job. It even made him smile. If you look compared to the original photo, you know, the model probably thinks clown smile, so his eyes are open a little wider. His mouth's smiling. So really, really impressive. Like I said, there's a demo. Play around with it. Let me know what you think.
But I think you're gonna start seeing GlideGen and this instruct picks to picks, you know, are really gonna make some revolutionary changes in these diffusion models. And then the last thing that came out, and this came out in the last twenty four hours, is Deep Floyd I f by this Deep Floyd team combined with Stability AI, which if you'll remember, is the team behind stable diffusion. So what's cool about this deep Floyd? Where did it come from? What's neat about it? Couple things. So if you look, you will notice, a, the level of photo realism in these photos. Very realistic. And the other thing you will notice is that the text is not gibberish. It's readable English text. This is not DALL What if it is more than text? So how did they do this? What is this? So here's some of the background.
You can look at it, read it if you want. I'll give you kind of the where it came from TLDR version. So if you remember back in, I think it was week one, we talked about Google's Imagen model, and that it wasn't released and you couldn't access it, and Google had concerns that this model was so good that its images could could confound most discriminators. They couldn't tell that they were AI generated, and they released a paper on it. So Stability AI said we're going to reverse engineer Google's paper, and that's exactly what they did. They said we're gonna have a language understanding that can create text. We're gonna have a level of photorealism, and we're gonna put a team on this, and they did. And they came out with this model yesterday. You can run the model. It'll fit in the free tier of Google Colab.
It requires about 16 to 24 gigs of v RAM, so it's not as lightweight as stable diffusion, but it's lightweight enough that it can be run on consumer hardware. They haven't released the research paper yet. They're promising that it'll come out soon, but I think this is gonna be a game changer. So let's look at what it can do. So first of all, it can dream. It's the text to image model. Here we have a prompt, ultra close-up color photo portrait of a rainbow owl with deer horns in the woods, and we have an incredibly realistic cohesive photo of a rainbow colored owl with deer horns. It can also do image to image translation, style transfer, kind of control net, you know, some of the things we were just looking at with the imagine pics to pics. So here's one, and I just took a couple of the examples.
Let's see. But it was a yeah. In the style of a classic anime from 1990 and in the style of Legos. And, you know, amazing what it can do. We also have its ability to upscale and restore photos is very impressive because of the way the model works. Here's a quick demo of that. And then it also does zero shot in painting really well. So kind of what we're just talking about with GlideGen, this also does but in a simpler, more standard text to image way. And so here we have oil art, a man in a hat, and then without the hat. And that's it. That's our news for today. I will, as usual, ask if anybody has any questions, and I will also normally, I have quotes here, but instead of quotes, I'm gonna give you a piece of news about something I'm working on that's kinda neat in the AI world.
So if you look at the top left and the top right, these are both generative art. A friend of mine's an artist. He's been working in MATLAB and p five JS and generating these circular, you know, patterns procedurally randomly. I'm then taking these as seeds and inputting them into diffusion models to generate outputs. And if you look, the one on the bottom right is generated from the one on the top right, and the other three are generated from the one on the top left. So we're playing with, you know, what where is the intersection between AI generated art, procedurally generated art, and then the interaction of the artist with that. So something something kinda fun. I think you'll see more projects like this down the road as these things sort of merge. And then where do we go from here?
To be continued, like I said, there is a huge amount of resources this week. I'm eager to see what y'all do with some of the demos in here. Pop on the Discord, show us what y'all make, and how you're experimenting with these things.
Ben really cool stuff, going on both in the field of GANs and coming up just in the field of generated AI art in general. Look forward to seeing y'all sometime in the next few months when we kick off some more lessons. And as Wout and I were talking about beforehand, maybe we'll even, in the meantime, have a session or two where we, yeah, do some do some workshops and just create some art collaboratively and learn together.
Speaker 0 That would be amazing, Ben. Thank thank you again for for such a great session, and I think all of your sessions are are really amazing. I re really recommend anybody watching on demand to go back to the previous sessions. If you wanna do anything with AI art, this has been an amazing series. I learned a lot myself, and I'm experimenting a lot with AI art. As are the people in the audience, I see I also learned a little bit from Ned the other day, and I think we can all learn together by doing an on hands on session. So that will be amazing. And, yeah, I really wanna thank you, Ben. I see one more question pop up by Nat, and she asked, in the future, can we have a session on CPP NS?
Ben So CPP NNs are compositional pattern producing networks. They're variation of neural networks. I do not know a whole lot about them. Ned, you may know more than I do. But, yeah, we could certainly learn about it, experiment with it, and do a session on it. Why not?
Speaker 0 Incredible. Like, yeah, let's learn together. There's, there's everyday new things coming out, so I'm also looking forward in, two months to see what new features are, and apps have been coming up and to see how we can use those. So definitely, let's stay in touch, Ben, and I'm looking forward to have you back on again. Thank you everyone for joining in the audience, and talk soon. See you in the Discord. I'll post by the way, I'll post the link that Ben has up here in the resources in the resource channel. Thank you so much. Have a wonderful night, everyone, and we see you at the next one. Bye, Ben. Thank you.