StyleGAN 3

StyleGAN 3

November 29, 2021• 11 min read

Key points

The StyleGAN neural network architecture has long been considered the cutting edge in terms of artificial image generation, in particular for generating photo-realistic images of faces. Now researchers from NVIDIA and Aalto University have released the latest upgrade, StyleGAN 3, removing a major flaw of current generative models and opening up new possibilities for their use in video and animation.

The "texture sticking" problem

For those who have spent much time looking at video generated using previous version of StyleGAN, one of the most unnatural visual flaws was the way that texture like hair or wrinkles would appear stuck to the screen, and not move naturally with the rest of the object. See this video comparing StyleGAN2 and StyleGAN3, notice how beards and hair in particular seem to be stuck to the screen rather than the face.

Karras et al. suggest that the root cause of this problem is the StyleGAN Generator making use of unintended positional information which is present. It then uses these as a basis for generating textures, rather than the objects to which they should be "attached". By removing this positional information they hoped that textures would move properly with the generated object rather than with the pixel locations.

To eliminate all sources of positional information requires a thorough analysis of all parts of StyleGAN's architecture. Technically this ensures that the network is translationally equivariant, even for sub-pixel translations[^1]. Once all this unwanted positional information has been eliminated the network can no-longer make use of the pixel grid as a reference system, so must create its own based on the positions of generated objects in the scene. This means that textures move naturally during videos and gives a quite striking effect.

Making it equivariant (the alias-free GAN)

Architecture

Although the overall architecture stays broadly the same, StyleGAN 3 introduces a number of changes to the building blocks, all motivated by a desire to remove unwanted positional information:

Datasets

StyleGAN has shown famously good results on the FFHQ dataset of people's faces. Although being challenging for realism, as we humans are so well tuned to recognise a face (and any poor results end up deep in the uncanny valley), images in this dataset are perfectly aligned to ensure that facial features are all located at the same points in the image. To better show off the new capabilities of StyleGAN3 the researchers introduce a new version of this dataset FFHQ-Unaligned, which lets the orientation and position of the faces vary.

Training

After applying all these changes, training is conducted in largely the same manner as StyleGAN 2, resulting in a network which is slightly computationally more expensive to run. One significant modification is the removal of Perceptual Path Length (PPL) regularisation as this actually penalises translational equivariance which is what they want to achieve.

Results

So after going to all that trouble to remove any way the network can use absolution positional information in generating images, what do we get out? Well as you can see from the videos, something that looks strikingly more natural in motion than previous models. But of course there is some price to pay, and when it comes to absolute image quality (at least as measured by the standard FID metric) the new StyleGAN3 architecture can't beat that of StyleGAN2 for the original FFHQ (faces) dataset. it does however perform slightly better for data with less strict alignment of FFHQ-U.

The real benefits however, are in motion. Texture sticking has been one of the most noticeable artifacts in GAN generated videos, so finding a way to overcome this with only a marginal cost to overall quality is a big step forward. New measures of these metrics in the paper show that StyleGAN3 significantly outperforms StyleGAN2 in this regard.

Training your own

So what do you need to train your own StyleGAN 3 model? Well the code has been released on GitHub along with many pre-trained models, of course on top of that you'll need some data and GPUs to run the training.

As the author's point out, the new StyleGAN is slightly more computationally expensive to run given the architecture changes above. They also provide a handy breakdown of the expected memory requirements and example training times on V100 and A100 GPUs at various configurations.

To train the largest models (1024x1024 pixels) from scratch (25 million images) will take about 6 days on 8x A100 GPUs, but in general you won't need to go these efforts. As in many other areas of deep learning, transfer learning is an effective approach for training a new StyleGAN model. Using a high quality starting point (like one of the existing FFHQ models) you can get to reasonable quality results within a few hundred thousand images.

What next?

Although StyleGAN3 doesn't quite match previous efforts in terms of absolute image quality in some cases, there is an interesting hint at future directions. As previously noted by others[^6] scaling up StyleGAN by increasing the number of channels can dramatically improve its generative abilities, the StyleGAN3 paper also shows this improvement for a smaller (256x256) model. Interestingly, there has been little research in simply "scaling up" StyleGAN in terms of numbers of layers as is common with other neural network architectures and it remains to be seen if this is feasible with GANs which are notoriously tricky to train.