Field Notes
Ideas & Learning

VAEs and GANs

How variational autoencoders and adversarial networks generate data.

9 min read

Variational Autoencoder (VAE) is a deep generative model proposed by Diederik P. Kingma and Max Welling in 2014¹. It combines the autoencoder (AE) with variational Bayesian methods to learn the latent distribution of data and generate new images.

Let me explain how a VAE works and what it's for:

  1. The Autoencoder (AE) model:

    • An AE consists of two parts: an encoder and a decoder.
    • The encoder maps the input data to a code vector in the latent space, where each dimension represents a feature learned about the data.
    • The decoder takes the code vector and tries to reconstruct the original input.
  2. How a VAE works:

    • A VAE also has an encoder and a decoder, but unlike an AE, it uses probability distributions to describe the latent space.
    • The VAE's encoder maps the input data to a probability distribution over the latent space rather than to a single code vector.
    • The decoder samples from that distribution and decodes the sample into a generated image.
  3. Why use probability distributions?

    • In practice, we prefer to represent each latent feature as a range of possible values rather than a single value.
    • The VAE describes observations of a feature as a "probability distribution over values," so every latent feature is itself a probability distribution.
  4. VAE architecture:

    • A VAE uses two neural networks to build the probabilistic density model:
      • Inference Network: performs variational inference on the raw input data, producing a variational probability distribution over the latent variables.
      • Generative Network: takes the variational distribution of the generated latent variables and reconstructs an approximate probability distribution of the original data.

In short, a VAE describes the latent space with probability distributions, which makes generated images richer and sharper. If you're interested in the mathematical derivations, see the references¹.²³⁴

Source: Conversation with Bing, 5/17/2024 (1) Understanding Variational Autoencoders (VAE) in One Article - Zhihu. https://zhuanlan.zhihu.com/p/64485020. (2) What is the VAE in Stable Diffusion? - Stable Diffusion CN. https://www.stablediffusion-cn.com/sd/sd-knowledge/1198.html. (3) Stable Diffusion | What exactly is a VAE? - Zhihu Column. https://zhuanlan.zhihu.com/p/631452035. (4) Variational autoencoder - Wikipedia, the free encyclopedia. https://zh.wikipedia.org/wiki/变分自编码器. (5) en.wikipedia.org. https://en.wikipedia.org/wiki/Variational_autoencoder.

Generative Adversarial Network (GAN) is a class of neural networks made up of two key components: a generator and a discriminator. The two networks compete against each other in adversarial training, sampling from complex probability distributions to produce things such as images, text, and speech.

Let me explain how a GAN works and what it's for:

  1. Generator:

    • The generator takes a random sample from the latent space as input.
    • Its goal is to produce data — images, text, and so on — that looks as close as possible to real samples from the training set.
  2. Discriminator:

    • The discriminator receives either real samples or fake data produced by the generator.
    • Its job is to tell the generator's output apart from real samples.
  3. Adversarial training:

    • Generator and discriminator are trained alternately, playing against each other.
    • The generator tries to fool the discriminator into being unable to distinguish generated data from real data.
    • The discriminator in turn keeps improving its ability to separate real data from generated fakes.
  4. Applications:

    • GANs are widely used to generate images, video, natural language, music, and more.
    • For example, GANs can generate photorealistic images and video, or even produce a singing video from a single photo.

In short, a GAN is a powerful generative model that, through adversarial learning, produces data with highly realistic characteristics.²⁴⁵

Source: Conversation with Bing, 5/17/2024 (1) What is a GAN, Generative Adversarial Network - AI Encyclopedia. https://ai-bot.cn/what-is-gan-generative-adversarial-network/. (2) Generative adversarial network - Wikipedia, the free encyclopedia. https://zh.wikipedia.org/zh-hans/生成对抗网络. (3) Understanding Generative Adversarial Networks in Four Days (I) — A Plain-Language Guide to the Classic GAN - Zhihu. https://zhuanlan.zhihu.com/p/307527293. (4) A Plain-Language Guide to Generative Adversarial Networks (GAN) - Zhihu. https://zhuanlan.zhihu.com/p/33752313. (5) What is Gallium Nitride? A Detailed Look at GaN Semiconductors | EPC. https://epc-co.com/epc/cn/氮化镓/what-is-gan. (6) Generative Adversarial Network (GAN) - Zhihu. https://bing.com/search?q=生成式对抗网络(Generative+Adversarial+Network). (7) Generative Adversarial Network (GAN) - Zhihu. https://www.zhihu.com/topic/20070859/intro. (8) Understanding Generative Adversarial Networks (GANs) - Zhihu. https://zhuanlan.zhihu.com/p/97015788. (9) en.wikipedia.org. https://en.wikipedia.org/wiki/Generative_adversarial_network.

Convolutional Neural Networks (CNNs) are a class of deep learning algorithms used primarily in computer vision. They are applied across many domains, including image and video recognition, natural language processing, and even game playing. CNNs have revolutionized computer vision, delivering state-of-the-art performance on tasks such as object detection, image segmentation, and face recognition²³⁴.

Let me explain in more detail how CNNs work and what they're for:

  1. Definition:

    • A CNN is a deep learning model, or a multilayer perceptron similar to an artificial neural network.
    • It is specialized in processing data with a grid-like structure, such as images.
  2. How it works:

    • The main building block of a CNN is the convolutional layer, whose purpose is to learn spatially hierarchical features from the input data automatically and adaptively.
    • Convolutional layers extract image features through local connectivity, sliding windows, and convolution operations.
    • Nonlinear (activation) layers apply a nonlinear mapping to the convolutional layer's output; the most common activation function is ReLU.
    • Pooling layers reduce the amount of data and parameters, helping to reduce overfitting.
  3. Main applications:

    • Image recognition: CNNs perform excellently on image classification, object detection, and image segmentation.
    • Video analysis: used for object tracking, activity recognition, and more in video.
    • Natural language processing: for example text classification and sentiment analysis.

In short, a CNN is a powerful neural network specialized in images and other grid-like data, with broad applications and excellent results.²⁴⁵

Source: Conversation with Bing, 5/17/2024 (1) What is a Convolutional Neural Network (CNN)? Definition, How It Works, and Main Applications - AI Encyclopedia .... https://bing.com/search?q=CNN+是什么. (2) What is a Convolutional Neural Network (CNN)? Definition, How It Works, and Main Applications - AI Encyclopedia .... https://ai-bot.cn/what-is-convolutional-neural-network/. (3) Introduction to CNN (Convolutional Neural Networks) - Zhihu. https://zhuanlan.zhihu.com/p/67206089. (4) [Survey] A One-Stop Guide to Convolutional Neural Networks (CNN) - Zhihu. https://zhuanlan.zhihu.com/p/561991816. (5) Machine Learning Algorithms — Explaining the Principles of Convolutional Neural Networks (CNN) - Zhihu. https://zhuanlan.zhihu.com/p/156926543.

An Artificial Neural Network (ANN) is a mathematical or computational model that mimics the structure and function of biological neural networks, used to estimate or approximate functions. It abstracts the neuron network of the human brain from the perspective of information processing, building a simple model and composing different networks from different connection patterns. In engineering and academia it is also often referred to simply as a neural network or neural-net-like model²³.

Let me explain in more detail how an ANN works and what it's for:

  1. Inspired by biological neural networks:

    • The ANN takes its inspiration from biological neural networks, particularly the neurons in the brain.
    • The brain's neural network consists of roughly 100 billion neurons, each connected to others through synapses.
    • Electrical currents and chemical reactions between neurons allow the brain to process vast amounts of information in complex ways.
  2. Components of an artificial neural network:

    • An ANN is made up of artificial neurons, each receiving inputs from other neurons.
    • A neuron multiplies its inputs by the assigned weights, sums them, and passes the result on to one or more neurons.
    • Some neurons may apply an activation function before producing an output.
  3. Network structure:

    • An ANN typically consists of an input layer, hidden layers, and an output layer.
    • The input layer receives external data, the hidden layers process it, and the output layer produces the results of the network's computations.
  4. Training:

    • An ANN first assigns random values to the connection weights between neurons.
    • By "training" the network on labeled examples, the weights are adjusted so it can pick up specific patterns in the data.
    • Once trained, an ANN can perform complex tasks such as image classification or speech recognition.

Although ANNs play an important role in deep learning, they still have limitations, such as requiring large amounts of data and being opaque. Even so, they remain one of the key technologies changing how we interact with the world.²³⁴

Source: Conversation with Bing, 5/17/2024 (1) Artificial Neural Network – ANN - easyAI. https://easyai.tech/ai-definition/ann/. (2) Neural Network Algorithms: A Complete Overview of ANN, CNN, RNN, Attention, Encoder-Decoder .... https://zhuanlan.zhihu.com/p/689992042. (3) CNN vs. RNN vs. ANN — A Brief Look at Three Neural Networks in Deep Learning - Zhihu. https://zhuanlan.zhihu.com/p/107993566. (4) What is an Artificial Neural Network (ANN)? - Zhihu. https://zhuanlan.zhihu.com/p/111288383.

A Large Language Model (LLM) is an AI system that understands and generates human language by processing large amounts of text data, thereby performing a wide range of natural language processing tasks. These models are trained on massive text corpora and can carry out a broad set of tasks, including text summarization, translation, and sentiment analysis²³.

The characteristics of LLMs are as follows:

  • Enormous scale: LLMs contain billions of parameters, which helps them learn complex patterns in language data.
  • Built on deep learning architectures: typically based on deep learning architectures such as the Transformer, which helps them achieve impressive results on a variety of natural language processing (NLP) tasks.
  • In-context learning: LLMs can learn from context, generating coherent, natural text based on the input context.
  • Instruction following: by fine-tuning on multi-task datasets described in natural language, LLMs can perform new tasks by understanding instructions, without explicit examples, which improves generalization.
  • Step-by-step reasoning: LLMs can solve complex tasks involving multiple reasoning steps using chain-of-thought reasoning strategies.

Some well-known large language models include:

  • GPT-3 (OpenAI): 175 billion parameters, showing strong performance in text generation, translation, and other tasks.
  • BERT (Google): uses a bidirectional approach to capture context from both the left and right of a word, improving performance on tasks such as sentiment analysis and named entity recognition.
  • T5 (Google): casts every NLP task as a text-to-text problem, with strong performance.

In short, LLMs are one of the key technologies transforming the field of natural language processing; their capabilities keep evolving, giving us ever more powerful tools for working with text.²³

If you have any other questions, feel free to keep asking! 🙂

Source: Conversation with Bing, 5/17/2024 (1) What are Large Language Models (LLMs)? | IBM. https://www.ibm.com/cn-zh/topics/large-language-models. (2) Decoding Large Language Models in One Article - Tencent Cloud. https://cloud.tencent.com/developer/article/2357760. (3) What is an LLM (Large Language Model)? From Quantitative to Qualitative Change - Zhihu. https://zhuanlan.zhihu.com/p/622518771.

Natural Language Processing (NLP) is an interdisciplinary field spanning computer science, artificial intelligence, and linguistics. It focuses on giving computers the ability to understand, process, generate, and simulate human language, thereby enabling natural conversation with people³⁴.

The goal of NLP is to give machines human-like intelligence so they can:

  • Understand the meaning, context, and grammatical structure of text.
  • Analyze and process large volumes of natural language data, such as text, speech, and images.
  • Perform tasks like text classification, sentiment analysis, machine translation, question answering, and speech recognition.

Over the past few decades NLP has made enormous progress, moving from the earliest rule-based approaches to today's deep learning models such as BERT and GPT. These models let computers interact with people more naturally, changing the way we engage with technology and information.²³⁴

If you have any other questions, feel free to keep asking! 🙂

Source: Conversation with Bing, 5/17/2024 (1) What is Natural Language Processing? This One Article Is All You Need - Zhihu Column. https://zhuanlan.zhihu.com/p/634689142. (2) Natural Language Processing (language processing method) - Baidu Baike. https://baike.baidu.com/item/自然语言处理/365730. (3) What is Natural Language Processing (NLP)? | The Complete NLP Guide | Elastic. https://www.elastic.co/cn/what-is/natural-language-processing. (4) What's the simplest, most down-to-earth way to understand NLP? - Zhihu. https://www.zhihu.com/question/433756594. (5) Getty Images. https://www.gettyimages.com/detail/illustration/natural-language-processing-concept-business-royalty-free-illustration/1193264709.