Gumbel-Softmax: Approximating Discrete Sampling with Continuous Distributions to Solve Backpropagation Challenges
In stochastic neural networks, categorical variables are the natural choice for representing discrete structures.
작성 방식표준
In stochastic neural networks, categorical variables are the natural choice for representing discrete structures. However, because backpropagation cannot be performed through samples, such networks rarely use categorical latent variables. The Gumbel-Softmax estimator proposed by Jang et al. aims to address this obstacle in gradient computation.
The theoretical foundation of this method stems from the Gumbel-Max trick, which provides a simple and efficient way to draw samples from a categorical distribution with class probabilities. The authors introduce a new continuous distribution, the Gumbel-Softmax distribution, which is a continuous distribution on the simplex that can approximate categorical samples, and whose parameter gradients can be easily computed via the reparameterization trick.
A key characteristic of the Gumbel-Softmax distribution is its ability to smoothly anneal into a categorical distribution. As the softmax temperature τ approaches 0, samples from the Gumbel-Softmax distribution become one-hot encoded, and the distribution becomes identical to the categorical distribution. This interpolation capability allows models to flexibly adjust the degree of discreteness during training.
In practical applications, there is a trade-off: at low temperatures, samples are close to one-hot but have high gradient variance; at high temperatures, samples are smooth but have low gradient variance. Therefore, in practice, it is common to start at a high temperature and anneal down to a small non-zero temperature. If the temperature is treated as a learnable parameter rather than a fixed schedule, this can be interpreted as entropy regularization, allowing the distribution to adaptively adjust its "confidence".
For scenarios where discrete values must be sampled (such as action spaces in reinforcement learning), the paper proposes the Straight-Through (ST) Gumbel Estimator. It uses arg max to discretize y during the forward pass, but uses the continuous approximation to estimate gradients during the backward pass. This allows samples to maintain sparsity even at high temperatures.