Here we are introducing Neural Networks and Backpropagation.
Neural Networks
We will start with a 2-layer example:
The $max$ here creates some non-linearity between $W1$ and $W2$. It is called the activation function.
In practice, we will usually add a learnable bias at each layer as well.
Here’s an example of a 3-layer one.
The networks we are talking about here are more accurately called fully-connected networks, or sometimes called multi-layer perceptrons (MLP).
Activation Functions
The most important characteristic of activation functions is to create some non-linearity.
There are multiple activation functions to choose from.
The function $max(0,x)$ is called ReLU. It is a good default choice for most problems.
We also have:
- Leaky ReLU: $\max(0.1x, x)$ — Allows a small gradient for negative inputs. Avoids dead neurons.
- Sigmoid: $\sigma(x) = \frac{1}{1+e^{-x}}$ — S-shaped curve squashing output between 0 and 1.
- Tanh: $\tanh(x)$ — S-shaped curve squashing output between -1 and 1.
- ELU: Exponential Linear Unit.
- GELU & SiLU: Newer functions often used in modern architectures like Transformers.
Choosing an activation function is very empirical.
Architecture
In a Neural Network, we have an input layer, an output layer, and several hidden layers.
Sample Code
1 | import numpy as np |
The number of neurson in the hidden layers is often related to the capacity of information. More neurons, more capacity.
A regularizer keeps the network simple and prevents overfitting. Do not use size of neural network as a regularizer. Use stronger regularization instead.
Derivatives
The derivative on each variable tells you the sensitivity of the whole expression on its value.
- add gate: gradient distributor
- mul gate: “swap multiplier”
- copy gate: gradient adder
- max gate: gradient router
Backpropagation
The problem now is to calculate the gradients in order to learn the parameters.
Obviously we can’t simply derive it on paper, for it can be very tedious and not feasible for complex models.
Here’s the better idea: Computational graphs + Backpropagation.
Backpropagation is generally chain rules.
1 | ================================================================================ |
Modularize: Forward/Backward API
Gate / Node / Function object: Actual PyTorch code
1 | class Multiply(torch.autograd.Function): |
Aside: Vector Derivatives
For vector derivatives:
- Vector to Scalar
The derivative is a gradient.
- Vector to Vector
The derivative is a Jacobian Matrix.
Aside: Jacobian Matrix
A Jacobian Matrix is the derivative for a function that maps a Vector to a Vector.
Definition
Let $\mathbf{f} : \mathbb{R}^n \to \mathbb{R}^m$ be a function such that each of its first-order partial derivatives exists on $\mathbb{R}^n$. This function takes a point $\mathbf{x} = (x_1, \ldots, x_n) \in \mathbb{R}^n$ as input and produces the vector $\mathbf{f}(\mathbf{x}) = (f_1(\mathbf{x}), \ldots, f_m(\mathbf{x})) \in \mathbb{R}^m$ as output.
Then the Jacobian matrix of $\mathbf{f}$, denoted $\mathbf{J}_{\mathbf{f}}$, is the $m \times n$ matrix whose $(i, j)$ entry is $\frac{\partial f_i}{\partial x_j}$; explicitly:
where $\nabla^T f_i$ is the transpose (row vector) of the gradient of the $i$-th component.
Backpropagation with Vectors
1 | ================================================================================ |
Jacobians can be sparse. For element-wise operations, the off-diagonal elements are always zero; the Jacobian becomes a diagonal matrix.
Four Fundamental Equations of Backpropagation
These allow us to propagate error ($\delta$) from the output backward through the network. We get them from simple chain rules.
Notation:
- $\odot$: Hadamard (element-wise) product
- $\sigma’$: Derivative of the activation function
- $\delta^l$: Error term for layer $l$
(BP1) Error at the Output Layer
Interpretation: How much the output layer “missed” the target, scaled by the activation’s sensitivity.
(BP2) Propagating Error to Hidden Layers
Interpretation: The error at layer $l$ is the weighted sum of errors from the next layer ($l+1$), pulled backward through the transpose of the weights.
(BP3) Gradient for Biases
Interpretation: The gradient for a bias is exactly equal to the error at that neuron.
(BP4) Gradient for Weights
Interpretation: The gradient for a weight is the product of the input activation ($a^{l-1}$) and the output error ($\delta^l$).