Deep Learning 2 - Multilayer Perceptrons
Welcome back - pull up a chair. In Deep Learning 1, every prediction came from one affine transformation. That gave us regression, logistic regression, and softmax regression, but it also left us with a hard limitation: one linear classifier can draw only one hyperplane.
Today we will break that limitation. We will begin with XOR, prove why no straight boundary can solve it, and then build a network that learns a new representation in which the same problem becomes easy. That shift—from choosing features by hand to learning them—is the real beginning of deep learning.
Why One Linear Boundary Is Not Enough
XOR and linear separability
The XOR function returns 1 when exactly one of its two binary inputs is 1:
| XOR target | ||
|---|---|---|
| 0 | 0 | 0 |
| 0 | 1 | 1 |
| 1 | 0 | 1 |
| 1 | 1 | 0 |
The positive points
We can make that statement precise with convexity. A set
A linear classifier divides space into two half-spaces, and each half-space is convex. If both positive XOR examples lie in the positive half-space, their midpoint must also lie there:
If both negative examples lie in the negative half-space, their midpoint must lie there too:
The same point cannot belong exclusively to both sides. This contradiction proves that no linear threshold rule solves XOR.
A broader invariance problem
The problem is not limited to four toy points. Imagine representing a small binary image directly by its pixels and asking a classifier to distinguish two patterns at every wrap-around translation. If every translation of pattern A has the same average pixel vector as every translation of pattern B, convexity creates the same contradiction: a linear separator would be forced to assign the shared average to both classes.
This matters because vision often requires translation invariance. A pattern should retain its identity when it moves. Raw pixels do not automatically give a linear classifier the features required to express that rule.
Key idea: a linear model is not weak because optimization failed. It is weak because the correct decision function may not exist anywhere in its hypothesis class.
Feature Maps: A Temporary Escape Hatch
As in polynomial regression, we can manually transform the input. For XOR, consider
The interaction feature
Checking all four inputs shows that this rule implements XOR. In the transformed feature space, the target is linearly separable.
But this is not a general solution. Someone had to know that
A neural network makes a different proposal: learn the feature map from data at the same time as the final predictor.
Building a Multilayer Perceptron
A fully connected layer
A layer with
The activation
A feed-forward neural network connects layers as a directed acyclic graph. Information travels from input to output without returning to an earlier unit. A recurrent network, by contrast, may contain cycles or a time-unrolled recurrence.
Composing layers
For an
Equivalently,
This composition gives neural networks modularity. Each layer can be implemented as a function with a forward computation. Later, backpropagation will supply a backward computation that passes derivatives through the same modules in reverse.
Activation Functions
The activation function decides what kind of transformation follows each affine map.
Linear and hard threshold
The identity activation is
It is useful for regression outputs, but stacking identity-activated layers does not create a nonlinear model. A hard threshold,
creates discrete logic but is not differentiable at zero and has zero derivative elsewhere. It is useful for reasoning about representational power, not convenient for ordinary gradient-based training.
ReLU and softplus
The rectified linear unit is
ReLU is piecewise linear but not globally linear. Its change of slope at zero is enough to let compositions form complex piecewise-linear decision surfaces. Its derivative is 0 on the negative side and 1 on the positive side; a convention must be chosen at exactly zero.
A smooth approximation is softplus:
Its derivative is the sigmoid. Softplus approaches zero for large negative
Sigmoid and hyperbolic tangent
The logistic sigmoid is
with output in
with output in
Do not ask which activation is universally best. Ask what range, differentiability, sparsity, saturation, and optimization behavior the layer needs.
Solving XOR with a Hidden Layer
A hidden layer can carve the input space into intermediate regions and let the output recombine them. One construction uses two hard-threshold hidden units:
which behaves like OR, and
which behaves like AND. The output is
Check the four inputs:
| 0 | 0 | 0 | |
| 1 | 0 | 1 | |
| 1 | 0 | 1 | |
| 1 | 1 | 0 |
The output unit remains a linear threshold unit, but it no longer sees the raw coordinates. It sees two learned-or-designed features whose combination makes XOR linearly separable. The hidden layer did not merely add parameters; it changed the representation of the problem.
Feature Learning
The feature-map view makes an MLP easy to interpret. If the hidden stack produces
the final layer is an ordinary linear regressor or classifier applied to learned features:
Consider a
The weight vector
The word learned matters. We do not specify that unit 17 should detect a diagonal stroke. We specify an architecture and objective; optimization adjusts all layers so that the internal representation supports the final task.
This also explains why depth can be useful. One layer may combine pixels into local patterns, another may combine patterns into parts, and a later layer may combine parts into a class-relevant representation. The precise hierarchy is learned rather than guaranteed, but composition creates the possibility.
Why Nonlinearity Is Essential
Suppose every layer is linear and biases are omitted for clarity:
Let
Then
Once a nonlinear activation is placed between the matrices,
we can no longer collapse the composition into one matrix. Even ReLU—linear on each side of zero—is nonlinear enough because different inputs activate different linear regions.
Universal Approximation
What the theorem says
Under suitable conditions, a feed-forward network with a nonlinear activation and enough hidden units can approximate any continuous function on a compact domain arbitrarily well. Related constructions exist for threshold, sigmoid, tanh, and ReLU activations.
For binary inputs, there is an especially direct construction. Create one hidden threshold unit for each input configuration on which the target is 1. Each unit activates only for its assigned configuration; a linear output combines those indicators. Since there are
Sigmoid units can approximate the hard thresholds by scaling their logits:
away from the boundary. This connects a clean Boolean construction to differentiable units that can be optimized.
Boolean logic gives another perspective. Threshold units can implement AND, OR, and NOT. Any Boolean circuit built from those gates can therefore be translated into a feed-forward network.
What the theorem does not say
Universal approximation is an existence statement, not a complete learning guarantee. It does not tell us that:
- the required network is small;
- gradient descent will find the desired parameters;
- finite data is sufficient to identify the function;
- the learned model will generalize;
- the computation is efficient or numerically stable.
The naive binary construction can require exponentially many hidden units. A network expressive enough to memorize arbitrary training labels can simply overfit. The interesting question is not whether a huge network can represent a function, but whether the architecture and training process can discover a compact, reusable representation from realistic data.
From Representation to Learning
This lecture establishes three steps in the argument for neural networks:
- A linear decision boundary cannot express important functions such as XOR.
- Nonlinear hidden layers can transform the input into a representation where a simple output layer succeeds.
- Sufficiently wide nonlinear networks are universal approximators, but useful learning still requires compactness, optimization, and generalization.
We have described what an MLP can compute, but not yet how to train all of its layers. The missing tool is the multivariable chain rule applied systematically through a composition of modules. That procedure is backpropagation, and it will turn the forward equations in this article into gradients for every matrix and bias.
- Title: Deep Learning 2 - Multilayer Perceptrons
- Author: Gavin0576
- Created at : 2026-09-24 14:00:00
- Updated at : 2026-09-24 20:34:02
- Link: https://jiangpf2022.github.io/blog/2026/09/24/Deep-Learning-2-Multilayer-Perceptrons/
- License: This work is licensed under CC BY-NC-SA 4.0.