intermediate · ~20 min

Convolutions Visualized

Slide a kernel across an image by hand and watch the exact math tf.conv2d runs internally.

This module builds on Your First Neural Net. Feel free to jump ahead anyway.

A convolution slides a small grid of numbers — a kernel — over an image, multiplying and summing at every position to produce a new grid: the feature map. Different kernels detect different things — edges, blurs, corners — and in a real CNN those kernel values are learned, not hand-picked.

Beginner tip

Try each preset on the "Plus" image below and see if you can guess why it produces what it does before reading on. Blur averages every pixel with its neighbors, smoothing out sharp changes. Sharpen does roughly the opposite: it boosts a pixel relative to its neighbors, exaggerating edges. Sobel (vertical or horizontal) is a directional edge detector — it only lights up where intensity changes along one specific axis, which is why the two Sobel presets highlight different edges of the same plus shape. Edge detect is the omnidirectional version: it lights up wherever intensity changes sharply in any direction.

Change the stride and padding and watch the output grid resize — the formula underneath always matches what you see, because it's computing the exact same thing tf.conv2d does internally.

Production note

A real convolutional layer applies many kernels at once (one per output channel) and learns their values via backpropagation, the same gradient descent you saw in the Linear Regression Playground — just with thousands of parameters instead of two.

🔍 Deep dive: Why 'same' padding keeps the output size the input size

'valid' padding only computes outputs where the kernel fully overlaps the input, so the output shrinks by kernel_size − 1 per dimension. 'same' padding adds just enough zeros around the border so the output width/height equals ceil(input / stride) — with stride = 1 that's exactly the input size, which is why 'same' is the default choice when you want to stack many conv layers without the feature map vanishing.

🔍 Deep dive: From one kernel to many: what a real Conv2D layer's output shape is

This playground applies one kernel at a time so the math stays visible on screen, and the mini project below chains two kernels in sequence — but that's a stand-in for stacking two whole layers, not what one real layer does. One real Conv2D layer applies many kernels in parallel against the same input and stacks their outputs along a new channel axis: for filters = 32, a (28, 28, 1) grayscale input becomes a (28, 28, 32) output (with 'same' padding) — 32 different feature maps, one per learned kernel, computed side by side in a single layer.

🔍 Deep dive: Receptive fields: why stacking convolutions lets a network 'see' more

A single 3×3 kernel only ever looks at a 3×3 patch of its input — its receptive field is 3×3. Stack a second 3×3 conv layer on top, though, and each of its output positions depends on a 3×3 patch of the first layer's output — which itself each depended on a 3×3 patch of the original input. The receptive field compounds: two stacked 3×3 layers see a 5×5 region of the original input, three see 7×7, and so on. This is the real reason deep CNNs stack many small kernels rather than using one enormous one — it's a cheaper way (fewer parameters) to build up the same field of view.

Production note

This module only covers the convolution operation itself. Real CNN architectures almost always follow each conv layer with a pooling layer (MaxPooling2D or AveragePooling2D), which downsamples the feature map — typically halving its width and height — to reduce computation and grow the receptive field faster than convolution alone.

Playground

Input (10×10)

0
0
0
0
1
1
0
0
0
0
0
0
0
0
1
1
0
0
0
0
0
0
0
0
1
1
0
0
0
0
0
0
0
0
1
1
0
0
0
0
1
1
1
1
1
1
1
1
1
1
1
1
1
1
1
1
1
1
1
1
0
0
0
0
1
1
0
0
0
0
0
0
0
0
1
1
0
0
0
0
0
0
0
0
1
1
0
0
0
0
0
0
0
0
1
1
0
0
0
0

Kernel

Output (8×8)

0
0
-3
3
3
-3
0
0
0
0
-3
3
3
-3
0
0
-3
-3
-5
2
2
-5
-3
-3
3
3
2
1
1
2
3
3
3
3
2
1
1
2
3
3
-3
-3
-5
2
2
-5
-3
-3
0
0
-3
3
3
-3
0
0
0
0
-3
3
3
-3
0
0
1

out = floor((in − k + 2p) / stride) + 1 → 8×8 — matches the output grid above.

Mini project

Real CNNs stack many convolutional layers back to back. Chain two filters here and see that order changes the result — exactly why layer order in a CNN architecture is a real design decision, not an arbitrary one.

A then B

0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
1
1
1
1
1
1
0
0
0
0
1
0
0
0
0
1
0
0
0
0
1
0
0
0
0
1
0
0
0
0
1
0
0
0
0
1
0
0
0
0
1
0
0
0
0
1
0
0
0
0
1
1
1
1
1
1
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0

input

0
0
0
0
0
0
0
0
0
0
0
0.11
0.22
0.33
0.33
0.33
0.33
0.22
0.11
0
0
0.22
0.33
0.44
0.33
0.33
0.44
0.33
0.22
0
0
0.33
0.44
0.56
0.33
0.33
0.56
0.44
0.33
0
0
0.33
0.33
0.33
0
0
0.33
0.33
0.33
0
0
0.33
0.33
0.33
0
0
0.33
0.33
0.33
0
0
0.33
0.44
0.56
0.33
0.33
0.56
0.44
0.33
0
0
0.22
0.33
0.44
0.33
0.33
0.44
0.33
0.22
0
0
0.11
0.22
0.33
0.33
0.33
0.33
0.22
0.11
0
0
0
0
0
0
0
0
0
0
0

after A

-0.11
-0.33
-0.67
-0.89
-1
-1
-0.89
-0.67
-0.33
-0.11
-0.33
0.11
0.33
1
0.89
0.89
1
0.33
0.11
-0.33
-0.67
0.33
0.00
0.67
-0.33
-0.33
0.67
0
0.33
-0.67
-0.89
1
0.67
1.89
0.33
0.33
1.89
0.67
1
-0.89
-1
0.89
-0.33
0.33
-1.89
-1.89
0.33
-0.33
0.89
-1
-1
0.89
-0.33
0.33
-1.89
-1.89
0.33
-0.33
0.89
-1
-0.89
1
0.67
1.89
0.33
0.33
1.89
0.67
1
-0.89
-0.67
0.33
0.00
0.67
-0.33
-0.33
0.67
0.00
0.33
-0.67
-0.33
0.11
0.33
1
0.89
0.89
1
0.33
0.11
-0.33
-0.11
-0.33
-0.67
-0.89
-1
-1
-0.89
-0.67
-0.33
-0.11

after B

B then A

0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
1
1
1
1
1
1
0
0
0
0
1
0
0
0
0
1
0
0
0
0
1
0
0
0
0
1
0
0
0
0
1
0
0
0
0
1
0
0
0
0
1
0
0
0
0
1
0
0
0
0
1
1
1
1
1
1
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0
0

input

0
0
0
0
0
0
0
0
0
0
0
-1
-2
-3
-3
-3
-3
-2
-1
0
0
-2
6
5
6
6
5
6
-2
0
0
-3
5
-5
-3
-3
-5
5
-3
0
0
-3
6
-3
0
0
-3
6
-3
0
0
-3
6
-3
0
0
-3
6
-3
0
0
-3
5
-5
-3
-3
-5
5
-3
0
0
-2
6
5
6
6
5
6
-2
0
0
-1
-2
-3
-3
-3
-3
-2
-1
0
0
0
0
0
0
0
0
0
0
0

after B

-0.11
-0.33
-0.67
-0.89
-1
-1
-0.89
-0.67
-0.33
-0.11
-0.33
0.11
0.33
1
0.89
0.89
1
0.33
0.11
-0.33
-0.67
0.33
0
0.67
-0.33
-0.33
0.67
0.00
0.33
-0.67
-0.89
1.00
0.67
1.89
0.33
0.33
1.89
0.67
1.00
-0.89
-1
0.89
-0.33
0.33
-1.89
-1.89
0.33
-0.33
0.89
-1
-1
0.89
-0.33
0.33
-1.89
-1.89
0.33
-0.33
0.89
-1
-0.89
1
0.67
1.89
0.33
0.33
1.89
0.67
1.00
-0.89
-0.67
0.33
0.00
0.67
-0.33
-0.33
0.67
0.00
0.33
-0.67
-0.33
0.11
0.33
1.00
0.89
0.89
1.00
0.33
0.11
-0.33
-0.11
-0.33
-0.67
-0.89
-1
-1
-0.89
-0.67
-0.33
-0.11

after A