Full explanation with concrete values • mobile‑first
The code implements a Multilayer Perceptron (MLP) — the simplest artificial neural network.
The network learns through backpropagation each time you press "Train X".
Matrix dimensions:
| Layer | Matrix | Size | Parameters |
|---|---|---|---|
| 1 (hidden) | W₁ | 32×196 | 6,304 |
| 2 (output) | W₂ | 10×32 | 330 |
| Total | 6,634 | ||
Each parameter is initialized randomly between -0.5 and 0.5.
[0.0, 1.0, 0.8, 0.0]4 pixels → 2 hidden → 2 outputs. Input: x = [0.0, 1.0, 0.8, 0.0]
| Variable | Value |
|---|---|
| W₁[0] | [0.30, -0.50, 0.10, -0.20] |
| b₁[0] | 0.10 |
| Σ xᵢWᵢ | 0·0.30+1·(-0.50)+0.8·0.10+0 = -0.42 |
| z₁ | -0.42+0.10 = -0.32 |
| a₁ = σ(z₁) | 1/(1+e0.32) ≈ 0.421 |
| Variable | Value |
|---|---|
| W₁[1] | [0.15, 0.40, -0.30, 0.05] |
| b₁[1] | -0.05 |
| a₂ = σ(0.16-0.05) | σ(0.11) ≈ 0.527 |
👉 Hidden output: [0.421, 0.527]
| Neuron | W₂ | Bias | Σ | a = σ(z) |
|---|---|---|---|---|
| y₀ (digit 0) | [0.60, -0.40] | 0.08 | 0.042 | 0.530 |
| y₁ (digit 3) | [0.20, 0.70] | -0.10 | 0.453 | 0.587 |
📊 y₁ (0.587) > y₀ → predicted digit 3 (low confidence).
📌 Why? Squashes numbers into (0,1), introduces non‑linearity, and has a simple derivative.
Fastest learning at x=0 (σ'(0)=0.25).
Target for digit "3": [0,0,0,1,0,0,0,0,0,0]
Error: errorᵢ = targetᵢ − outputᵢ
η (learning rate) = 0.15
Output: [0.530, 0.587], target: [0, 1]
| Neuron | Output | Target | Error | σ'(z) | δᴸ |
|---|---|---|---|---|---|
| y₀ | 0.530 | 0 | -0.530 | 0.249 | -0.132 |
| y₁ | 0.587 | 1 | +0.413 | 0.242 | +0.100 |
| Hidden | W₂ᵀ×δᴸ | σ'(zᴴ) | δᴴ |
|---|---|---|---|
| h₁ | -0.059 | 0.244 | -0.014 |
| h₂ | 0.123 | 0.249 | +0.031 |
| Weight | Old | ΔW | New |
|---|---|---|---|
| W₂[y₀→h₁] | 0.60 | -0.0083 | 0.592 |
| W₂[y₁→h₂] | 0.70 | +0.0079 | 0.708 |
📈 Weights to the correct answer increase, to the wrong one – decrease.
| Step | Action | Value |
|---|---|---|
| 1 | Draw "3" on the 14×14 canvas | ~30‑40 pixels >0 |
| 2 | Normalization + anti‑aliasing | grid[45]=0.8, grid[60]=1.0… |
| 3 | Forward pass: 196→32→10 | y₃ = 0.82 |
| 4 | Find max | Index 3 → "3" |
The canvas is 14×14 pixels (visually scaled to 280×280 via CSS).
This creates a thicker, smoother stroke → easier recognition.
A single‑hidden‑layer MLP can approximate any continuous function. 32 neurons are enough for digits.
Without it the network would be linear and couldn't separate complex digit shapes.
Backprop computes the direction where error decreases fastest and moves weights that way.
200 epochs over all saved examples (with shuffling) help the network generalize.
| Operation | Formula | In code |
|---|---|---|
| Sigmoid | σ(x)=1/(1+e⁻ˣ) | activate(x) |
| σ derivative | σ'(x)=σ(x)(1−σ(x)) | activatePrime(x) |
| Forward hidden | a₁=σ(W₁·x+b₁) | predict() |
| Forward output | a₂=σ(W₂·a₁+b₂) | predict() |
| Output error | δ₂=(t−a₂)⊙σ'(z₂) | train() |
| Hidden error | δ₁=(W₂ᵀ·δ₂)⊙σ'(z₁) | train() |
| Weight update | ΔW=η·δ·a_prevᵀ | train() |
🧠 That's the whole magic — a neural network from scratch, no libraries.