🧠 How the MLP Recognizer Works

Full explanation with concrete values • mobile‑first

14×14 px 196 inputs 32 hidden 10 outputs Sigmoid Backprop

📋 Code Overview

The code implements a Multilayer Perceptron (MLP) — the simplest artificial neural network.

  • Input layer: 196 neurons (14×14 pixels)
  • Hidden layer: 32 neurons with sigmoid
  • Output layer: 10 neurons (digits 0–9)

The network learns through backpropagation each time you press "Train X".

🏗 Network Architecture

Input (196)
p₁
p₂
···
p₁₉₆
⟹ W₁
32×196
Hidden (32)
h₁
h₂
···
h₃₂
⟹ W₂
10×32
Output (10)
y₀
y₁
···
y₉

Matrix dimensions:

LayerMatrixSizeParameters
1 (hidden)W₁32×1966,304
2 (output)W₂10×32330
Total6,634

Each parameter is initialized randomly between -0.5 and 0.5.

➡️ Forward Pass – Prediction

Step by step

  1. Input vector – 196 values (0=empty, 1=filled).
    Example (4 pixels): [0.0, 1.0, 0.8, 0.0]
  2. Hidden layer: zⱼ = b₁ⱼ + Σ (xᵢ × W₁ⱼᵢ)
  3. Sigmoid: aⱼ = 1/(1+e⁻ᶻⱼ)
  4. Output layer: same with W₂ and b₂
  5. Output: 10 numbers – largest = predicted digit

🔢 Concrete example (simplified)

4 pixels → 2 hidden → 2 outputs. Input: x = [0.0, 1.0, 0.8, 0.0]

Hidden neuron h₁:

VariableValue
W₁[0][0.30, -0.50, 0.10, -0.20]
b₁[0]0.10
Σ xᵢWᵢ0·0.30+1·(-0.50)+0.8·0.10+0 = -0.42
z₁-0.42+0.10 = -0.32
a₁ = σ(z₁)1/(1+e0.32) ≈ 0.421

Hidden neuron h₂:

VariableValue
W₁[1][0.15, 0.40, -0.30, 0.05]
b₁[1]-0.05
a₂ = σ(0.16-0.05)σ(0.11) ≈ 0.527

👉 Hidden output: [0.421, 0.527]

Output layer:

NeuronW₂BiasΣa = σ(z)
y₀ (digit 0)[0.60, -0.40]0.080.0420.530
y₁ (digit 3)[0.20, 0.70]-0.100.4530.587

📊 y₁ (0.587) > y₀ → predicted digit 3 (low confidence).

📈 Sigmoid Function

σ(x) = 1 / (1 + e⁻ˣ)
σ'(x) = σ(x)·(1−σ(x))

📌 Why? Squashes numbers into (0,1), introduces non‑linearity, and has a simple derivative.

🧮 Calculator:

Fastest learning at x=0 (σ'(0)=0.25).

⬅️ Backpropagation – Learning

🎯 Goal: minimize error

Target for digit "3": [0,0,0,1,0,0,0,0,0,0]

Error: errorᵢ = targetᵢ − outputᵢ

📐 Steps

  1. Output error: δᴸ = (target−output) ⊙ σ'(zᴸ)
  2. Propagate back: δᴴ = (W₂ᵀ × δᴸ) ⊙ σ'(zᴴ)
  3. Update weights: ΔW = η × δ × a_prevᵀ
  4. Update biases: b_new = b_old + η × δ

η (learning rate) = 0.15

🔢 Concrete example (continued)

Output: [0.530, 0.587], target: [0, 1]

Step 1: δᴸ

NeuronOutputTargetErrorσ'(z)δᴸ
y₀0.5300-0.5300.249-0.132
y₁0.5871+0.4130.242+0.100

Step 2: δᴴ

HiddenW₂ᵀ×δᴸσ'(zᴴ)δᴴ
h₁-0.0590.244-0.014
h₂0.1230.249+0.031

Step 3: Updating W₂ (η=0.15)

WeightOldΔWNew
W₂[y₀→h₁]0.60-0.00830.592
W₂[y₁→h₂]0.70+0.00790.708

📈 Weights to the correct answer increase, to the wrong one – decrease.

🎯 Full Example: Drawing a "3"

StepActionValue
1Draw "3" on the 14×14 canvas~30‑40 pixels >0
2Normalization + anti‑aliasinggrid[45]=0.8, grid[60]=1.0…
3Forward pass: 196→32→10y₃ = 0.82
4Find maxIndex 3 → "3"

🖌 Canvas & Preprocessing

From pixel to number

The canvas is 14×14 pixels (visually scaled to 280×280 via CSS).

// On every move: function addInk(col, row, intensity) { const idx = row * 14 + col; grid[idx] = Math.min(1, grid[idx] + intensity); } // Main stroke + anti‑aliasing: addInk(col, row, 1.0); // center addInk(col-1, row, 0.4); // left neighbor addInk(col+1, row, 0.4); // right neighbor ...

This creates a thicker, smoother stroke → easier recognition.

💡 Why the Code Works

1. Universal Approximation

A single‑hidden‑layer MLP can approximate any continuous function. 32 neurons are enough for digits.

2. Sigmoid = Non‑linearity

Without it the network would be linear and couldn't separate complex digit shapes.

3. Gradient Descent

Backprop computes the direction where error decreases fastest and moves weights that way.

W_new = W_old − η × ∂Error/∂W

4. Dataset + Repetition

200 epochs over all saved examples (with shuffling) help the network generalize.

⚡ Summary

1. Draw → 14×14 → 196 numbers 2. Forward: 196→32→10 3. Train X → target=[0..1..0] 4. Backprop: error → gradient → update W,b 5. Repeat 200× on all saved examples 6. The network becomes more accurate! 🎉

📐 Formula Quick Reference

OperationFormulaIn code
Sigmoidσ(x)=1/(1+e⁻ˣ)activate(x)
σ derivativeσ'(x)=σ(x)(1−σ(x))activatePrime(x)
Forward hiddena₁=σ(W₁·x+b₁)predict()
Forward outputa₂=σ(W₂·a₁+b₂)predict()
Output errorδ₂=(t−a₂)⊙σ'(z₂)train()
Hidden errorδ₁=(W₂ᵀ·δ₂)⊙σ'(z₁)train()
Weight updateΔW=η·δ·a_prevᵀtrain()

🧠 That's the whole magic — a neural network from scratch, no libraries.