Built from scratch in pure JavaScript

How does this
MLP digit recognizer work?

Complete mathematical explanation behind the code.
With real numbers and interactive examples.

Educational companion to the Mobile MLP Digits Demo

SECTION 1

Network Architecture

The actual network in the demo

INPUT LAYER
196
neurons
14 × 14 pixels
HIDDEN LAYER
32
neurons
Sigmoid
OUTPUT LAYER
10
neurons
Digits 0–9
Every pixel from your drawing becomes an input value (0 or 1). The network transforms these 196 numbers into 10 output values — one for each digit.

Why these exact sizes?

  • • 14×14 — detailed enough for digits, but small (196 inputs) so it can train quickly in the browser.
  • • 32 hidden neurons — good balance. Enough capacity to recognize patterns without slowing down training in JavaScript.
  • • 10 outputs — one neuron per digit (one-hot style).
SECTION 2

The Sigmoid Activation Function

Sigmoid squashes any number into a value between 0 and 1. This is crucial because:

  • It lets the network smoothly "turn neurons on or off".
  • Its derivative is easy to compute (used heavily in backpropagation).
  • Output values can be interpreted as "how confident the network is".
FORMULA
σ(z) = 1 / (1 + e-z)
Derivative (used in backprop):
σ'(a) = a × (1 − a)   (where a = σ(z))
INTERACTIVE CALCULATOR
-6
0
+6
σ(z)
0.500
σ'(z)
0.250
SECTION 3

Forward Propagation (Predict)

This is when the network makes a prediction. Input data flows through the weights and activations and finally produces 10 numbers.

1 How is a single neuron calculated?

For every neuron in the next layer:

z = bias + Σ (weightk × activationk from previous layer)
a = sigmoid(z)
In the code:
sum = bias[j]
sum += current[k] * weights[i][j][k]

next.push( activate(sum) )
SECTION 4 • INTERACTIVE

Play with a tiny network (2 → 2 → 1)

Change the inputs and watch the numbers flow through the network in real time.

INPUT VALUES
x₁ 0.70
01
x₂ 0.25
01
WEIGHTS & BIASES (watch them change when training)
W1 (input → hidden layer) + bias1
h₁ h₂
x₁
x₂
bias
W2 (hidden → output) + bias2
out
h₁
h₂
bias
FORWARD PASS — LIVE RESULTS
Hidden Layer
h₁ (z → a)
h₂ (z → a)
Output
out (z → a)
Prediction
0.000
0.85
SECTION 5

Backpropagation Explained

This is how the network learns. It compares what it predicted with what it should have predicted, then adjusts the weights.

Steps in the code (simplified)

  1. 1
    Calculate output error
    error = target − output_activation
  2. 2
    Propagate error backwards
    For each previous layer: error = Σ (error_next × weight)
  3. 3
    Compute gradient
    gradient = error × activatePrime(activation)
  4. 4
    Update weights and biases
    weight += learningRate × gradient × prev_activation
    bias += learningRate × gradient

Why does it work?

The network uses gradient descent. It finds the direction where the error decreases fastest and takes a small step in that direction (learning rate = 0.15 in the original code).

Every time you press "Train X" in the original demo, the network runs 200 full epochs over all examples you have drawn so far. This is why it starts recognizing digits quickly even with very few examples.
SECTION 6

How does the original demo code actually work?

1
You draw → the 14×14 grid is updated in real time (with slight line thickening for better visibility).
2
You press "Train 5" → the current grid is saved as an example + one-hot target [0,0,0,0,0,1,0,0,0,0].
Then the network is trained for 200 epochs on all examples collected so far (with shuffling each epoch).
3
While drawing → guess() is called on every movement, which runs predict() and shows the most likely digit.