A multilayer perceptron written in NumPy. Forward, reverse-mode autodiff, SGD with momentum, He / Xavier init, fused softmax-cross-entropy, and a raw IDX parser for MNIST. No PyTorch. No TensorFlow. No Keras.
The whole engine lives in one notebook: Neural_Network_From_Scratch.ipynb.
loss.backward() is one line. The reason it works is three matrix products and a Hadamard product, and most people who ship models have never written them.
This repo is that missing page:
- XOR with tiny Gaussian weights sits at prediction
0.5. Same net, Xavier init, same 1000 steps, loss ~2e-3. - A numerical gradient check on
dWlands at relative error8.7e-12. - MNIST from the raw IDX bytes, 40 epochs. The notebook in this tree finished at 98.04 percent test accuracy and 100 percent train accuracy. That gap is overfitting, not a mystery.
The notebook is allowed to redefine Dense, Activation, and SoftmaxCrossEntropy as you go. Early cells are the historical artifact. Later cells are the engine you would actually keep.
git clone https://github.com/hrnrxb/DFNN
cd DFNN
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
jupyter notebook Neural_Network_From_Scratch.ipynbThen Run All. First launch downloads four MNIST .gz files and checks them against the TensorFlow Datasets SHA-256 checksums.
The learning code is NumPy. Matplotlib draws figures. Manim Community is optional; the MP4s already live in visualizations/rendered/.
| Path | What it is |
|---|---|
Neural_Network_From_Scratch.ipynb |
The course. 18 sections, XOR, gradient check, MNIST. |
visualizations/manim_scenes.py |
Five Pango-only Manim scenes. No LaTeX. |
visualizations/diagrams.py |
Static posters used by the notebook and this README. |
figures/ |
Last frames and matplotlib PNGs. |
LICENSE |
MIT. |
Do not commit *.gz, uploads/, or assembler scrap. .gitignore already blocks the gzip files.
Here are some best Vizs; Open the notebook if you want the rest next to the code that earned them.
XOR is two diagonals. A line cannot cut them. A fold can.
Same 2-3-1 tanh net. Left: W ~ 0.01 N(0,1), MSE = 0.25. Right: Xavier, XOR solved.
Backprop is a protocol. One Jacobian-vector product per layer.
The only product whose shape is W:
Softmax + cross-entropy never allocates the C x C Jacobian. The fused gradient is (P - Y) / N.
The calculus is not a vibe. Finite differences vs dW:
MNIST topology that actually trains: 784 → 256 (He, ReLU) → 64 (He, ReLU) → 10 (Xavier), softmax fused into the loss.
Five short films, 720p30, Pango Text only.
| File | Scene | Point |
|---|---|---|
01_affine_loom.mp4 |
AffineLoom | Linear maps shear. Tanh folds. XOR becomes linearly separable. |
02_blame_telegraph.mp4 |
BlameTelegraph | Cache X and Z going forward. Walk the same path backward. |
03_jacobian_collapse.mp4 |
JacobianCollapse | diag(P) - P P^T times -Y/P is P - Y. |
04_chain_rule_contract.mp4 |
ChainRuleContract | (Fin, N) @ (N, Fout) = (Fin, Fout). |
05_dead_signal.mp4 |
DeadSignal | Tiny init is not a saddle. It is a network that has not moved. |
If ffmpeg is missing, imageio-ffmpeg ships a binary. The notebook prepends it to PATH when needed.
Dense, Z = X W + b, batch axis 0:
dW = X.T @ dZ # (Fin, Fout)
db = dZ.sum(axis=0, keepdims=True)
dX = dZ @ W.T
Activation:
dX = dZ * f'(Z) # Hadamard, not a matmul
Fused softmax + mean categorical cross-entropy:
dZ = (P - Y) / N
MSE with np.mean matches PyTorch MSELoss(reduction="mean"): divide by the number of elements, not by batch size alone. Docs.
Xavier (Glorot normal) uses sqrt(2 / (fan_in + fan_out)). He uses sqrt(2 / fan_in). Keras glorot_normal is the truncated-normal cousin of the same scale.
The notebook's SGD is an EMA smoother v = β v + (1-β) g. It is not bit-identical to Polyak heavy ball or to torch.optim.SGD(momentum=β).
Papers this code actually implements, with a URL you can click:
- Rumelhart, Hinton, Williams (1986). Learning representations by back-propagating errors. Nature. doi:10.1038/323533a0
- Glorot, Bengio (2010). Understanding the difficulty of training deep feedforward neural networks. AISTATS. PDF
- He, Zhang, Ren, Sun (2015). Delving deep into rectifiers. ICCV. PDF
- Kingma, Ba (2015). Adam: A method for stochastic optimization. ICLR. arXiv:1412.6980
- Srivastava, Hinton, Krizhevsky, Sutskever, Salakhutdinov (2014). Dropout. JMLR. HTML · PDF
- Polyak (1964). Some methods of speeding up the convergence of iteration methods. doi:10.1016/0041-5553(64)90137-5
- Cybenko (1989). Approximation by superpositions of a sigmoidal function. doi:10.1007/BF02551274
- Hornik (1991). Approximation capabilities of multilayer feedforward networks. doi:10.1016/0893-6080(91)90009-T
- Dauphin et al. (2014). Identifying and attacking the saddle point problem. NeurIPS. arXiv:1406.2572
- LeCun, Cortes, Burges. The MNIST database. yann.lecun.com/exdb/mnist
- Bridle (1990). Probabilistic interpretation of feedforward classification network outputs. doi:10.1007/978-3-642-76153-9_28
- TensorFlow Datasets, MNIST checksums:
mnist.txt - PyTorch
MSELossreductionmean: docs
Visual primers, not papers: 3Blue1Brown, Neural Networks · Karpathy, micrograd
MIT.









