Signal archive

Signal

Issue 12 · September 29, 2026 · 7 min read

Neural Networks, Explained Simply (With Gems for Pros)

How a neural network works in plain English: neurons, layers, backprop, hidden layer sizing and batch size, with a gem in every section for experienced readers.

"Neural net training is a leaky abstraction." - Andrej Karpathy (A Recipe for Training Neural Networks)

Every AI headline this year, from chatbots to self-driving cars, runs on one idea from the 1950s: multiply, add, squash, repeat.

This week we're going back to basics with how a neural network works, in plain English. If you already train these for a living, look for the Gem after each section, and yes, this can get you a very comfortable job.

Here's the idea: a neural network is a function built from layers of simple units called neurons. Each neuron multiplies its inputs by weights, adds them up, and passes the result through a simple nonlinear function. Training means nudging millions of those weights, a little at a time, until the network's outputs match the examples it has seen.

One neuron

Think of a neuron as a weighted vote. Each input gets a weight that says how much it matters, the neuron adds everything up plus a bias, and then it applies a rule. The most common rule, ReLU, just says: if the total is negative, output zero. That is the entire neuron. Frank Rosenblatt built one in hardware in 1958 and called it the perceptron.

Gem: in 1969 Minsky and Papert showed a single perceptron cannot learn XOR. The fix is a hidden layer plus a nonlinearity. Remove the nonlinearity and any stack of layers collapses back into one linear function, so depth buys you nothing.

Layers, and how the network learns

Stack neurons into layers and the outputs of one layer become the inputs of the next. To learn, the network makes a prediction, a loss function scores how wrong it was, and gradient descent moves every weight a small step in the direction that reduces the loss. Backpropagation is how the network works out that direction for millions of weights at once: it is the chain rule from calculus, applied one layer at a time from the output back to the input. Rumelhart, Hinton and Williams popularized it in 1986, and in 2024 Hinton shared a Nobel Prize for this line of work.

Gem: for a softmax output trained with cross-entropy, the gradient with respect to the logits is simply the predicted probabilities minus the one-hot label. That one clean line is why this pairing is everywhere. We covered it in issue 8, and the full derivation lives on the backpropagation page.

Is a neural network the same as AI, machine learning or deep learning?

No, they are nested. AI is the broad goal of machines doing tasks that need intelligence. Machine learning is the part of AI that learns from data. A neural network is one family of machine learning models. Deep learning means neural networks with many layers. An LLM is a very large deep neural network, a transformer, trained to predict the next token.

TermWhat it isExample
AIMachines doing tasks that normally need intelligenceA chess engine, a chatbot
Machine learningModels that learn patterns from dataA spam filter, a price forecast
Neural networkA model built from layers of weighted sums and nonlinearitiesA small MLP scoring loan risk
Deep learningNeural networks with many layersAn image classifier like ResNet
LLMA very large transformer trained on textGPT, Claude, Llama

How many hidden layers and neurons do you need?

Fewer than you think. For tabular data, start with one or two hidden layers whose width sits between the number of inputs and a few times that. Train, check the validation score, and grow only while it keeps improving. More capacity with no improvement means you are memorizing. There is no formula, only the validation curve.

Gem: the universal approximation theorem (Cybenko, 1989) says one hidden layer is enough to approximate any continuous function, in principle. In practice that layer may need to be absurdly wide, and depth gets the same result with far fewer neurons. That gap is the real argument for deep learning.

What is batch size in a neural network?

Batch size is how many training examples the network looks at before it takes one gradient step, usually 32 to 512. Bigger batches give smoother gradients and keep a GPU busy. Smaller batches are noisier, and that noise often helps the model generalize.

Gem: when you multiply the batch size by k, multiply the learning rate by k too, and warm it up over the first few epochs. That linear scaling rule is how Goyal et al. trained ImageNet in one hour in 2017. Change the batch size alone and your old learning rate is quietly wrong.

Now I'm going to explain what this means for you:

The Skill Employers EXPECT

Learn to build a small neural network by hand. Neural network interview questions start here, and you should be able to:

  • Write the forward pass of a two-layer network in NumPy, with the shapes of every matrix right
  • Explain loss, gradient, learning rate and batch size in one sentence each
  • Say why a network with no nonlinearity is just a linear model
  • Train a small MLP in Python with scikit-learn or PyTorch and read its training and validation curves

If you can explain XOR, you can explain why neural networks exist.

The Skill That Separates You

Learn to debug training instead of rerunning it. The gems for the experienced crowd:

  • Overfit one batch first. If the loss will not go to nearly zero on 32 examples, the bug is in your code, not your data.
  • Initialize for your activation. He initialization sets weight variance to 2 divided by the number of inputs so ReLU signals neither vanish nor explode.
  • Watch for dead ReLUs: a neuron whose input is negative for every example gets zero gradient forever, and a high learning rate can kill a large share of them.
  • Remember where the parameters are. In a standard transformer block roughly two-thirds of the weights sit in the plain MLP layers, not in attention.
  • Know that big networks are mostly slack. The lottery ticket hypothesis (Frankle and Carbin, 2019) found sparse subnetworks at 10 to 20% of the original size that train to the same accuracy.

The Project To Learn These Skills This Week

Go from a raw table to a trained, regularized neural network, and prove it beats a straight line. This is a graded Pro project on GRADuateML: Neural Network from Scratch Path. The table has a decision boundary no straight line can draw, the test set is hidden, and every milestone is graded instantly. Not on Pro? The free challenge below covers the forward pass by hand.

Requirements (each one is a graded milestone):

  • Profile the table, then standardize features using training statistics only.
  • Ship a logistic regression baseline and measure how far a straight line gets.
  • Choose the architecture before you train, and write down why.
  • Train an MLP to at least 0.85 accuracy, 0.80 F1 and 0.90 ROC-AUC on hidden data.
  • Regularize it, show it still clears the bar, and document your choices for your portfolio.

This project teaches a practical lesson for production ML:

A neural network earns its place only when it beats the simple baseline you were honest enough to build first.

If you want to sharpen your machine learning skills even more, I also selected a free challenge problem for you this week. It is the two-neuron network that solves XOR:

If you learned something from this newsletter, make sure to forward it to a friend.

Two-Layer Network Forward Pass (and Why It Solves XOR)

Easy · ~15 min

Concept: neural network basics

GRADuateML

The full Neural Network from Scratch Path, graded against a hidden test set, is part of GRADuateML Pro.

Ready to practice?

Turn weekly insights into hands-on ML skills on GRADuateML.