Signal
Issue 8 · August 26, 2026 · 4 min read
Cross-Entropy Is the Loss Behind Vision, Language, and Robot Policies
ResNet, GPT next-token training, YOLO class heads, and DeepMind RT-2 action tokens all optimize the same -log(p_correct). Learn the production form from logits.
"The maximum likelihood estimator is equivalent to minimizing the cross-entropy between the training data and the model's predictions." - Ian Goodfellow, Yoshua Bengio, and Aaron Courville (Deep Learning)
Every serious classifier ends up asking the same question: how surprised am I by the correct label?
This week we're talking about cross-entropy loss, the objective behind most vision, language, and robot-policy models in production, and yes, this can get you a very comfortable job.
Here's the idea: cross-entropy charges your model Predict 0.9 on the right class and you pay about 0.105. Predict 0.01 and you pay about 4.6. Being confidently wrong is expensive on purpose. That single formula shows up everywhere teams ship probability models. ImageNet-style CNNs (ResNet and friends) train with categorical cross-entropy over thousands of classes. Ultralytics YOLO-family detectors still lean on binary cross-entropy for class and objectness heads while they regress boxes with a separate localization loss. Autoregressive language models, including every GPT-style stack, train with next-token cross-entropy over a vocabulary that can be 50k+ tokens. Perplexity, the number people quote when they compare language models, is just of the average per-token cross-entropy.
Robotics caught the same habit. Google DeepMind's RT-2 turns continuous robot actions into discrete tokens (256 bins per dimension), then co-fine-tunes a vision-language model with ordinary next-token cross-entropy on those action strings mixed with web VQA data. Their evaluation ran about 6,000 robot trials and reported roughly 2x better generalization than prior baselines on novel objects, symbols, and reasoning-style commands. Same loss. New embodiment.
To see why this matters in practice, watch what happens when people swap in MSE for classification. Squared error on probabilities is bounded and soft on confident mistakes. Cross-entropy is not. That is why production stacks (PyTorch CrossEntropyLoss, TensorFlow sparse categorical CE) fuse softmax and NLL into one numerically stable op and keep log-softmax away from zero. A loss in training is usually a zero probability hitting a log, not a mysterious CUDA bug.
So the statement is blunt: if a system outputs a distribution over discrete choices, whether ImageNet labels, next tokens, or robot action bins, you will almost certainly train it with cross-entropy. Learn that loss cold and you can read half the training code in modern ML without guessing.
Now I'm going to explain what this means for you:
The Skill Employers EXPECT
Learn binary and categorical cross-entropy by hand. Interview loops and code reviews assume you can:
- Write BCE for sigmoid outputs and CE for one-hot / class-index labels
- Explain why punishes confident errors harder than MSE
- Clip probabilities or use log-softmax so never appears
- Match PyTorch / NumPy results on a tiny batch within floating-point noise
- Convert between mean loss, sum loss, and per-token / per-example reporting
If you only ever call without knowing the scalar, you will struggle when the curve misbehaves.
The Skill That Separates You
Learn softmax cross-entropy from logits (the production form).
Libraries rarely take probabilities. They take raw scores. You should know how to:
- Implement a stable softmax (subtract max before exp)
- Compute CE as without materializing an unstable intermediate
- Derive that the gradient w.r.t. logits collapses to
- Apply label smoothing and class weights when the label distribution is skewed
- Read a YOLO / LLM / VLA training config and point to which head is on CE vs a regression loss
The Project To Learn These Skills This Week
Ship a tiny multi-class trainer that only uses your CE.
Dataset: Fashion-MNIST (or CIFAR-10 if you want more pain). 10 classes is enough.
Requirements:
- Implement
softmax_cross_entropy(logits, y)from scratch in NumPy (or PyTorch ops you wrote, notnn.CrossEntropyLoss). Stable log-softmax required - Train a small MLP or CNN for at least 5 epochs. Log train CE, val CE, and accuracy each epoch
- Hit at least 85% test accuracy on Fashion-MNIST (or 60% on CIFAR-10) with your hand-rolled loss
- Add one ablation: train the same net with MSE on one-hot targets for 5 epochs. Plot CE-trained vs MSE-trained val accuracy on one chart
- Break it on purpose once: feed unclipped probabilities of exactly 0 into a naive path, show the nan, then show the stable fix
- Write a half-page note mapping your loss to one industrial system: ResNet classification head, GPT next-token head, or RT-2 action tokens. Same math, different vocabulary size
This project teaches a practical lesson for production ML:
Cross-entropy is how discrete decisions get a gradient. Vision, language, and robot policies all rent the same meter.
If you want to sharpen your machine learning skills even more, I also selected a challenge problem for you this week:
If you learned something from this newsletter, make sure to forward it to a friend.