Neural Network
Neural Network
Definition: A model made of layers of interconnected nodes (“neurons”) that transform input data through weighted connections and nonlinear activation functions to learn complex patterns.
How It Works
- Input layer receives features, hidden layers transform them via weights + activation functions, output layer produces the prediction
- Each connection has a learnable weight; training (via backpropagation + gradient descent) adjusts these weights to reduce loss
- Stacking more layers (“deep” learning) lets the network learn increasingly abstract representations — early layers detect edges/simple patterns, later layers combine them into higher-level concepts
- Every neuron computes a weighted sum of its inputs plus a bias term, then passes the result through a nonlinear activation function before forwarding it to the next layer
- Training alternates two passes: a forward pass that produces a prediction and a loss value, and a backward pass that computes gradients of that loss with respect to every weight and updates them
- A batch of examples (not just one) is typically passed through at once — this “mini-batch” approach smooths out noisy gradient estimates while still updating weights far more often than processing the entire dataset per step
Network Architecture, Visualized
The canonical picture of a feedforward network — an input layer, one or more hidden layers, and an output layer, with every neuron in one layer connected to neurons in the next. The diagram below shows a 3-input, 4-hidden-neuron, 2-output network; full connectivity between adjacent layers would mean 12 edges between the input and hidden layers alone, so only a representative subset is drawn here for legibility:
Every arrow is a weighted connection — a learnable number that scales the signal flowing along it — and every node except the input layer also adds a bias and applies an activation function before passing its result forward (see “A Single Neuron’s Internals” below). Real networks draw every edge between adjacent layers, not just the subset shown; a fully connected 3-to-4-to-2 network like this one has 3*4 + 4*2 = 20 weighted edges plus 6 biases (one per hidden and output neuron), all of them learned during training rather than hand-set.
Under the Hood
- A single neuron computes
z = w1*x1 + w2*x2 + ... + wn*xn + b, then applies an activationa = f(z)— common choices are ReLU, sigmoid, tanh, or GELU - A full layer is just this computation vectorized:
Z = W @ X + b, whereWis a weight matrix andXis the batch of input vectors — this is why GPUs (built for matrix multiplication) accelerate neural nets so well - The network as a whole is a composed function:
y_hat = f_L(W_L * f_{L-1}(... f_1(W_1*x + b_1) ...) + b_L)— depth is literally function composition - Parameter count grows fast: a single dense layer mapping 784 inputs to 256 outputs already has
784 * 256 + 256 = 200,960learnable parameters - Backpropagation computes gradients efficiently via the chain rule, propagating the error signal from the output layer back to the input layer in a single backward pass, reusing intermediate computations rather than recomputing derivatives layer by layer from scratch
- The loss surface of a deep network is highly non-convex, with many local minima and saddle points — in practice, stochastic optimizers like Adam or SGD with momentum still reliably find “good enough” minima that generalize well, which is one of deep learning’s more empirically-surprising properties
A Single Neuron’s Internals
Every node in the diagram above — regardless of layer — runs the exact same small computation. Zooming into just one hidden neuron:
This is z = w1*x1 + w2*x2 + w3*x3 + b followed by a = f(z), drawn out as a pipeline rather than a formula — the same computation described in the bullets above, just made spatially explicit.
Forward and Backward Passes, Visualized
Training a network means running this per-neuron computation forward through every layer, then running an error signal back through the same layers in reverse. As a sequence of steps rather than a static diagram:
The forward arrows and backward arrows in this diagram traverse the exact same connections, just in opposite directions and carrying different cargo — activations going forward, gradients going backward. This is also why Backpropagation is cheap relative to naively recomputing every derivative from scratch: it reuses the forward pass’s intermediate values instead of discarding them.
Worked Example: A Tiny Forward Pass by Hand
Concrete numbers make the diagrams above easier to trust. Take a network with 2 inputs, a hidden layer of 2 ReLU neurons, and 1 linear output neuron:
- Inputs:
x1 = 1.0,x2 = 0.5 - Hidden neuron 1: weights
0.6, -0.3, bias0.1 - Hidden neuron 2: weights
0.2, 0.8, bias-0.2 - Output neuron: weights
0.5, -0.4(applied to h1, h2), bias0.05, linear activation
| Step | Computation | Result |
|---|---|---|
| Hidden neuron 1, pre-activation | 0.6(1.0) + (-0.3)(0.5) + 0.1 | 0.55 |
| Hidden neuron 1, activation | ReLU(0.55) | 0.55 |
| Hidden neuron 2, pre-activation | 0.2(1.0) + 0.8(0.5) - 0.2 | 0.40 |
| Hidden neuron 2, activation | ReLU(0.40) | 0.40 |
| Output, pre-activation | 0.5(0.55) + (-0.4)(0.40) + 0.05 | 0.165 |
| Output (linear activation) | identity(0.165) | 0.165 |
That 0.165 is the network’s full prediction for this input — nine learned numbers (six weights, three biases) and two nested function calls, computed exactly as the “Under the Hood” formulas describe. The interactive sandbox in the Code Example section below runs this identical computation; running it should reproduce 0.165 exactly. Backpropagation’s job, not shown here, is to compare that prediction against a target label, compute a loss, and push a gradient back through this same graph to adjust all nine numbers toward a better prediction next time.
Types
- Feedforward (MLP) — information flows strictly forward, no loops; good for tabular data and as the “head” on top of other architectures
- Convolutional Neural Network (CNN) — shares weights across spatial locations via convolution kernels; the default for images and grids
- Recurrent Neural Network (RNN) / LSTM (Long Short-Term Memory) — shares weights across time steps and maintains a hidden state; built for sequences
- Transformer — replaces recurrence with self-attention, letting every position attend to every other position directly; now dominant for text, and increasingly vision and audio
- Autoencoder — trained to reconstruct its own input through a compressed bottleneck, useful for dimensionality reduction and anomaly detection
- GAN (Generative Adversarial Network) — two networks (generator, discriminator) trained adversarially against each other
Why It Matters
- The foundational architecture behind virtually all modern AI breakthroughs — CNNs, RNNs, and transformers are all neural networks with specialized structure
- Universal approximation theorem: a network with even a single sufficiently wide hidden layer can in theory approximate any continuous function on a bounded domain — depth is what makes this practical rather than just theoretical, since deep narrow networks approximate complex functions with far fewer total parameters than shallow wide ones
- Differentiable end-to-end: because every operation in the network is differentiable, the whole pipeline — feature extraction included — can be learned jointly instead of hand-engineered
- Representation learning: rather than hand-crafting features (edges, corners, word frequencies), the network discovers its own internal representations directly from raw data, which is why deep learning displaced most classical feature-engineering pipelines in vision and NLP
- Transfer learning depends on this property directly — a network’s early/middle layers learn broadly reusable representations that can be repurposed for a new task with far less data than training from scratch, see Transfer Learning
Code Example
import torch
import torch.nn as nn
class SimpleNet(nn.Module):
def __init__(self, in_features=784, hidden=128, out_classes=10):
super().__init__()
self.net = nn.Sequential(
nn.Linear(in_features, hidden),
nn.ReLU(),
nn.Linear(hidden, hidden),
nn.ReLU(),
nn.Linear(hidden, out_classes),
)
def forward(self, x):
return self.net(x) # raw logits; apply softmax/CrossEntropyLoss externally
model = SimpleNet()
optimizer = torch.optim.Adam(model.parameters(), lr=1e-3)
loss_fn = nn.CrossEntropyLoss()
# one training step
logits = model(batch_x)
loss = loss_fn(logits, batch_y)
loss.backward()
optimizer.step()
optimizer.zero_grad()
The same tiny forward pass from the Worked Example above, translated to plain JavaScript so it runs directly on this page — press Run and compare the printed output to the 0.165 computed by hand:
Change any weight, bias, or input above and re-run to see how the prediction shifts — this is the entire “forward pass” half of training; backpropagation is what computes how to change those numbers intelligently instead of by hand.
Comparison
| Architecture | Weight sharing | Best for | Struggles with |
|---|---|---|---|
| Feedforward (MLP) | None | Tabular data, small fixed-size inputs | Images, sequences, high-dim raw data |
| CNN | Spatial (kernel slides across input) | Images, grids, local patterns | Long-range/global dependencies |
| RNN/LSTM | Temporal (same weights per step) | Sequences, time series | Long sequences, parallelization |
| Transformer | None (attention is dynamic, not shared weights) | Text, long-range dependencies, large-scale pretraining | Compute/memory cost on very long sequences |
| Autoencoder | None (symmetric encoder/decoder) | Dimensionality reduction, denoising, anomaly detection | Generating genuinely novel (not just reconstructed) samples |
| GAN | None (two adversarially trained networks) | Realistic sample generation, data augmentation | Stable training, mode collapse |
Real-World Example
- Image recognition — a CNN like ResNet classifies photos into thousands of categories by learning hierarchical filters: edges in early layers, textures and shapes in middle layers, whole objects in later layers
- Machine translation — transformer-based encoder-decoder networks map a sentence in one language to another, attending over the entire source sentence at each output step
- Recommendation systems — embedding layers (a specialized neural network component) learn dense vector representations of users and items, so that similarity in that vector space predicts likely engagement
- Speech recognition — networks combining convolutional and recurrent (or transformer) layers convert raw audio waveforms into text transcriptions
- Game-playing agents — AlphaGo and its successors pair a neural network (evaluating board positions and candidate moves) with tree search, trained via Reinforcement Learning against itself
Named Case Studies
- AlexNet (2012) — Alex Krizhevsky, Ilya Sutskever, and Geoffrey Hinton’s 8-layer CNN won the ImageNet Large Scale Visual Recognition Challenge with a 15.3% top-5 error rate against the runner-up’s 26.2% — a nearly 11-point margin so large it convinced much of the computer vision field to switch to deep learning within about two years
- AlphaFold 2 (2020) — DeepMind’s system, built around attention-based neural networks layered with domain-specific geometric reasoning, predicted 3D protein structures at CASP14 with a median accuracy close to experimental methods, a result widely reported as resolving a roughly 50-year-old open problem in structural biology
- Google’s Neural Machine Translation (2016) — Google Translate’s switch from phrase-based statistical translation to a deep LSTM-based encoder-decoder network (GNMT) cut translation errors by a reported 55-85% depending on the language pair, one of the clearest public before/after demonstrations of neural networks displacing a mature classical pipeline
Deeper Dive: Why Depth Helps
- Each layer can be thought of as re-representing the input in a new coordinate space — a shallow network has to carve up the input space with a limited number of decision boundaries in one shot, while a deep network builds up complex boundaries incrementally, layer by layer
- Empirically, doubling depth tends to be more parameter-efficient than doubling width for most vision and language tasks, though very deep networks need architectural help (residual/skip connections, normalization layers) to remain trainable
- Depth also enables hierarchical feature reuse — a CNN’s later layers combine early-layer edge detectors in different arrangements to detect eyes, wheels, or text, without re-learning “what an edge is” from scratch at every layer
Common Pitfalls
- Assuming bigger networks are automatically better — they need proportionally more data and compute, or they overfit
- Skipping activation functions between layers, which collapses the whole network into a single linear transformation no matter how many layers you stack
- Initializing all weights to zero (or the same value) — every neuron in a layer then computes identical gradients and the network never breaks symmetry
- Ignoring input scaling/normalization — features on wildly different scales slow convergence and can destabilize training
- Choosing a learning rate without a schedule or warmup for deep networks, causing early divergence
- Treating architecture choice as more important than data quality — a mediocre architecture with clean, plentiful data usually beats a state-of-the-art architecture trained on noisy or scarce data
- Reusing the same random seed across every experiment and mistaking a lucky initialization for a genuinely better architecture or hyperparameter choice
- Evaluating only on the training set, or on a validation set that leaked into training through preprocessing decisions made before the split — either one produces metrics that look good and mean nothing
- Deploying a network trained on clean, curated data directly against messier real-world inputs without checking for distribution shift between the two
Best Practices
- Normalize or standardize inputs before feeding them into the network
- Use He initialization with ReLU-family activations, Xavier/Glorot with tanh/sigmoid — matching init to activation prevents vanishing/exploding activations at the start of training
- Add Batch Normalization or layer normalization in deeper networks to stabilize and speed up training
- Start with a known-good architecture and optimizer (Adam, lr around 1e-3) before customizing
- Monitor train vs validation loss every epoch to catch Overfitting vs Underfitting early
- Checkpoint model weights periodically during long training runs so a crash or a bad late-stage update doesn’t destroy hours of progress
- Version both the code and the exact data/preprocessing used for a trained model — reproducing a specific network’s behavior later requires both
- Split data into train/validation/test sets before any preprocessing decisions are made, so no information from validation or test ever leaks into training, see Cross-Validation
- Fix a random seed while debugging so a training run is reproducible, then rerun with multiple seeds before trusting a result — a single lucky or unlucky initialization can make a mediocre configuration look great or a good one look bad
- Profile where training time actually goes (data loading, forward pass, backward pass) before optimizing — it’s common to spend effort speeding up compute when a slow data pipeline was the real bottleneck
History
- McCulloch-Pitts threshold neuron (1943) — Warren McCulloch and Walter Pitts’ paper “A Logical Calculus of the Ideas Immanent in Nervous Activity” modeled a neuron as a weighted sum compared against a threshold, showing that networks of such units could in principle compute any logical function — 15 years before anyone tried to train one from data
- Perceptron (1958, Rosenblatt) — single-layer, could only learn linearly separable functions, which stalled research after Minsky and Papert’s 1969 critique
- Backpropagation popularized for multi-layer networks in 1986 (Rumelhart, Hinton, Williams), enabling training of deeper models
- 1990s-2000s “AI winter” for neural nets specifically — support vector machines and other classical methods often outperformed neural nets given the data and compute available at the time
- Deep learning’s modern resurgence began around 2012 with AlexNet’s ImageNet win, driven by GPU compute, large labeled datasets, and ReLU activations replacing sigmoid/tanh
- “Attention Is All You Need” (2017) introduced the transformer, eventually displacing RNNs as the default architecture for sequence tasks and later much of vision and audio too
- LeCun et al.’s LeNet-5 (1998) applied backpropagation-trained convolutional networks to handwritten digit recognition at production scale for check processing, years before “deep learning” was a common term
- He et al.’s residual networks, ResNet (2015), introduced skip connections that made it practical to train networks hundreds of layers deep, directly addressing the degradation problem plain stacking runs into
FAQ
Q: How many layers make a network “deep”? No strict cutoff — informally, more than one hidden layer counts as deep. Modern networks routinely use dozens to hundreds of layers.
Q: Why not just use one giant hidden layer instead of many small ones? A single wide layer can approximate the same functions in theory but typically needs exponentially more neurons than a deep, narrow stack to represent the same function.
Q: Do neural networks need labeled data? Not necessarily — Supervised Learning uses labels, but the same architectures train under Unsupervised Learning (autoencoders) and Reinforcement Learning (policy/value networks).
Q: Can a neural network have zero hidden layers? Yes — that’s just linear (or logistic) regression, a single layer mapping inputs directly to outputs with no intermediate representation. “Neural network” as a useful, distinct term generally implies at least one hidden layer.
Common Interview Questions
Q: Why do we need non-linear activation functions? Without them, any stack of linear layers collapses algebraically into a single linear transformation, so the network could only ever learn linear decision boundaries regardless of depth.
Q: What’s the difference between a parameter and a hyperparameter? Parameters (weights, biases) are learned automatically from data during training. Hyperparameters (learning rate, layer count, batch size) are set by the practitioner before training starts and control how learning happens.
Q: What happens if you remove all the biases from a network? Every neuron’s output is forced to pass through the origin when inputs are zero, which restricts the family of functions the network can represent and typically hurts fit quality, especially in shallow networks.
Q: What’s the difference between a neuron’s “weights” and its “activation”? Weights are the learned numbers that scale each input signal; activation is the output value produced after summing the weighted inputs and passing the result through a nonlinear function. Weights are fixed after training (until fine-tuned); activations change with every new input.
Q: Why do deep networks sometimes train worse than shallower ones, despite having more capacity? Naively stacking many layers can cause vanishing or exploding gradients, making early layers train slowly or unstably. Techniques like residual connections (skip connections), batch normalization, and careful initialization specifically exist to make very deep networks trainable at all.
Q: How would you explain backpropagation to someone who only knows calculus, not deep learning?
It’s the chain rule applied systematically: the loss is a function of the output, the output is a function of the last layer’s weights, those activations are a function of the previous layer’s weights, and so on back to the input. Backpropagation computes dLoss/dWeight for every weight by multiplying local derivatives along this chain, reusing shared intermediate results instead of recomputing them for every weight independently.
Related Terms
- Backpropagation
- Activation Function
- Convolutional Neural Network (CNN)
- Recurrent Neural Network (RNN)
- Gradient Descent
- Vanishing-Exploding Gradient
- Loss Function
- Epoch, Batch, and Iteration
- Hyperparameter Tuning
- Regularization (L1, L2, Dropout)
- Autoencoder
- GAN (Generative Adversarial Network)
Example
A simple network with one hidden layer can learn to classify handwritten digits (0-9) from pixel values, given enough labeled examples. Feed it a 28x28 grayscale image (784 pixel values flattened into a vector), pass it through a hidden layer of, say, 128 neurons with ReLU activation, then an output layer of 10 neurons (one per digit) with softmax — the network outputs a probability distribution over the 10 possible digits, and training adjusts its ~100K+ weights until those probabilities line up with the correct labels across thousands of examples. Run the same trained network on a digit it has never seen before, written in an unfamiliar handwriting style, and it generalizes because it learned patterns like “loops and curves in this arrangement” rather than memorizing exact pixel layouts.
This same basic recipe — flatten or embed the input, pass it through learned linear transformations and nonlinearities, compare the output to a target, and adjust weights to reduce the error — scales up essentially unchanged from this 100K-parameter digit classifier to a large language model with hundreds of billions of parameters. What changes between them is architecture (dense layers vs. attention vs. convolution), scale, and training data — not the core mechanism.
Referenced by
- Activation Function
- Autoencoder
- Backpropagation
- Batch Normalization
- Computer Vision
- Convolutional Neural Network (CNN)
- Cosine Similarity
- Embeddings
- Epoch, Batch, and Iteration
- Explainable AI (XAI)
- GAN (Generative Adversarial Network)
- Learning Rate
- Loss Function
- LSTM (Long Short-Term Memory)
- Machine Learning and Deep Learning Terms MOC
- Object Detection
- Overfitting vs Underfitting
- Recurrent Neural Network (RNN)
- Regularization (L1, L2, Dropout)
- Reinforcement Learning
- Supervised Learning
- Transfer Learning
- Transformer Architecture
- Turing Test