SPB Git

spb/neural-networks Public

The Complete Taxonomy of Neural Networks — equation-level reference from McCulloch-Pitts (1943) to diffusion transformers, Mamba, KAN and JEPA (2026).

3.8 KB

# The Complete Taxonomy of Neural Networks

A comprehensive, equation-level reference covering every major family of neural network architectures — from the McCulloch–Pitts neuron (1943) to diffusion transformers, Mamba, KAN, and JEPA (2026).

Compiled from a large-scale parallel web research sweep (six independent research passes, each validating equations, authors, dates, and founding papers against primary sources — arXiv, NeurIPS/ICML/ICLR proceedings, Nature/Science, and authoritative references).


# Table of Contents

Part I — Foundations of Artificial Neural Networks Biological vs. artificial neuron · McCulloch–Pitts (1943) · Perceptron (Rosenblatt 1958) & convergence theorem · XOR / Minsky–Papert · ADALINE/MADALINE (delta rule) · Multilayer Perceptron & forward propagation · Universal Approximation Theorems (Cybenko, Hornik, Leshno) · Backpropagation (full derivation) · Activation functions (sigmoid → Mish, with derivatives) · Loss functions (MSE, MAE, Huber, cross-entropy, hinge) · Optimizers (SGD → AdamW, full update equations) · Regularization (L1/L2, dropout, BatchNorm, LayerNorm) · RBF networks

Part II — Convolutional Neural Networks Neocognitron (1980) · Convolution math (stride, padding, dilation, output-size formula) · Pooling · LeNet-5 · AlexNet · VGG · GoogLeNet/Inception · ResNet (residual equation & gradient analysis) · DenseNet · MobileNet (depthwise separable convolution) · EfficientNet (compound scaling) · ConvNeXt · Object detection (R-CNN family, YOLO loss, SSD) · Segmentation (FCN, U-Net, Mask R-CNN) · 1D & 3D CNNs

Part III — Recurrent Networks and Sequence Models Elman/Jordan RNNs · BPTT & vanishing/exploding gradients (eigenvalue analysis) · LSTM (complete gate equations) · GRU · Bidirectional & stacked RNNs · Seq2Seq · Bahdanau & Luong attention · Echo State Networks / reservoir computing / Liquid State Machines · Hopfield networks (classical & modern) · Boltzmann machines, RBMs, contrastive divergence, Deep Belief Networks · Temporal Convolutional Networks · Neural Turing Machines & Differentiable Neural Computers

Part IV — Transformers and Modern Attention Scaled dot-product & multi-head attention · Sinusoidal / RoPE / ALiBi positional encodings · BERT, GPT, T5, LLaMA (RMSNorm, SwiGLU, GQA) · Efficient attention (Sparse, Linformer, Performer, FlashAttention, sliding window, MQA/GQA) · Mixture of Experts (gating & load balancing, Switch, Mixtral) · Vision Transformers (ViT, DeiT, Swin) · Multimodal (CLIP InfoNCE, Flamingo, LLaVA) · State Space Models (S4, Mamba, RWKV, Hyena) · Scaling laws (Kaplan, Chinchilla) · Kolmogorov–Arnold Networks (KAN)

Part V — Generative Models Autoencoders (denoising, sparse, contractive) · VAE (ELBO derivation, reparameterization, β-VAE, VQ-VAE) · GANs (minimax game, DCGAN, cGAN, WGAN/WGAN-GP, StyleGAN 1–3, Pix2Pix, CycleGAN) · Normalizing flows (change of variables, RealNVP, Glow, MAF/IAF) · Diffusion (DDPM full equations, DDIM, score-based SDE, classifier-free guidance, latent diffusion, flow matching) · Autoregressive (PixelCNN, WaveNet) · Energy-based models · DALL·E, Imagen, DiT, Sora, consistency models

Part VI — Specialized and Emerging Architectures Graph Neural Networks (message passing, GCN, GraphSAGE, GAT, GIN, AlphaFold) · Spiking Neural Networks (LIF, Hodgkin–Huxley, Izhikevich, STDP, neuromorphic hardware) · Self-Organizing Maps (Kohonen) · Capsule Networks (dynamic routing) · Neural ODEs (adjoint method) · Physics-Informed Neural Networks · NeRF (volume rendering) · SIREN / implicit representations · Deep RL networks (DQN, REINFORCE, actor-critic, PPO, AlphaZero/MuZero) · Siamese networks & metric learning (contrastive, triplet) · ELM, DEQ, HyperNetworks, NAS/DARTS, binarized networks, Bayesian NNs, Liquid Neural Networks, World Models & JEPA