Few decisions in neural network design have attracted as much revision as the choice of activation function. It is where the comparison with biology is most literal, deciding what it means for an artificial neuron to fire. The early answer favored biological resemblance: Warren Mc Culloch and Walter Pitts proposed the first formal neuron model in 1943, and Frank Rosenblatt's perceptron in 1957 added adjustable weights. The sigmoid function, with its smooth, saturating curve, was taken to mirror graded neural firing, and biological plausibility served as both a design principle and a source of prestige.
In the early 2010s, the Rectified Linear Unit (ReLU) displaced the sigmoid because it made deep networks trainable where the sigmoid had not. The sigmoid's derivative causes gradients to vanish through backpropagation, as the product of small derivatives shrinks with each layer. ReLU solved this by avoiding saturation, enabling effective training of deeper networks.
The ReLU revolution revealed that the activation function was never a fixed biological commitment but a working hypothesis revised under empirical pressure. The shift from biological inspiration to pragmatic performance calls into question how much weight biology ever carried in neural network design.
Source: Towards Data Science · Summarized by HeadlinesBriefing