HeadlinesBriefing HeadlinesBriefing.com

Autoencoders vs PCA: Linear Method Wins Rigged Test

Towards Data Science •
×

A theoretical advantage that didn't survive contact with a real benchmark. The theoretical case for autoencoders is clean: train the network only on normal data, and it learns to reconstruct normal patterns well. Unlike PCA, which can only capture linear relationships, an autoencoder can learn nonlinear ones — so it should catch anomalies that violate a nonlinear structure. I built two experiments — one easy case, and one deliberately designed to be the autoencoder's best shot — and measured whether the theoretical advantage actually shows up.

Setup: synthetic data, trained on normal samples only, evaluated on a held-out mix of normal and anomalous samples. Three methods compared: an autoencoder (MLPRegressor with a 3-dimensional bottleneck), PCA reconstruction error (also reduced to 3 components), and Isolation Forest. Experiment 1: normal data from Gaussian clusters, anomalies with shifted mean and higher variance. PCA matched or beat the autoencoder — 0.885 F1 for both, catching every anomaly (recall = 1.0). Isolation Forest came in at 0.870.

Experiment 2: rigging the test in the autoencoder's favor. I built anomalies specifically invisible to a linear method: normal data where two features follow a nonlinear relationship (y = sin(3x) + noise), and anomalies that keep the same individual range but violate the relationship between them. This is close to a best-case scenario for the autoencoder.

Source: Towards Data Science · Summarized by HeadlinesBriefing