The kernel trick and support vectors
The authors showed that a maximum-margin linear classifier can be applied to nonlinear problems without ever constructing the high-dimensional space explicitly.
Why it matters
A method with proven guarantees appeared that beat networks in practice, and networks were set aside a second time for a decade.
A kernel function computes an inner product in feature space without building the features, so the cost does not depend on that space's dimension. The training problem becomes convex, meaning it has a single optimum, unlike networks where the result depends on initial weights and chance. That reproducibility is what made support vectors the standard of the 1990s.