Back to timeline

Research · April 2012

Apache Spark

Zaharia and colleagues proposed keeping intermediate data in the cluster's memory, and repeated computations sped up tenfold and more.

Why it matters

Machine learning on large data became feasible: algorithms that pass over the data many times stopped being bound by the disk.

MapReduce wrote every step's result to disk, which suits a single pass and ruins iterative methods, meaning most learning algorithms. Spark's resilient distributed datasets live in memory and are rebuilt after a failure by recomputing from their lineage. That made a cluster usable for training models rather than only for counting.

Event record

Event date
April 2012
Timeline date
Event date
Verification
Sources gathered automatically · September 17, 2026
Lines
ID
evt-0256

Presented at the NSDI conference in April 2012.

Sources

Related events