← All posts / Research

SLAC's AI Compressor Shrinks Scientific Data 100x Without Losing the Details That Matter

SLAC and Stanford researchers built a neural network that compresses experimental data 10-100x while preserving the fine-grained signals science depends on — and lets scientists decompress only what they need.

SLAC's AI Compressor Shrinks Scientific Data 100x Without Losing the Details That Matter

The next generation of scientific instruments has a problem that has nothing to do with physics: the experiments are about to drown in their own data. Researchers at the Department of Energy’s SLAC National Accelerator Laboratory, working with Stanford, UT Austin, UC Davis, and Carnegie Mellon, have now published an AI-based answer — a neural network method that compresses massive experimental datasets by 10 to 100 times while keeping the subtle, science-rich details that conventional compression throws away. The work appeared in Nature Machine Intelligence on August 24, 2026.

A flood of data with nowhere to go

The driving force behind the project is SLAC’s Linac Coherent Light Source (LCLS), an ultrafast X-ray free-electron laser that takes “snapshots” of atoms and molecules in motion. As LCLS is upgraded, it will eventually fire up to one million X-ray pulses per second — a data rate measured in terabytes per second. No storage system deployed today can simply absorb that stream and sort it out later.

“There is going to be such a flood of data that there’s really no way to handle it in the way we’ve done before,” said Joshua Turner, a lead scientist at SLAC and the Stanford Institute for Materials and Energy Sciences, and principal investigator on the work. “There are many applications in science now where data storage and analysis speed are really important problems, and I think this method is a clever way to solve them.”

Instruments that demand this kind of throughput include the X-ray photon fluctuation spectroscopy (XPFS) instrument, designed to study how particles move inside exotic topological and quantum materials. For these experiments, the scientifically valuable signal isn’t a clean, obvious feature — it’s statistical texture buried in the noise.

Why ordinary compression fails science

Standard file compression is built for photographs, video, and audio, where the human eye or ear tolerates small losses. Science is different. The tiny speckles scattered across X-ray images of molecules carry information about how a material is arranged, how disordered it is, and how it changes over time.

“Those speckles often reflect the underlying arrangement, disorder or dynamics of a material,” said Yuan Ni, a research associate at SLAC and lead author of the paper. “If we lose them, we would lose unique scientific insights, like how a material is structured and how that changes over time.”

A generic compressor optimizes for perceptual quality. When it discards “unimportant” information, it has no way of knowing that a faint speckle pattern is exactly the measurement a physicist spent beamtime collecting.

Wavelets first, neural networks second

The SLAC method’s core idea is to stop treating a dataset as one undifferentiated blob. The pipeline works in two stages:

  1. Separate by scale. A mathematical tool called wavelet analysis decomposes the data into features of different sizes — coarse structure, mid-scale patterns, and fine detail.
  2. Compress each scale separately. A neural network then learns a compact representation of those scale-separated features, so the fine-grained features are never averaged away by generalized compression.

The result is a compressed encoding that preserves detail across the entire scale range, rather than optimizing for a single notion of “quality.”

Just as important is what happens at decode time. Because features are mapped onto a neural network representation at multiple scales, users can select a specific region of interest and recover it at a chosen resolution — without touching the rest of the dataset. “If you are using a conventional compressor, you would need to decompress the entire file, which could take you minutes, hours or days,” said Zhantao Chen, now an assistant professor at the University of Texas at Austin, who developed the method while a SLAC research associate. “This method can decompress only the region of interest rather than the entire dataset, so it’s much more efficient.”

That selective-decoding property is what turns compression from a storage hack into an analysis accelerator: a researcher probing one corner of a massive dataset gets answers in seconds instead of scheduling an overnight job.

10-100x, across domains

“Depending on the underlying data and the desired quality/fidelity, we can typically achieve 10- to 100-fold reductions in file size,” Ni said. The team didn’t limit validation to X-ray science. They tested the method on measurements of molecules and materials from several experimental techniques, on solar magnetic field measurements, and even on ordinary photographs. Across data types, the network adapted, learning which features matter for each kind of measurement.

The researchers are also explicit about scope: the method is designed to work alongside existing data-reduction techniques — such as pre-filtering to keep only interesting events — not to replace them. “Rather than replacing existing compression methods,” Ni noted, “our work provides an additional AI-based approach.”

Training the networks required serious compute in its own right: the team used Perlmutter, the flagship supercomputer at the National Energy Research Scientific Computing Center (NERSC) at Lawrence Berkeley National Laboratory. The project was supported by the DOE Office of Science and SLAC’s Laboratory Directed Research and Development program.

Why this matters beyond the lab

This work is a preview of a constraint every data-intensive field is about to hit. Radio astronomy’s next-generation observatories, fusion diagnostics, particle physics detectors, and large-scale climate simulations all face the same asymmetry: data generation is scaling faster than storage, bandwidth, and analysis capacity. The choices are stark — throw data away at the instrument, or get smarter about what is kept and how it can be recovered later.

SLAC’s approach suggests a third path that is uniquely suited to AI: learn a representation of the data that keeps multi-scale structure intact, is orders of magnitude smaller, and supports random-access decoding. It reframes compression as a learned, queryable encoding of an experiment rather than a shrink-wrapped archive of it.

There’s a second, quieter implication. For years, AI in science has mostly meant analysis — models that find patterns in data humans already collected. This is AI operating one step upstream, deciding how scientific memory itself is written. If the encoding determines what future researchers can recover from an experiment, then the compression model becomes part of the scientific record. Getting it right — with fidelity controls, scale-aware preservation, and selective retrieval — is not an engineering nicety. It is what makes a one-in-a-million speckle pattern, captured in a fraction of a second, still usable by a scientist who hasn’t been born yet.

The flood is coming. SLAC just demonstrated that machine learning can be the levee — and the index.