Speaker
Dr. Jeff Humphreys, Kummer Endowed Professor of Data Science, Missouri Science and Technology
Title
Mathematics Seminar
Subtitle
Title: Data Science as Generalized Rate-Distortion Theory
Physical Location
Allen Hall 14
Abstract: A useful way to think about machine learning is as lossy compression. A model replaces a complicated data set by a simpler representation while preserving the information needed for a specified task. Rate-distortion theory formalizes exactly this tradeoff between information retained and fidelity achieved.
In this talk, I will argue that this viewpoint extends beyond machine learning. Statistical inference, model fitting, dimensionality reduction, causal inference, and other parts of data science can be placed within a common variational framework. We first choose a structured model class, impose observational or interventional evidence, specify the allowable distortion for the task, and then select the least-informative model compatible with those constraints. Mathematically, this leads naturally to constrained minimization of relative entropy.
Within this framework, ordinary statistical inference appears as the zero-distortion case, while machine learning allows positive distortion in exchange for a simpler representation. Structural assumptions, including parametric models, graphical structure, symmetry, and physical laws, enter as information about the admissible model class; observations and interventions supply additional constraints.
The resulting perspective suggests a broad organizing principle: data science is the problem of preserving exactly the information required by structure, evidence, and task fidelity, while discarding everything else.