Data Visualization

16 posts

datadog3 min readCurated summary

Robust statistical distances for machine learning

Statistical distances provide quantitative ways to measure how similar or different data distributions are, complementing visual tools such as histograms and Q-Q plots. The post compares the Kolmogorov-Smirnov, Earth Mover’s, and Cramér-von Mises distances, showing that each responds differently to local changes, long tails, and shifts in distribution. No single metric is universally best; the appropriate choice depends on which differences matter most. ## Visual Inspection and Q-Q Plots - Histograms offer a quick comparison of: - Minimum and maximum values - Center or average - Spread and overall shape - Q-Q plots sort both datasets and plot corresponding values against one another. - Points close to the first bisector indicate that the datasets likely come from similarly shaped distributions. - Visual methods are useful heuristics but do not provide a precise quantitative distance. ## Kolmogorov-Smirnov Distance - The KS distance compares empirical cumulative distribution functions (CDFs). - It is the largest absolute difference between the two CDFs at any point. - It is a true metric: - It is nonnegative. - It is zero only for identical distributions. - It is symmetric. - It satisfies the triangle inequality. - The distance is bounded between 0 and 1. - This makes it convenient for determining whether distributions are similar, but less useful for measuring how far apart very different distributions are. - For normal distributions with equal standard deviations and increasingly separated means, KS quickly levels off rather than growing proportionally. ## Earth Mover’s Distance - Earth Mover’s Distance (EMD), or the first Wasserstein distance, measures the minimum work required to transform one distribution into another. - “Work” is the amount of mass moved multiplied by the distance it travels. - EMD is equivalent to the area between the two empirical CDFs. - It is particularly useful for distributions with long or significant tails because distant mass contributes substantially to the result. - Unlike KS, EMD is unbounded and can grow with the physical separation between distributions. ## Cramér-von Mises Distance - The Cramér-von Mises (CM) distance sums the squared differences between empirical CDFs and takes the square root. - Its relationship to EMD resembles the relationship between L2 and L1 norms. - CM is less sensitive to isolated local changes than KS but more sensitive than EMD. - For normal distributions with separated means: - EMD grows linearly with the separation. - KS rapidly reaches a plateau. - CM grows approximately like the square root of the separation. ## Sensitivity to Distribution Changes - KS focuses on the maximum CDF difference, making it highly sensitive to local deformations. - EMD averages absolute CDF differences, so localized changes may have little effect. - CM provides a compromise, detecting local changes without reacting as strongly as KS. - In examples involving shifted or localized probability mass: - KS can increase dramatically from a small local deformation. - EMD may barely change. - CM usually shows a moderate increase. - Conversely, changes in distant high-percentile regions may have little effect on KS while producing much larger changes in EMD and CM. Choose the distance according to the type of distributional difference you need to detect: KS for maximum local deviation, EMD for overall displacement and long tails, and CM for a balanced sensitivity to both global and local changes.

Read original(opens in new tab)