Remove 2009 Remove Data Science Remove Knowledge Discovery Remove Measurement
article thumbnail

ML internals: Synthetic Minority Oversampling (SMOTE) Technique

Domino Data Lab

Working with highly imbalanced data can be problematic in several aspects: Distorted performance metrics — In a highly imbalanced dataset, say a binary dataset with a class ratio of 98:2, an algorithm that always predicts the majority class and completely ignores the minority class will still be 98% correct. References. link] Fisher, R.

article thumbnail

Explaining black-box models using attribute importance, PDPs, and LIME

Domino Data Lab

but it generally relies on measuring the entropy in the change of predictions given a perturbation of a feature. PDPs for the bicycle count prediction model (Molnar, 2009). Courville, Pascal Vincent, Visualizing Higher-Layer Features of a Deep Network, 2009. Conference on Knowledge Discovery and Data Mining, pp.

Modeling 139