You are on page 1of 15

# WHAT ARE OUTLIERS?

• A database may contain data objects that do not comply with the general behavior or model of the data. • the outliers may be of particular interest . • Outliers can be caused by measurement or execution error. These data objects are outliers.

Applications: • Fraud detection • Medicine • Public health • Sports statistics • Detecting measurement errors .

OUTLIER DETECTION METHODS • Statistical Distribution-Based Outlier Detection • Distance-Based Outlier Detection • Density-Based Local Outlier Detection • Deviation-Based Outlier Detection .

.Statistical Distribution-Based Outlier Detection • assumes a distribution for the given data set • identifies outliers with respect to the model using a discordancy test • requires knowledge of the data set parameters • knowledge of distribution parameters • expected number of outliers.

How does the discordancy testing work? • This test examines two hypotheses: • working hypothesis • alternative hypothesis .

• Verifies whether oi is <> in relation to F • Assume T is some statistic used as discordancy test • Assume value of the statistic for object oi is vi • Then distribution T is constructed • SP(vi)=Prob(T > vi). n. … .• A working hypothesis. F. where i = 1. H. is a statement that the entire data set of n objects comes from an initial distribution model. that is. • H : oi E F. 2. is evaluated • If SP(vi) is small H is rejected .

is adopted. H. G. . which states that oi comes from another distribution model. • The result is very much dependent on which model F is chosen because oi may be an outlier under one model and a perfectly valid value under another.• An alternative hypothesis.

H : oi E (1-mu)F +muG. n • Mixture alternative distribution G. • Inherent alternative distribution H’ : oi E G. : : : . where i = 1. : : : . • Slippage alternative distribution . where i = 1. 2. 2.• kinds of alternative distributions. n.

of the objects in D lie at a distance greater than dmin from o.dmin)-outlier. if at least a fraction. a DB(pct. pct. in a data set. D. .Distance-Based Outlier Detection • An object. o. is a distancebased (DB) outlier with parameters pct and dmin.11 that is.

algorithms for mining distance-based outliers • Index-based algorithm • Nested-loop algorithm • Cell-based algorithm .

.Density-Based Local Outlier Detection • Distance-based outlier detection is based on global distance distribution • It encounters difficulties to identify outliers if data is not uniformly distributed.

Deviation-Based Outlier Detection • it identifies outliers by examining the main characteristics of objects in a group • two techniques for deviation-based outlier detection • Sequential Exception Technique • OLAP Data Cube Technique .

Sequential Exception Technique .

OLAP Data Cube Technique .