Professional Documents
Culture Documents
Data Mining: BITS Pilani
Data Mining: BITS Pilani
BITS Pilani
Pilani|Dubai|Goa|Hyderabad M5: Association Analysis
BITS Pilani
Pilani|Dubai|Goa|Hyderabad
Source Courtesy: Some of the contents of this PPT are sourced from materials provided by publishers of prescribed books
Association Rule Mining
• Given a set of transactions, find rules that will predict the
occurrence of an item based on the occurrences of other
items in the transaction
Data Mining
Data Mining
• Two-step approach:
1. Frequent Itemset Generation
– Generate all itemsets whose support minsup
2. Rule Generation
– Generate high confidence rules from each frequent itemset, where each rule
is a binary partitioning of a frequent itemset
Data Mining
Data Mining
• But confidence of rules generated from the same itemset has an anti-
monotone property
• e.g., L = {A,B,C,D}:
Data Mining
• Candidate rule is generated by merging two rules that share the same
prefix
in the rule consequent
CD=>AB BD=>AC
• join(CD=>AB,BD=>AC)
would produce the candidate
rule D => ABC
Data Mining
Support
distribution of
a retail data set
Data Mining
Data Mining
Data Mining
Data Mining
Coffee Coffee
Tea 15 5 20
Tea 75 5 80
90 10 100
Data Mining
P (Y | X )
Lift
P (Y )
P( X , Y )
Interest
P ( X ) P (Y )
Data Mining
Coffee Coffee
Tea 15 5 20
Tea 75 5 80
90 10 100
Data Mining
Y Y Y Y
X 10 0 10 X 90 0 90
X 0 90 90 X 0 10 10
10 90 100 90 10 100
0.1 0.9
Lift 10 Lift 1.11
(0.1)(0.1) (0.9)(0.9)
Statistical independence:
If P(X,Y)=P(X)P(Y) => Lift = 1
Data Mining
• Objective measure:
• Rank patterns based on statistics computed from data
• e.g., 21 measures of association (support, confidence,
Laplace, Gini, mutual information, Jaccard, etc).
• Subjective measure:
• Rank patterns according to user’s interpretation
• A pattern is subjectively interesting if it contradicts the
expectation of a user
• A pattern is subjectively interesting if it is actionable
Data Mining
+ - Expected Patterns
- + Unexpected Patterns
Data Mining