You are on page 1of 57

Data Science for Business

Lecture 1 - Overview of Data Science for Business and Finance1

Vinh Vo (Ph.D)

Department of Economic Mathemetics


Ho Chi Minh city University of Banking

March 31, 2023

1
Materials used: “Fundamentals of Machine Learning for Predictive Data Analytics”
by J.D. Kelleher, B. Mac Namee and A. D’Arcy (MIT, 2015)
Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 1 / 50
Outline

1 What is Predictive Data Analytics?

2 What is Data Mining and Machine Learning?

3 Machine Learning for Data Analytics

4 How Does Machine Learning Work?

5 What Can Go Wrong With ML?

6 The Predictive Data Analytics Project Lifecycle: CRISP-DM

7 Summary

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 2 / 50


What is Predictive Data Analytics?

Outline

1 What is Predictive Data Analytics?

2 What is Data Mining and Machine Learning?

3 Machine Learning for Data Analytics

4 How Does Machine Learning Work?

5 What Can Go Wrong With ML?

6 The Predictive Data Analytics Project Lifecycle: CRISP-DM

7 Summary

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 3 / 50


What is Predictive Data Analytics?

What is Predictive Data Analytics?

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 4 / 50


What is Predictive Data Analytics?

What is (and isn’t) analytics?


From The Institute for Operations Research and the Management Sciences
(INFORMS):

INFORMS defines analytics as the scientific process of transforming data


into insights for the purpose of making better decisions. Analytics is
always an action-driven approach. There is always a decision to be made
when we look at doing analytics. Coming from a data science background
and working with a lot of statisticians, data scientists love to analyze
data just for the sake of analyzing it. However, it is important to
ensure our analysis is driving business action. Ultimately, we want
analytics to empower an organization’s vision.

Predictive Data Analytics encompasses the business and data processes


and computational models that enable a business to make data-driven
decisions.

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 5 / 50


What is Predictive Data Analytics?

Figure: Predictive data analytics moving from data to insights to decisions.

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 6 / 50


What is Predictive Data Analytics?

Example Applications:

❖ Price Prediction:
➢ Businesses such as hotel chains, airlines, and online retailers need to
constantly adjust their prices to maximize returns based on factors such
as seasonal changes, shifting customer demand, and the occurrence of
special events. Predictive analytics models can be trained to predict
optimal prices based on historical sales records. Businesses can then
use these predictions as an input into their pricing strategy decisions.
❖ Dosage Prediction:
➢ Doctors and scientists frequently decide how much of a medicine or
other chemical to include in a treatment. Predictive analytics models
can be used to assist this decision making by predicting optimal
dosages based on data about past dosages and associated outcomes.

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 7 / 50


What is Predictive Data Analytics?

Example Applications:

❖ Risk Assessment:
➢ Risk is one of the key influencers in almost every decision an
organization makes. Predictive analytics models can be used to predict
the risk associated with decisions such as issuing a loan or underwriting
an insurance policy. These models are trained using historical data
from which they extract the key indicators of risk. Output from risk
prediction models can be used by organizations to make better risk
judgements.

❖ Propensity Modeling:
➢ Most business decision making would be made much easier if we could
predict the likelihood, or propensity, of individual customers to take
different actions. Predictive analytics models can be used to build
models that predict future customer actions based on historical
behavior.

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 8 / 50


What is Predictive Data Analytics?

Example Applications:

❖ Diagnosis:
➢ Doctors, engineers, and scientists regularly make diagnoses as part of
their work. Typically, these diagnoses are based on their extensive
training, expertise, and experience. Predictive analytics models can
help professionals make better diagnoses by leveraging large collections
of historical examples at a scale beyond anything one individual would
see over his or her career. These diagnoses made by predictive analytics
models usually become an input into the professional’s existing
diagnosis process.

❖ Document Classification:
➢ Predictive analytics models can be used to automatically classify
documents into different categories. Examples include email spam
filtering, new sentiment analysis, customer complaint redirection, and
medical decision making. The definition of a document can be
expanded to include images, sounds, and videos, all of which can be
classified using predictive data analytics models.

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 9 / 50


What is DM & ML?

Outline

1 What is Predictive Data Analytics?

2 What is Data Mining and Machine Learning?

3 Machine Learning for Data Analytics

4 How Does Machine Learning Work?

5 What Can Go Wrong With ML?

6 The Predictive Data Analytics Project Lifecycle: CRISP-DM

7 Summary

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 10 / 50


What is DM & ML?

What is Data Mining and Machine


Learning?

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 11 / 50


What is DM & ML?

Data Mining
“Data-driven discovery of models and patterns from massive
observational data sets”.

Machine Learning
“(Supervised) Machine Learning techniques automatically learn a model
of the relationship between a set of descriptive features and a target
feature from a set of historical examples”.

Data Science
“Data Science is a set of fundamental principles that support and guide
the principled extraction of information and knowledge from data”.

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 12 / 50


What is DM & ML?

Figure: Data science tasks [Kotu & Deshpande, 2019]

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 13 / 50


ML for Data Analytics

Outline

1 What is Predictive Data Analytics?

2 What is Data Mining and Machine Learning?

3 Machine Learning for Data Analytics

4 How Does Machine Learning Work?

5 What Can Go Wrong With ML?

6 The Predictive Data Analytics Project Lifecycle: CRISP-DM

7 Summary

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 14 / 50


ML for Data Analytics

Machine Learning for Data Analytics

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 15 / 50


ML for Data Analytics

Figure: Using machine learning to induce a prediction model from a training dataset.

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 16 / 50


ML for Data Analytics

Figure: Using the model to make predictions for new query instances.

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 17 / 50


ML for Data Analytics

Loan-Salary
ID Occupation Age Ratio Outcome
1 industrial 34 2.96 repaid
2 professional 41 4.64 default
3 professional 36 3.22 default
4 professional 41 3.11 default
5 industrial 48 3.80 default
6 industrial 61 2.52 repaid
7 professional 37 1.50 repaid
8 professional 40 1.93 repaid
9 industrial 33 5.25 default
10 industrial 32 4.15 default

❖ What is the relationship between the descriptive features


(Occupation, Age, Loan-Salary Ratio) and the target
feature (Outcome)?
Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 18 / 50
ML for Data Analytics

if Loan-Salary Ratio > 3 then


Outcome=’default’
else
Outcome=’repay’
end if

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 19 / 50


ML for Data Analytics

if Loan-Salary Ratio > 3 then


Outcome=’default’
else
Outcome=’repay’
end if
➢ This is an example of a prediction model.

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 19 / 50


ML for Data Analytics

if Loan-Salary Ratio > 3 then


Outcome=’default’
else
Outcome=’repay’
end if
➢ This is an example of a prediction model.

➢ This is also an example of a consistent prediction model.

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 19 / 50


ML for Data Analytics

if Loan-Salary Ratio > 3 then


Outcome=’default’
else
Outcome=’repay’
end if
➢ This is an example of a prediction model.

➢ This is also an example of a consistent prediction model.

➢ Notice that this model does not use all the features and the feature
that it uses is a derived feature (in this case a ratio): feature design
and feature selection are two important topics in ML.

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 19 / 50


ML for Data Analytics

❖ What is the relationship between the descriptive features and the


target feature (Outcome) in the following dataset?

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 20 / 50


Loan-
Salary
ID Amount Salary Ratio Age Occupation House Type Outcome
1 245,100 66,400 3.69 44 industrial farm stb repaid
2 90,600 75,300 1.2 41 industrial farm stb repaid
3 195,600 52,100 3.75 37 industrial farm ftb default
4 157,800 67,600 2.33 44 industrial apartment ftb repaid
5 150,800 35,800 4.21 39 professional apartment stb default
6 133,000 45,300 2.94 29 industrial farm ftb default
7 193,100 73,200 2.64 38 professional house ftb repaid
8 215,000 77,600 2.77 17 professional farm ftb repaid
9 83,000 62,500 1.33 30 professional house ftb repaid
10 186,100 49,200 3.78 30 industrial house ftb default
11 161,500 53,300 3.03 28 professional apartment stb repaid
12 157,400 63,900 2.46 30 professional farm stb repaid
13 210,000 54,200 3.87 43 professional apartment ftb repaid
14 209,700 53,000 3.96 39 industrial farm ftb default
15 143,200 65,300 2.19 32 industrial apartment ftb default
16 203,000 64,400 3.15 44 industrial farm ftb repaid
17 247,800 63,800 3.88 46 industrial house stb repaid
18 162,700 77,400 2.1 37 professional house ftb repaid
19 213,300 61,100 3.49 21 industrial apartment ftb default
20 284,100 32,300 8.8 51 industrial farm ftb default
21 154,000 48,900 3.15 49 professional house stb repaid
22 112,800 79,700 1.42 41 professional house ftb repaid
23 252,000 59,700 4.22 27 professional house stb default
24 175,200 39,900 4.39 37 professional apartment stb default
25 149,700 58,600 2.55 35 industrial farm stb default
ML for Data Analytics

if Loan-Salary Ratio < 1.5 then


Outcome=’repay’
else if Loan-Salary Ratio > 4 then
Outcome=’default’
else if Age < 40 and Occupation =’industrial’ then
Outcome=’default’
else
Outcome=’repay’
end if

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 22 / 50


ML for Data Analytics

if Loan-Salary Ratio < 1.5 then


Outcome=’repay’
else if Loan-Salary Ratio > 4 then
Outcome=’default’
else if Age < 40 and Occupation =’industrial’ then
Outcome=’default’
else
Outcome=’repay’
end if
➢ The real value of machine learning becomes apparent in situations like
this when we want to build prediction models from large datasets with
multiple features.

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 22 / 50


How Does ML Work?

Outline

1 What is Predictive Data Analytics?

2 What is Data Mining and Machine Learning?

3 Machine Learning for Data Analytics

4 How Does Machine Learning Work?

5 What Can Go Wrong With ML?

6 The Predictive Data Analytics Project Lifecycle: CRISP-DM

7 Summary

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 23 / 50


How Does ML Work?

How Does Machine Learning Work?

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 24 / 50


How Does ML Work?

❖ Machine learning algorithms work by searching through a set of


possible prediction models for the model that best captures the
relationship between the descriptive features and the target feature.

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 25 / 50


How Does ML Work?

❖ Machine learning algorithms work by searching through a set of


possible prediction models for the model that best captures the
relationship between the descriptive features and the target feature.

❖ An obvious search criteria to drive this search is to look for models


that are consistent with the data.

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 25 / 50


How Does ML Work?

❖ Machine learning algorithms work by searching through a set of


possible prediction models for the model that best captures the
relationship between the descriptive features and the target feature.

❖ An obvious search criteria to drive this search is to look for models


that are consistent with the data.

❖ However, because a training dataset is only a sample, ML is an


ill-posed problem.

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 25 / 50


How Does ML Work?

Table: A simple retail dataset

ID Baby Alcohol Organic Grp


1 no no no couple
2 yes no yes family
3 yes yes no family
4 no no yes couple
5 no yes yes single

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 26 / 50


How Does ML Work?

Table: A full set of potential prediction models before any training data becomes available.

Bby Alc Org Grp M1 M2 M3 M4 M5 ... M6 561


no no no ? couple couple single couple couple couple
no no yes ? single couple single couple couple single
no yes no ? family family single single single family
no yes yes ? single single single single single couple
...
yes no no ? couple couple family family family family
yes no yes ? couple family family family family couple
yes yes no ? single family family family family single
yes yes yes ? single single family family couple family

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 27 / 50


How Does ML Work?

Table: A sample of the models that are consistent with the training data

Bby Alc Org Grp M1 M2 M3 M4 M5 ... M6 561


no no no couple couple couple single couple couple couple
no no yes couple single couple single couple couple single
no yes no ? family family single single single family
no yes yes single single single single single single couple
...
yes no no ? couple couple family family family family
yes no yes family couple family family family family couple
yes yes no family single family family family family single
yes yes yes ? single single family family couple family

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 28 / 50


How Does ML Work?

Table: A sample of the models that are consistent with the training data

Bby Alc Org Grp M1 M2 M3 M4 M5 ... M6 561


no no no couple couple couple single couple couple couple
no no yes couple single couple single couple couple single
no yes no ? family family single single single family
no yes yes single single single single single single couple
...
yes no no ? couple couple family family family family
yes no yes family couple family family family family couple
yes yes no family single family family family family single
yes yes yes ? single single family family couple family

➢ Notice that there is more than one candidate model left! It is because
a single consistent model cannot be found based on a sample training
dataset that ML is ill-posed.

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 28 / 50


How Does ML Work?

❖ Consistency ≈ memorizing the dataset.

❖ Consistency with noise in the data isn’t desirable.

❖ Goal: a model that generalises beyond the dataset and that isn’t
influenced by the noise in the dataset.

❖ So what criteria should we use for choosing between models?

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 29 / 50


How Does ML Work?

❖ Inductive bias the set of assumptions that define the model selection
criteria of an ML algorithm.
❖ There are two types of bias that we can use:
1 restriction bias
2 preference bias

❖ Inductive bias is necessary for learning (beyond the dataset).

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 30 / 50


How Does ML Work?

How ML works (Summary)

❖ ML algorithms work by searching through sets of potential models.


❖ There are two sources of information that guide this search:
1 the training data,
2 the inductive bias of the algorithm.

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 31 / 50


Underfitting/Overfitting

Outline

1 What is Predictive Data Analytics?

2 What is Data Mining and Machine Learning?

3 Machine Learning for Data Analytics

4 How Does Machine Learning Work?

5 What Can Go Wrong With ML?

6 The Predictive Data Analytics Project Lifecycle: CRISP-DM

7 Summary

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 32 / 50


Underfitting/Overfitting

What Can Go Wrong With ML?

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 33 / 50


Underfitting/Overfitting

❖ No free lunch! – there is no particular inductive bias that on average


is the best one to use.
❖ What happens if we choose the wrong inductive bias:
1 underfitting: the prediction model is too simplistic to represent the
underlying relationship in the dataset between descriptive features and
the target feature.
2 overfitting: the prediction model is so complex that the model fits to
the dataset too closely and becomes sensitive to noise in the data.

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 34 / 50


Underfitting/Overfitting

Table: The age-income dataset.

ID Age Income
1 21 24,000
2 32 48,000
3 62 83,000
4 72 61,000
5 84 52,000

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 35 / 50


Underfitting/Overfitting

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 36 / 50


Underfitting/Overfitting

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 37 / 50


Underfitting/Overfitting

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 38 / 50


Underfitting/Overfitting

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 39 / 50


Underfitting/Overfitting

(a) Dataset (b) Underfitting (c) Overfitting (d) Just right

Figure: Striking a balance between overfitting and underfitting when trying to predict age from
income.

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 40 / 50


Lifecycle

Outline

1 What is Predictive Data Analytics?

2 What is Data Mining and Machine Learning?

3 Machine Learning for Data Analytics

4 How Does Machine Learning Work?

5 What Can Go Wrong With ML?

6 The Predictive Data Analytics Project Lifecycle: CRISP-DM

7 Summary

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 41 / 50


Lifecycle

The Predictive Data Analytics Project


Lifecycle: CRISP-DM

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 42 / 50


Figure: A diagram of the CRISP-DM process which shows the six key phases and indicates the
important relationships between them [Wirth and Hipp, 2000].
Lifecycle

Business Understanding

During this initial phase of any analytics project, the primary goal of the
data analyst is to fully understand the project objectives and requirements
from a business perspective, and then to design a data analytics solution
to achieve the objectives.

Data Understanding

Once the manner in which predictive data analytics will be used to address
a business problem has been decided, it is important that the data analyst
fully understand the different data sources available within an organization
and the different kinds of data contained in these sources.

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 44 / 50


Lifecycle

Data Preparation

Building predictive data analytics models requires specific kinds of data


organized in a specific kind of structure known as an analytics base table
(ABT). This phase of CRISP-DM includes all activities required to convert
the disparate data sources available into a well-formed ABT from which
machine learning models can be induced.

Modeling

The modeling phase of CRISP-DM process is when the machine learning


work occurs. Different machine learning algorithms are used to build a
range of prediction models from which the best model will be selected for
deployment.

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 45 / 50


Lifecycle

Evaluation

Before models can be deployed for use, it is important that they are fully
evaluated and proved to be fit for the purpose. This phase of CRISP-DM
covers all the evaluation tasks required to show that a prediction model
will be able to make accurate predictions after being deployed and that
does not suffer from overfitting or underfitting.

Deployment

Machine learning models are built to serve a purpose within an


organization, and the last phase of CRISP-DM covers all the work that
must be done to successfully integrate a machine learning model into the
processes within the organization.

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 46 / 50


Summary

Outline

1 What is Predictive Data Analytics?

2 What is Data Mining and Machine Learning?

3 Machine Learning for Data Analytics

4 How Does Machine Learning Work?

5 What Can Go Wrong With ML?

6 The Predictive Data Analytics Project Lifecycle: CRISP-DM

7 Summary

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 47 / 50


Summary

Summary

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 48 / 50


Summary

❖ Machine Learning techniques automatically learn the relationship


between a set of descriptive features and a target feature from a
set of historical examples.
❖ Machine Learning is an ill-posed problem:
1 generalize,
2 inductive bias,
3 underfitting,
4 overfitting.

❖ Striking the right balance between model complexity and simplicity


(between underfitting and overfitting) is the hardest part of machine
learning.

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 49 / 50


Summary

B B B B B

B B B B B

B B
Q&A B B

B B B B B

B B B B B

Vinh Vo (EM@HUB) Lecture 0 - Overview of Data Science March 31, 2023 50 / 50

You might also like