You are on page 1of 31

Mas S Mohktar

Email: mas_dayana@um.edu.my
Phone (office): 0379677681
 Sampling distributions
 The Central Limit Theorem
 Confidence Intervals
Basic statistical concepts:
1. Population is a set of measurements or items of interest
 A characteristic of the population is called a parameter
2. Sample is any subset from the population of interest
 A characteristic of the sample is called a statistic
 That is, we are interested in a particular characteristic of the
population (a parameter)
 To get an idea about the parameter, we select a (random)
sample and observe a related characteristic in the sample (a
statistic)
 Then, based on assumption of the behavior of this statistic we
make guesses about the related population parameter
 This is called inference, since we infer something about the
population
 Statistical inference is performed in two ways: Testing of
hypotheses and estimation
 Suppose that our focus is on estimating the mean value of some
continuous r.v. of interest
 The obvious approach would be to use the mean of the sample
as an estimate of the unknown population mean μ
 The quantity X is called an estimator of the parameter μ
 Suppose now that one obtains repeated samples from the
population and measures their means𝑥1 , 𝑥2 ,…, 𝑥𝑛
 If you consider each sample mean as a single observation, from
a r.v. 𝑋 , then it will have a probability distribution
 This is known as the sampling distribution of means of samples
of size n
 Given a underlying population with mean μ and standard
deviation σ, the distribution of sample means for sample of size
n has three important properties:
1. The mean of the sampling distribution is identical to the
population mean μ
2. The standard deviation of the sampling distribution is equal to
𝜎
which is known as the standard error of the mean
𝑛
3. Provided that n is large enough, the shape of the sampling
distribution is approximately normal
Cholesterol level in KL males 20-74 years old
 The serum cholesterol levels for all 20-74 year-old KL males
has mean μ = 211 mg/100 ml and the standard deviation is = 46
mg/100 ml
 That is, each individual serum cholesterol level is distributed
around μ = 211 mg/100 ml, with variability expressed by the
standard deviation
 If we select repeated samples of size 25 from the population,
what proportion of the samples will have a mean value of 230
mg/100 ml or above?
 If the sample size 25 is large enough, the central limit theorem
states that the distribution of sample means of size 25 is
approximately normal with mean μ =211 mg/100 ml, and the
standard error of the mean is
𝜎 46
= = 9.2 mg/100
𝑛 25
 If 𝑥= 230, then we have

 From the table, we find the area to the right of z = 2.07 is 0.019
 Therefore, only about 1.9% of the samples will have a mean
greater than 230 mg/100 ml
 What if?
 To calculate the upper and lower cutoff points enclosing the
middle 95% of the means of samples of size n = 25 drawn from
this population we work as follows:
 The cutoff points in the standard normal distribution are −1.96
and +1.96
 We can translate this to a statement about serum cholesterol
levels
= 193.0 ≤ 𝑥25 ≤ 229.0

 Approximately 95% of the sample means will fall between 193


and 229 mg/100 ml
 The goal is to describe or estimate some characteristic of a
continuous r.v. using the information contained in a sample of
observations
 Two methods of estimation are commonly used:
 point estimation and
 interval estimation

 Point estimation involves using the sample data to calculate a


single number to estimate the parameter of interest, e.g. using
the sample mean 𝑥 to estimate the population means μ
 However, two different samples are very likely to result in
different sample means, and thus there is some degree of
uncertainty involved
 Besides, a point estimate does not provide any information
about the variability of the estimator
 Consequently, interval estimation is often preferred
 This technique provides a range of reasonable values that are
intended to contain the parameter of interest with a certain
degree of confidence
 This range of values is called a confidence interval
 Given a r.v. X that has mean μ and standard deviation σ, the C.L.T.
states that

1) Z ∼ N(0, 1) if X is normally distributed; or


2) Z ∼a N(0, 1) if X is not normally distributed but n is sufficiently
large
 For a standard normal r.v. Z is found between -1.96 and 1.96
about 95% of the time, i.e.

 Equivalently, we have

 or
 This means that we are 95% confident that the interval

 will cover μ
 However, it is important to realize that the estimator 𝑋 is a r.v.,
where the parameter μ is a constant (maybe unknown)
 Therefore, the interval is random and has a 95% chance of
covering μ before a sample is selected
 Once a sample has been drawn and the CI has been calculated,
either μ is within the interval or not
 There is no longer any probability involved
 We do not say that there is a 95% probability that μ lies in
this single CI
 If 95% is not an acceptably high confidence, we may elect to
construct a 99% confidence interval
 Similarly to the last case, this interval will be of the form

 where z0.995 = 2.58, and consequently the 99% two-sided


confidence interval of the population mean is
 All else being equal therefore, higher confidence (say 99%
versus 95%) gets translated to a wider confidence interval
 This is intuitive, since the more certain we want to be that the
interval covers the unknown population mean, the more values
(i.e., wider interval) we must allow this unknown quantity to
take
 In general, all things being equal:
 Larger variability (larger standard deviation) is associated with
wider confidence intervals
 Larger sample size n results in narrower confidence intervals
 Higher confidence results in wider confidence intervals
 In estimation we obviously want the narrowest confidence
intervals for the highest confidence (i.e., wide confidence
intervals are to be avoided)
 Although a 95% CI is used most often in practice, we are not
restricted to this choice only
 The smaller the range of values we consider (say, a 90% CI),
the less confident we are that the interval covers μ
 If we wish to make an interval tighter without reducing the level
of confidence, we can select a larger sample (i.e. by increasing
n)
 A generic CI for μ can be constructed. The 100(1-α)% CI for μ is
 Suppose that we draw a sample of size 12 from the population
of hypertensive smokers and that these men have a serum
cholesterol level of 𝒙= 217 mg/100 ml. The population
standard deviation σ = 46 mg/100 ml.
 Calculate a 99% CI for μ
 Using these 12 hypertensive smokers, we find the limits to be
(183, 251)
 We are 99% confident that (183,251) covers the true mean serum
cholesterol level of the population
 The length of the 99% CI is 251-183 = 68 mg/100 ml
 How large a sample would we need to reduce the length of the
above 99% CI to only 20 mg/100 ml?
 Since the interval is symmetric about the sample mean 𝑥= 217
mg/100 ml, we are looking for the sample size necessary to
produced the interval
(217 − 10, 217 + 10) = (207, 227)
 As the 99% CI is of the form, we just need to find the required n
such that

 Therefore, we have n=140.8


 We can say that we need a sample of 141 men to reduce the
length of the 99% CI to 20 mg/100 ml
 In some situations, we are interested in either an upper limit for
the population mean or a lower limit of it, but not both
 With the same notations used in the previous section, an upper
100(1-α)% confidence bound for μ is

 and a lower 100(1-α)% confidence bound for μ is


 Suppose that we select a sample of 74
children who have been exposed to high
levels of lead from a population with an
unknown mean μ and standard deviation σ =
0.85 g/100 ml
 These 74 children have a mean hemoglobin
level of 𝒙=10.6 g/100 ml
 Based on this sample, a one-sided 95% CI for
μ - the upper bound only is?

You might also like