This action might not be possible to undo. Are you sure you want to continue?
https://www.scribd.com/doc/30278574/QuantitativeMethods
10/28/2015
text
original
These two measures of dispersion can be discussed in the same section because the
standard deviation is the positive square root of the variance. So even if you are asked to
find the standard deviation of a set of data, you will have to find the variance first.
The variance is of great theoretical importance in advanced statistical work, but all you need
to know about it is its definition, how to calculate it, and its relationship to the standard
deviation.
76 Measures of Dispersion
© ABE and RRC
Since the variance is a function of the squares of the observations, the unit in which it is
measured is the square of the unit in which the observations are measured. So its square
root, i.e. the standard deviation, is measured in the same unit as the observations.
Definition and Calculation
We will define the standard deviation first as you are more likely to be asked for this in your
examination than the variance.
The standard deviation is the positive square root of the mean of the squares of the
differences between all the observations and their mean. You will find this definition quite
easy to understand when you see it written in symbols.
Let σ (the small Greek letter "sigma") be the standard deviation. (You will sometimes find "s"
or "sd" used instead.)
The formula used to calculate the standard deviation of a set of data is:
n
x
x
σ
2
)

(
Formula (a)
where:x = values of the observations
x = mean of the observations
n = number of observations.
As the standard deviation = variance
2
σ = variance.
Example 1:
Using the data of Table 4.1, find the standard deviation of the monthly rainfall of Town A. The
data is shown again in the first two columns of Table 5.2.
Measures of Dispersion 77
© ABE and RRC
Month
Rainfall in
Inches
)
x
x
( i
2
i )
x
x
(
Jan
5.4
+0.51 0.2601
Feb
6.8
+1.91 3.6481
Mar
7.2
+2.31 5.3361
Apr
6.5
+1.61 2.5921
May
5.2
+0.31 0.0961
June
4.2
0.69 0.4761
July
2.1
2.79 7.7841
Aug
2.8
2.09 4.3681
Sep
3.9
0.99 0.9801
Oct
4.5
0.39 0.1521
Nov
4.8
0.09 0.0081
Dec
5.3
+0.41 0.1681
25.8692
Table 5.2: Monthly rainfall for Town A
Table 5.4 shows the calculation of 2
i )
x
x
(
using x = 4.89, which we have already
calculated and so can assume to be known, and n = 12.
Then, substituting in the formula gives:
47
.
1
16
.
2
12
8692
.
25
σ
variance = σ2
= (1.47)2
= 2.16 in.
By expanding the expression 2
i )
x
x
(
we can rewrite the formula in the following useful
way:
2
2
x
n
x
σ
Formula (b)
The choice of the formula to use depends on the information that you already have and the
information that you are asked to give. In any calculation you should keep the arithmetic as
simple as possible and the rounding errors as small as possible.
If you already know x and it is a small integer, there is no reason why you should not
use formula (a).
If you already know x but it is not an integer, as in Example 1, then formula (b) is the
best to use.
When you are given the data in the form of a simple or grouped frequency distribution then
the alternative formulae for calculating the standard deviation are:
78 Measures of Dispersion
© ABE and RRC
f
x
x
f
σ
2
)
(
Formula (c)
OR
2
2
f
fx
f
fx
σ
Formula (d)
Formula (c), like formula (a), is derived directly from the definition of σ. Formula (d) is
obtained by using the sigma notation as in formula (b) and this is the one to use in
calculations.
Example 2:
Using the data of Table 4.3, find the standard deviation of the heights of employees.
The first four columns of Table 5.3 are exactly the same as those in Table 4.3. Column five is
added so that we have all the sums required in formula (d). So from the table:
2,478,750
fx
14,070.0;
fx
;
80
f
n
2
Height
(cm)
Midpoint
(x)
Frequency
(f)
fx
fx2
160 – under 165 162.5
7
1,137.5 184,843.75
165 – under 170 167.5
11
1,842.5 308,618.75
170 – under 175 172.5
17
2,932.5 505,856.25
175 – under 180 177.5
20
3,550.0 630,125.00
180 – under 185 182.5
16
2,920.0 532,900.00
185 – under 190 187.5
9
1,687.5 316,406.25
Totals
80 14,070.0 2,478,750.00
Table 5.3: Heights of employees in cm
Then:
2
80
14070
80
2478750
σ
7.24
52.36
σ
Using a Calculator
As you can see, the arithmetic involved in calculating a standard deviation is extensive even
with a small number of classes. Modern calculators have statistics modes that enable you to
input sets of ungrouped and grouped data, and then read off the standard deviation directly.
Measures of Dispersion 79
© ABE and RRC
It is important that you read your calculator's manual carefully and practise calculating
standard deviations with your own calculator prior to the examination.
Interpreting the Standard Deviation
We have seen that the standard deviation is one of the measures used to describe the
variability of a distribution. However, the standard deviation is in general not easy to interpret.
Clearly larger values mean that the set of data is more spread out around the mean, while
smaller values mean that the set of data is less spread out around the mean. If the data set
has an approximately symmetrical and bellshaped distribution (i.e. is approximately normally
distributed), like the frequency curve illustrated in Fig 3.5, then we can say that:
Roughly 68% of the data will lie within ONE standard deviation of the mean.
Roughly 95% of the data will lie within TWO standard deviations of the mean.
More than 99% of the data will lie within THREE standard deviations of the mean.
Furthermore, the standard deviation has an additional use which makes it more important
than the other measures of dispersion. It may be used as a unit to measure the distance
between any two observations. For example, in the distribution of employees' heights, the
distance in the distribution between a height of 160 cm and a height of 189 cm is 29 cm. But
as the standard deviation is 7.24 cm, we can also say that the distance between the two
observations is (29 7.24), or just over 4 standard deviations.
Advantages and Disadvantages of the Standard Deviation
(a) Advantages
(i) It uses all the observations.
(ii) It is closely related to the most commonly used measure of location, i.e. the
mean.
(iii) It is easy to manipulate arithmetically.
(b) Disadvantages
(i) It is rather complicated to define and calculate.
(ii) Its value can be distorted by extreme values.
This action might not be possible to undo. Are you sure you want to continue?