You are on page 1of 49

Classic Exercises

Chapter 1:
1.1 During the presidential election of 1996, the news media regularly presented opinion polls that tracked the
fortunes of the three major candidates. One such poll, taken for TIME/CNN, showed the following results:
If the presidential election were held today, for whom would you voteBill Clinton, Bob
Dole, or Ross Perot?
Clinton53%
Dole35%
Perot7%
The results were based on a sample taken October 1011 of 1387 likely votersthose
most likely to cast ballots in November.
a.
b.
c.
d.

If the pollsters were planning to use these results to predict the outcome of the 1996
election, describe the population of interest to them.
Describe the actual population from which the sample was drawn.
What is the difference between registered voters and likely voters? Why is this
important?
Is the sample selected by the pollsters representative of the population described in part a? Explain.

1.2 The first 24 qualifying times for the inaugural race of the NASCAR Winston Cups California
500 are listed in the table:
Pontiac
181.607
181.493
180.991

Chevrolet
183.015 182.209
182.894 181.915
182.755 181.823

Ford
182.927
182.482
182.163

181.018
180.900
180.814

180.655

182.553

181.653
181.410
181.383

180.795
180.764
180.614

a.
b.
c.
d.

181.283

Draw a graph to compare the times of the top 24 qualifiers for each of the three car makes.
Construct a relative frequency histogram to describe the distribution of the 24 top
qualifying times. How do you describe the shape of the distribution?
Construct three stem and leaf plots to compare the qualifying times for the three makes
of race cars.
Write a paragraph summarizing the results of parts a, b, and c.

1.3 Refer to Exercise 1.2. The Minitab stem and leaf output is shown here for the three sets
of race times. Compare the output with the stem and leaf plots constructed in Exercise
1.2. Explain the similarities and differences in the computer and hand-generated plots.
Minitab output for Exercise 1.3
Character Stem-and-Leaf Displays
StemandleafofPontiacN=4
LeafUnit=0.10

StemandleafofFordN=12
LeafUnit=0.10

11806
21809
2181
2181
21814
11816

518067789
(3)181034
41816
318214
11829

StemandleafofChevroletN=8
LeafUnit=0.10
11812
318189
41822
4182578
11830
1.4 During the presidential election of 2000, the news media regularly presented opinion polls
that tracked the fortunes of the four major candidates. One such poll, taken for CNN/TIME,
showed the following results:
Suppose the election for President were being held today, and you had to choose from Al
Gore and Joseph Lieberman, the Democrats; George W. Bush and Dick Cheney, the
Republicans; Ralph Nader and Winona LaDuke, the Green Party candidates; and Pat
Buchanan, the Reform Party candidate. For whom would you vote Gore, Bush, Buchanan,
or Nader?
Bush
Gore
Nader
Buchanan

47%
45%
4%
1%

The results were based on a sample taken October 4 and 5, 2000, of 1244 adult
Americans, including 636 likely voters those most likely to cast ballots in November.
a. If the pollsters were planning to use these results to predict the outcome of the 2000
election, describe the population of interest to them.
b. Describe the actual population from which the sample was drawn.
c. What is the difference between registered voters and likely voters? Why is this
important?
d. Is the sample selected by the pollsters representative of the population described in
part a? Explain.
1.5 For all the speeches given, the commercials aired, the mud splattered in any election, a
hefty percentage of people decide whom they will vote for only at the last minute.
Suppose that a sample of people were asked, When did you finally decide whom to vote
for in the presidential election? These answers were given:
Answer
Percent
Only in the past 3 days
17%
In the past week
8
In the past 2 weeks (after the debates)
18
In early fall (after the convention)
24
Earlier than that
33
a. Would you use a pie chart or a bar chart to graphically describe the data? Why?
b. Draw the chart you chose in part a.
c. What does the chart tell you about the reliability of election polls taken early in the
election campaign?
1.6 The 1960s generation was never as radical as it was portrayed. According to an opinion
poll in The American Enterprise, when a group of 3040-year-olds were asked to describe
their political views in the 1960s and early 1970s, they gave these responses:
Conservative:
Moderate:
Liberal:
31%
Radical:
6%

28%
35%

a. Define the variable that has been measured in this survey.


b. Is the variable qualitative or quantitative?

c.
d.
e.
f.

What do the numbers represent?


Define the sample and the population of interest to the researchers.
Construct a pie chart to describe the data.
What percentage of 3040-year-olds described themselves as either moderate or
conservative during the 1960s and early 1970s?
g. What other types of questions might you want to investigate, based on the results of
part f?

1.7 The American population over the past 50 years has been grouped by nicknames that try
to describe their common traits. According to a recent magazine article, the numbers of
Americans in each of four age categories are as shown in the table:
Generation
Matures (born before 1946) 68.3
Baby boomers (born 1946-1964)
GenXers (born 1965-1976) 44.6
Others (born after 1976)

Number of Americans (in millions)


77.6
72.4

a. What graphical methods could you use to describe the data?


b. Select the method from part a that you think best describes the data.
c. Why are there fewer Americans in the GenX group than in the other three groups?
1.8 The table shows the average ticket prices for
of July 30-August 5, 2001:
Average
Ticket
Show
Price
1. The Producers
$78
2. The Lion King
71
3. Aida
67
4. 42nd Street
68
5. Beauty and the Beast
62
6. The Phantom of the
61
Opera
7. The Music Man
60
8. Riverdance
48
9. The Full Monty
71
50
10. Les Miserables

the top 20 Broadway shows during the week

Show
11. Annie Get Your Gun
12. Chicago
13. Kiss Me Kate
14. Rent
15. Cabaret
16. Contact
17. Proof
18. Fosse
19. Blast!
20. The Tale of the
Allergists Wife

Average
Ticket
Price
$53
53
55
49
64
62
58
53
41
48

a. Draw a stem and leaf plot for the data. Describe the shape of the distribution. Are
there any outliers?
b. Construct a dotplot for the data. Which of the two graphs is more informative?
Explain.
1.9 The average per-site gross weekend revenue (in dollars) for the top 20 movies were
measured for two consecutive weekends in August 2001.
August 3-5,
2001
$21,619
7,800
9,012
3,545
2,670
2,918
2,307
2,200
1,352
3,211

$1,042
1,315
902
1,062
1,196
5,129
758
882
941
14,895

August 10-12,
2001
$14,730
10,621
5,254
8,397
3,911
2,370
2,287
1,640
1,858
1,402

$1,781
941
871
1,130
1,021
974
2,935
3,313
20,202
630

a. Draw a relative frequency histogram to describe the distribution of the average per-site
revenues for these movies. How would you describe the shape of the distribution?
b. Construct two stem and leaf plots to compare the average per-site revenues for the
two weekends.
c. Write a short paragraph summarizing the results of parts a and b.
1.10 Refer to Exercise 1.9 and data set EX0109. The MINITAB stem and leaf output is shown
here for the two weekends. Compare the output with the stem and leaf plots constructed
in Exercise 1.9. Explain the similarities and differences in the computer- and handgenerated plots.
Stem-and-Leaf Display: Aug3-5
Stem-and-leaf of Aug3-5 N
Leaf Unit = 100
4
9
(4)
7
5
5
4
4
3
3

0
1
2
3
4
5
6
7
8
9

= 20

7899
00133
2369
25
1
8
0

HI 148, 216
Stem-and-Leaf Display: Aug10-12
Stem-and-leaf of Aug10-12 N
Leaf Unit = 100
4
10
10
7
5
5
4
4
4

0
1
2
3
4
5
6
7
8

= 20

6899
014678
239
39
2
3

HI 106, 147, 202


1.11 In July 2000, 22.4 million teenagers and young adults worked a substantial number more than in April, when
school is still in session. Many of these young people worked in amusement and theme parks, whose average
number of employees jumps dramatically during the summer months. Here are the most common injuries suffered
on the job by kids under 18:
Most Common Injury
Bruises and contusions
Cuts and lacerations
Fractures
Heat burns
Sprains and strains
a.
b.
c.
d.

Percentage
14%
13%
8%
9%
33%

Are all possible injuries accounted for in the table? Add another category if necessary.
Create a pie chart to describe the data.
Create a bar chart to describe the data.
Rearrange the bars in part c so that the categories are ranked from the largest percentage to the smallest.

1.12 There is an increasing demand among consumers for broadband Internet access, which will increase the speed at
which they can access the Internet. Cable currently leads the broadband market over DSL and up-and-coming

satellite access, and it is predicted that cable will continue to dominate the U.S. residential broadband market over
the next 4 years. 15
Broadband Access in 2005
(U.S. Year end 2005)
Technology
Subscribers
Cable
15.7 Million
DSL
10.5 Million
Satellite
4.5 Million
Fixed Wireless
359,000
Source: The Yankee Group
a. Use a pie chart to describe the predicted broadband market in 2005.
b. Use a bar chart to describe the predicted broadband market in 2005.
c. Which of the two graphical methods is less susceptible to distortion by the presenter?
1.13 Happy in the Air? Of the 6,229 consumer complaints against major U.S. airlines in a recent year, the
distribution by airline is shown in the table.
Airline
Number of Complaints
United
1172
America West
318
Northwest
765
Continental
563
Delta
1231
Source: Time Almanac 2004
a.
b.
c.

Airline
American
U.S. Airways
Alaska
American Eagle
Southwest

Number of Complaints
1212
532
129
71
236

Construct a pie chart to describe the total number of complaints by airline.


Order the airlines from the smallest to the largest number of complaints. Construct a Pareto chart to describe
the data. Which display is more effective?
Is there another variable that you could measure that might help to explain why some airlines have many
more complaints than others? Explain.

1.14 Top 20 Movies The table below shows the weekend gross ticket sales for the top 20 movies during the week of
January 23-25, 2004:

Movie
1. The Butterfly Effect
2. Along Came Polly
3. Win a Date with Tad Hamilton
4. Big Fish
5. Lord of the Rings: The Return of the King
6. Cheaper by the Dozen
7. Cold Mountain
8. Torque
9. Somethings Gotta Give
10. Mystic River
11. Calendar Girls
12. The Last Samurai
13. Monster
14. My Babys Daddy
15. Disneys Teachers Pet
16. Peter Pan
17. Paycheck
18. The Cooler
19. Mona Lisa Smile
20. 21 Grams
Source: Entertainment Weekly
a.
b.

Weekend
Gross
($ millions)
$ 17.1
16.4
7.3
7.1
6.8
6.4
5.0
4.5
4.0
3.4
2.4
2.2
2.1
2.0
1.4
1.2
1.2
0.9
0.9
0.9

Draw a stem and leaf plot for the data. Describe the shape of the distribution. Are there any outliers?
Construct a dotplot for the data. Which of the two graphs is more informative? Explain.

1.15 Princeton Review An advertising flyer for the Princeton Review, a review course designed for high school
students taking the SAT tests, presents the accompanying bar graph to show the average score improvements for
students using various study methods.
a.
b.

What graphical techniques did the Princeton Review use to make their average improvement figures look as
dramatic as possible?
If you were in charge of promoting a review course at your high school, how would you modify the graph to
make the average improvement for students using a school review course look more impressive?

1.16 Election Results The 2000 election was a close race, in which George W. Bush defeated Al Gore, Ralph Nader,
and Pat Buchanan by the closest of margins. The popular vote (in thousands) for George W. Bush in each of the
50 states is listed below:
AL
AK
AZ
AR
CA
CO
CT
DE
FL
GA
a.
b.
c.

941
167
782
473
4567
884
561
137
2913
1420

HI
ID
IL
IN
IA
KS
KY
LA
ME
MD

138
337
2019
1246
634
622
873
928
287
814

MA
MI
MN
MS
MO
MT
NE
NV
NH
NJ

879
1953
1110
573
1190
240
434
302
274
1284

NM
NY
NC
ND
OH
OK
OR
PA
RI
SC

286
2406
1631
175
2351
744
714
2281
131
787

SD
TN
TX
UT
VT
VA
WA
WV
WI
WY

191
1062
3800
515
120
1437
1109
336
1237
148

By just looking at the table, what shape do you think the data distribution for the popular vote by state will
have?
Draw a relative frequency histogram to describe the distribution of the popular vote for President Bush in the
50 states.
Did the histogram in part b confirm your guess in part a? Are there any outliers? How can you explain
them?

1.17 Election Results, continued Refer to Exercise 1.16. Listed here is the percentage of the popular vote received
by President Bush in each of the 50 states:
AL
AK
AZ
AR
CA
CO

57
60
50
52
42
51

HI
ID
IL
IN
IA
KS

38
70
43
58
48
59

MA
MI
MN
MS
MO
MT

33
47
46
57
51
59

NM
NY
NC
ND
OH
OK

47
36
56
61
50
61

SD
TN
TX
UT
VT
VA

61
51
60
67
41
53

CT
DE
FL
GA
a.
b.

39
42
49
56

KY
LA
ME
MD

57
53
44
40

NE
NV
NH
NJ

63
50
48
41

OR
PA
RI
SC

49
47
32
57

WA
WV
WI
WY

45
52
48
70

By just looking at the table, what shape do you think the data distribution for the percentage of the popular
vote by state will have?
Draw a relative frequency histogram to describe the distribution. Describe the shape of the distribution and
look for outliers. Did the graph confirm your answer to part a?

1.18 Election results, continued Refer to Exercises 1.16 and 1.17. The accompanying stem and leaf plots were
generated using MINITAB for the variables named Popular Vote and Percent Vote.
Stem-and-Leaf Display: Popular Vote, Percent Vote
Stem-and-leaf of
Popular Vote N = 50
Leaf Unit = 100
8
15
20
(6)
24
18
14
11
9
8
7
6
4

0
0
0
0
0
1
1
1
1
1
2
2
2
HI

a.
b.
c.

11111111
2222333
44555
667777
888899
0111
222
44
6
9
0
23
4

Stem and leaf of


Percent Vote N = 50
Leaf Unit = 1.0
2
5
12
22
(10)
18
9
3
2

3 23
3 689
4 0112234
4
5677788899
5
0001112233
5 667777899
6 001113
6 7
7 00

29, 38, 45

Describe the shapes of the two distributions. Are there any outliers?
Do the stem and leaf plots resemble the relative frequency histograms constructed in Exercises 1.16 and
1.17?
Explain why the distribution of the popular vote for President Bush by state is skewed while the percentage
of popular votes by state is mound-shaped.

1.19 Fear of Terrorism One week after the terrorist attack on the World Trade Center on September 11, 2001, the
Gallup poll presented a summary of various polls that had asked the question: How worried are you that you or
someone in your family will become a victim of a terrorist attack? The data are shown in the table below:
Date
April 1995
April 1996
July 1996
August 1998
April 2000
September 11, 2001
September 14-15, 2001
a.
b.
c.
d.

Very/Somewhat
Worried (%)
41
35
39
32
24
58
51

Not Too/Not at All


Worried (%)
57
65
61
67
75
40
48

Draw a line chart to describe the percentages that are very/somewhat worried. Use time as the horizontal
axis.
Superimpose another line chart on the one drawn in part b to describe the percentages that are not too/not at
all worried.
Use your line chart to summarize the changes in the polls during the period just after the attack on the World
Trade Center.
The following line chart was presented on the www.gallup.com Web page. How does it differ from the graph
that you drew? What characteristic of the Gallup line chart might cause distortion when the graph is
interpreted?

1.20 Getting Up to Speed Small businesses continue to rely on the telephone as their primary communication tool,
using the Internet primarily as an information source. A research survey conducted by Covad/Sprint/Equation
Research indicates that most small businesses are only beginning to adopt broadband Internet access.
Small Business Internet Access Methods
Dial-up/modem
31.7%
DSL/ADSL broadband
30.3%
Cable broadband
26.1%
T1 or higher
9.8%
Other
2.1%
Source: www.clickz.com
a.
b.

Do the percentages add up to 100%?


Use a pie chart to describe the Internet access methods used by small businesses.

1.21 Internet Hotspots The most visited site on the Internet is Yahoo!, which boasted 111,271 thousand unique
visitors in December 2003. The number of unique visitors at the top 25 sites are shown in the table:
a.
b.
c.

Look at the data. Can you guess the approximate shape of the data distribution?
Use the second applet in Building a Histogram to construct a relative frequency histogram for the data.
What is the shape of the distribution?
Are there any outliers in the set? If so, what sites boast an unusually large number of unique visitors?

Site
Yahoo! Sites
Time Warner Network
MSN-Microsoft Sites
EBay
Google Sites
Terra Lycos
Amazon Sites

Unique
Visitors
(Millions)
111.3
110.5
110.0
69.2
61.5
52.1
45.7

Site
Weather Channel
Real.com Network
Verizon Communications
Wal-Mart
Shopping.com Sites
Symantec
AT&T Properties

Unique
Visitors
(Millions)
23.8
22.3
22.1
21.4
21.3
19.9
17.5

About/Primedia
Excite Network
CNET Networks
Walt Disney Internet
Viacom Online
American Greetings
Source: cyberatlas.internet.com

42.6
25.1
25.1
25.1
24.7
24.4

InfoSpace Network
Monster Property
EA Online
SBC Communications
Sony Online

17.3
17.3
16.8
16.5
16.5

Chapter 2:
2.1 The top ten movies and their profits (in millions of dollars) from the weekend of May
2325, 1997, are presented here:
The Lost World: Jurassic Park
Addicted to Love
The Fifth Element
Austin Powers
Breakdown
a.
b.
c.

$90.2

Fathers Day
Liar Liar
Volcano
2.3
Night Falls on Manhattan
5.4
Anaconda

11.4
8.0
5.6

$4.7
3.0
2.1
1.73

Draw a stem and leaf plot for the data. Are the data skewed?
Calculate the mean profit for these ten movies for the May 2325 weekend. Calculate the median profit.
Which of the two measures in part b best describes the center of the data? Explain.

2.2 Does birth order have any effect on a persons personality? A report on a study by an MIT researcher indicates that
later born children are more likely to challenge the establishment, more open to new ideas, and more accepting of
change. In fact, the number of later born children is increasing. During the Depression years of the 1930s, families
averaged 2.5 children (59% later born), whereas the parents of baby boomers averaged 3 to 4 children (68% later
born). What does the author mean by an average of 2.5 children?
2.3 The Internet has captured the attention and interest of many Americans. In fact, the
number of hours spent online has been steadily increasing over the last few years. It is
projected that the average time spent online in the year 1999 will be 14 hours per person
per year. Suppose that the standard deviation of the distribution of x, the number of hours
per person per year spent online, is 20 hours.
a.
b.
c.

Visualize the distribution of hours per year spent online. Why do you suppose the yearly average for the
entire U.S. population is so low?
Do you think the distribution is relatively mound-shaped, skewed right, or skewed left? Explain.
Within what limits would you expect at least 3/4 of the measurements to lie?

2.4 Refer to Exercise 2.3. You can use the Empirical Rule to see why the distribution of hours per person per year
could not be mound-shaped.
a. Find the value of x that is exactly one standard deviation below the mean.
b. If the distribution is in fact mound-shaped, approximately what percentage of the
measurements should be less than the value of x found in part a?
c. Since the variable being measured is hours per year, is it possible to find any
measurements that are more than one standard deviation below the mean? Explain.
d. Use your answers to parts b and c to explain why the data distribution cannot be
mound-shaped.
2.5 In a review of several textbooks for use in high school biology instruction, the number of chapters and the number
of pages in each of seven textbooks were recorded as shown here:
Textbook

Chapters

Biology: The Study of Life


39
Biology: The Dynamics of Life
Biology Today
54
Biological Science: A Molecular Approach
Biological Science: An Ecological Approach
Health Biology
50

Pages
885

36

779
887

26
24

599
689
849

Modern Biology

53

831

a. Calculate z-scores for the number of chapters in Biology Today and Biological Science:
An Ecological Approach. Does either of the books have an unusually large or small
number of chapters?
b. Construct a box plot for the numbers of chapters in the seven books. Does the box plot
confirm your answer in part a?
c. Use a box plot to describe the distribution of the numbers of pages in the seven books.
2.6 The state of California not only has some of the most polluted metropolitan areas in the country but also has some
of the strictest pollution regulations in the country. As a result of an aggressive campaign by the South Coast Air
Quality Management District (AQMD), the amount of pollution in southern Californias atmosphere is beginning
to decrease. One of its regulations requires employers of 100 or more workers to prepare a ride-sharing plan that
offers incentives for workers to share a ride, telecommute, ride a bike, or take public transportation to work. The
goal is to have 1.5 workers per car during morning rush hours. What does the AQMD mean by this goal
statement?
2.7 American Demographics is a magazine that investigates various aspects of life in the United States using all sorts
of descriptive statistics. These facts were noted in a recent issue:

26% of all U.S. adults between the ages of 18 and 24 own five or more pairs of wearable sneakers.
61% of U.S. households have two or more television sets.

Identify any percentiles that can be determined from this information.


2.8 The top 20 movies and their gross revenues (in millions of dollars) from Labor Day weekend (August 31September 3, 2001) are shown below:
Weeke
nd
Gross
$15.8
11.7
11.0

Movie
Jeepers Creepers
Rush Hour 2
American Pie 2
The Others

a.
b.
c.

9.2
7.6
6.9
6.5

Planet of the Apes


Jurassic Park III
John Carpenters Ghosts of
Mars
The Curse of the Jade
Scorpion
The Deep End
Legally Blonde
Americas Sweethearts
American Outlaws

4.9
3.9

Bubble Boy
Pearl Harbor

10.2

Rat Race
The Princess Diaries
O
Jay and Silent Bob Strike
Back
Summer Catch
Captain Corellis Mandolin

Movie

Weeke
nd
Gross
$3.8
2.3
2.1
2.0
1.8
1.7
1.3
1.3
1.3
1.2

Can you tell by looking at the data whether it is roughly symmetric? Or is it skewed?
Calculate the mean and the median. Use these measures to decide whether or not the data are symmetric or
skewed.
Draw a box plot to describe the data. Explain why the box plot confirms your conclusions in part b.

2.9 The gross revenue (in millions of dollars) for the top 20 movies over the 2001 Labor Day weekend (Exercise 2.41
and data set EX0241) are reproduced below7:
15.8
3.8
a.
b.
c.

11.7
2.3

11.0
2.1

10.2
2.0

9.2
1.8

7.6
1.7

6.9
1.3

6.5
1.3

4.9
1.3

3.9
1.2

Calculate the mean and standard deviation for these 20 observations.


Calculate the z-score for the largest and smallest observations ( x 15.8 and x 1.2 ). Should either movie,
Jeepers Creepers or Pearl Harbor, be considered unusual?
Construct a box plot for the data, or refer to the box plot drawn in Exercise 2.41. Does the box plot confirm
your results in part b? [HINT: Since the z-score and the box plots are two unrelated methods for detecting

outliers, and use different types of statistics, they do not necessarily have to (but usually do) produce the
same results.]
2.10 LCD TVs The cost of televisions exhibits huge variation from $100-200 for a standard TV to $8000-10,000 for
a large plasma screen TV. Consumer Reports gives the prices for the top ten LCD standard definition TVs in the
14- to 20-inch category:
Brand
Sharp LC-20E1U
Sony KLV-15SR1
Panasonic TC-20LA1
Panasonic TC-17LA1
Gateway GTW-L18M103
Panasonic TC-14LA1
Gateway GTW-L17M103
Toshiba 14VL43U
Toshiba 20VL43U
Sharp LC-15E1U
a.
b.
c.

Price
$1200
800
1050
750
700
500
600
670
1200
650

What is the average price of these ten LCD televisions?


What is the median price of these ten LCD televisions?
As a consumer, would you be interested in the average cost of an LCD TV? What other variables would be
important to you?

2.11 Brett Favre The number of passes completed by Bret Favre, quarterback for the Green Bay Packers, was
recorded at each of the 16 regular season and 2 postseason games in the fall of 2003 (ESPN.com):
25
23
22
a.
b.
c.

15
18
23

23
14
22

21
13
12

19
10
26

25
23
15

Draw a stem and leaf plot to describe the data.


Calculate the mean and standard deviation for Brett Favres per-game pass completions.
What proportion of the measurements lie within two standard deviations of the mean?

2.12 Comparing NFL Quarterbacks How does Brett Favre, quarterback for the Green Bay Packers, compare to
Donovan McNabb, quarterback for the Philadelphia Eagles? The table below shows the number of completed
passes for each athlete during the 2003 NFL football season (ESPN.com):
Brett Favre
25
23
22
15
18
23
23
14
22
21
13
12
19
10
26
25
23
15
a.
b.
c.

Donovan McNabb
19 17 18
18 21 15
18 15 27
16 24 23
11 16 21
9 18 10

Calculative five-number summaries for the number of passes completed by both Brett Favre and Donovan
McNabb.
Construct box plots for the two sets of data. Are there any outliers? What do the box plots tell you about the
shapes of the two distributions?
Write a short paragraph comparing the number of pass completions for the two quarterbacks.

2.13 Internet Hotspots The most visited site on the Internet is Yahoo!, which boasted 111,271 thousand unique
visitors in December 2003 (Exercise 1.21 and data set EX0121). The number of unique visitors at the top 25 sites
are shown in the table:

Site
Yahoo! Sites
Time Warner Network
MSN-Microsoft Sites

Unique
Visitors
(Millions)
111.3
110.5
110.0

Site
Weather Channel
Real.com Network
Verizon Communications

Unique
Visitors
(Millions)
23.8
22.3
22.1

EBay
Google Sites
Terra Lycos
Amazon Sites
About/Primedia
Excite Network
CNET Networks
Walt Disney Internet
Viacom Online
American Greetings
Source: cyberatlas.internet.com
a.
b.
c.

69.2
61.5
52.1
45.7
42.6
25.1
25.1
25.1
24.7
24.4

Wal-Mart
Shopping.com Sites
Symantec
AT&T Properties
InfoSpace Network
Monster Property
EA Online
SBC Communications
Sony Online

21.4
21.3
19.9
17.5
17.3
17.3
16.8
16.5
16.5

Can you tell by looking at the data whether it is roughly symmetric? Or is it skewed?
Calculate the mean and the median. Use these measures to decide whether or not the data are symmetric or
skewed.
Draw a box plot to describe the data. Explain why the box plot confirms your conclusions in part b.

2.14 Snapshots Here are a few facts reported as snapshots in USA Today.

Fifty-two percent of Americans believe that the ideal family size is 2 or fewer children.

Seventy percent of Americans reheat leftovers in their microwave at least two times a week.

Fifty percent of Americans typically wait 15 minutes or less to have a prescription filled.
Identify the variable x being measured, and any percentiles you can determine from this information.
Chapter 3:
3.1 A Gallup poll surveyed 499 women and 502 men to assess the importance of family life in the United States. The
results shown in the table are responses to the question, How would you rate the importance of having a good
family life?
Very Important
Women
Men

459
432

Somewhat
Important
35
70

Very Unimportant

No Opinion

0
0

5
0

a. Define the sample and the population of interest to the researchers.


b. Describe the variables that have been measured in this survey. Are the variables
qualitative or quantitative? Are the data univariate or bivariate?
c. What do the entries in the cells represent?
d. Use comparative pie charts to compare the responses for men and women.
e. What other graphical techniques could be used to describe the data? Would any of these techniques be more
informative than the pie charts constructed in part d?
3.2 Does education really make a difference in how much money you will earn? A study summarized in American
Demographics reports the percentage of people classified as marginally rich, comfortably rich, or super rich
according to their educational attainment.
Education
No college
At least an
undergraduate
degree
Postgraduate
study/degree
a.
b.
c.

Marginally Rich
($7099 K)
32.0%
43.3%

Comfortably Rich
($100249 K)
20.5%
59.3%

Super Rich
($250 K or more)
23.4%
60.3%

12.3%

12.9%

16.2%

The percentages in each of the three categories should sum to 100%. What category has been left out? Add
that category to the table.
Create side-by-side pie charts to display the data.
Create a side-by-side comparative bar chart to display the data.

3.3 A study in Psychological Reports explored the relationship of mental health services (as reflected in per-capita
expenditures on psychiatric hospitals, number of beds per 1000 population, and length of stay in days) to the gross
domestic product (GDP) for selected countries. A portion of the results are shown here:

Country
Finland
The
Netherlands
Denmark
USA
Spain
Japan

GDP per
Capita
18,045.00
14,449.80

Expenditur
es per
Capita
83.53
71.38

Beds/1000
Population
2.6
1.7

Length of
Stay
27.2
35.6

19,842.15
17,670.15
7,563.70
19,763.63

64.79
41.21
9.17
54.98

0.9
0.7
0.8
2.8

11.3
12.3
156.8
310.8

a. What variables have been measured in this study?


b. Draw a scatterplot to examine the relationship between per-capita gross domestic
product and expenditures on psychiatric hospitals for these six countries. Describe the
relationship.
c. Draw a scatterplot to examine the relationship between beds per 1000 population and
length of stay for these six countries. Describe the relationship. Are there any
observations that seem unusual?
d. Comment on the values of r given in the Minitab output in light of parts b and c.
Minitab Output for Exercise 3.3
Correlations (Pearson)
CorrelationofBedsandStay=0.477CorrelationofGDPandExpend=0.708
3.4 Investors are becoming more and more concerned about securities fraud, especially involving initial public
offerings (IPOs). During a 6-year period, the number of federal securities-fraud class action suits has continued
to increase:
Year
Suits
a.
b.
c.

1996
110

1997
178

1998
236

1999
205

2000
211

2001
282

Plot the data using a scatterplot. How would you describe the relationship between year and number of class
action suits?
Find the least squares regression line relating the number of class action suits to the year being measured.
If you were to predict the number of class action suits in the year 2002, what problems might arise with your
predictions?

3.5 Is it harder for single parents to make ends meet than it is for two working parents? The monthly expenses for
families with two children in Riverside, San Bernardino, Orange, and Ventura, California, are displayed in the
side-by-side bar chart.

900

WkgParents
1
2

800
700
Expenses

600
500
400
300
200
100
0
1
g
in
us
Ho

e
ar
dc
hil
C
Tr

a.
b.
c.
d.

n
io
tat
or
p
s
an

od
Fo

e
ar
hc
alt
e
H

isc
M

us
eo
an
ell

1 2
s
xe
a
T

What variables have been measured in this study? Are the variables qualitative or quantitative?
Describe the population of interest. Do these data represent a population or a sample drawn from the
population?
What type of graphical presentation has been used? What other type could have been used?
If you wanted to make the increase in the expenses for families with two working parents look as dramatic as
possible, what changes would you make in the graphical presentation?

3.6 LCD TVs, again In Exercise 2.10, Consumer Reports gave the prices for the top ten LCD standard definition
TVs in the 14- to 20-inch category. Does the price of an LCD TV depend on the size of the screen? The table
below shows the ten costs again, along with the screen size in inches.
Brand
Sharp LC-20E1U
Sony KLV-15SR1
Panasonic TC-20LA1
Panasonic TC-17LA1
Gateway GTW-L18M103
Panasonic TC-14LA1
Gateway GTW-L17M103
Toshiba 14VL43U
Toshiba 20VL43U
Sharp LC-15E1U
a.
b.

Price
$1200
800
1050
750
700
500
600
670
1200
650

Size
20
15
20
17
18
14
17
14
20
15

Which of the two variables (price and size) is the independent variable, and which is the dependent variable?
Construct a scatterplot for the data. Does the relationship appear to be linear?

3.7 LCD TVs, continued Refer to Exercise 3.6. Suppose we assume that the relationship between x and y is linear.
a.
b.
c.
d.

Find the correlation coefficient, r. What does this value tell you about the strength and direction of the
relationship between size and price?
What is the equation of the regression line used to predict the price of the TV based on the size of the screen?
The Sony Corporation is introducing a new 18 LCD TV. What would you predict its price to be?
Would it be reasonable to try to predict the price of a 30 LCD TV? Explain.

3.8 Movie Money Does the opening weekend adequately predict the success or failure of a new movie? In a recent
summer, 36 movies were investigated in Entertainment Weekly, and the following variables were recorded.

The movies first weekends gross earnings (in millions)


The movies total gross earnings in the United States (in millions)

a.
b.
c.
d.

How would you describe the relationship between the first weekends gross and the total gross?
Are there any outliers? If so, explain how they do not fit the pattern of the other movies.
Which dot represents the movie with the best opening weekend? Did it also have the highest total gross?
The film Pearl Harbor opened on a 3-day weekend (Memorial Day). Does that help to explain its position in
relation to the other data points?

3.9 Movie Money, continued The data from Exercise 3.8 were entered into a MINITAB worksheet, and the
following output was obtained.
Covariances: 1st Gross, Total Gross
1st Gross
Total Gross
a.
b.
c.
d.

1st Gross
412.528
1232.231

Total Gross
4437.109

Use the MINITAB output or the original data to find the correlation between first weekend and total gross.
Which of the two variables would you classify as the independent variable? The dependent variable?
If the average first weekend gross is 25.66 million dollars and the average total gross is 86.71 million dollars,
find the regression line for predicting total gross as a function of the first weekends gross.
If another film was released and grossed $30 million on the first weekend, what would you predict that its
total gross earnings will be?

3.10 Brett Favre, again The number of passes completed and the total number of passing yards were recorded for
Brett Favre for each of the 16 regular season and 2 postseason games in the fall of 2003:
Week
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15

Completions
25
15
23
21
19
25
23
18
14
13
10
23
22
23
22

Total Yards
248
132
245
179
185
272
268
194
109
92
138
296
210
278
399

16
12
17
26
18
15
Source: ESPN.com
a.
b.
c.
d.
e.

116
319
180

Draw a scatterplot to describe the relationship between number of completions and total passing yards for
Brett Favre.
Describe the plot in part a. Do you see any outliers? Do the rest of the points seem to form a pattern?
Calculate the correlation coefficient, r, between the number of completions and total passing yards.
What is the regression line for predicting total number of passing yards y based on the total number of
completions x?
If Brett Favre had 20 pass completions in his next game, what would you predict his total number of passing
yards to be?

3.11 Happy in the Air? Of the 6,229 consumer complaints against major U.S. airlines in a recent year, the
distribution by airline is shown in the table, along with the system wide passenger boardings (millions of
passengers).
Airline
United
American West
Northwest
Continental
Delta
American
U.S. Airways
Alaska
American Eagle
Southwest
a.
b.
c.
d.

Number of Complaints
1172
318
765
563
1231
1212
532
129
71
236

Passengers (millions)
68.6
19.5
52.8
40.0
89.9
94.1
47.2
14.2
11.8
72.5

Construct a scatterplot for the data.


Describe the form, direction, and strength of the pattern in the scatterplot.
Are there any outliers in the scatterplot? If so, which airline does this outlier represent?
Does the outlier from part c indicate that this airline is doing better or worse than the other airlines with respect
to customer satisfaction?

3.12 CD Players The table below shows the prices of eight portable CD players along with their overall score (on a
scale of 0-100) in a consumer rating survey presented by Consumer Reports.
Brand and Model
Sony D-EJ611
Panasonic SL-SX280
Aiwa XP-V713
Aiwa XP-SP911
Panasonic SL-CT470
Phillips AZ9213
GPX C3948B1
RCA RP-2360FM
a.
b.
c.

Price
$80
50
70
80
100
80
60
65

Overall Score
70
66
60
65
59
60
47
42

Calculate the correlation coefficient r between price and overall score. How would you describe the
relationship between price and overall score?
Use the applet called Correlation and the Scatterplot to plot the eight data points. What is the correlation
coefficient shown on the applet? Compare with the value you calculated in part a.
Describe the pattern that you see in the scatterplot. Are there any clusters or outliers? If so, how would you
explain them?

3.13 SAT Scores Is there a correlation between Math and Verbal SAT test scores? That is, do students who do well on
the Math portion typically do well on the Verbal portion of the test? The data below (score 400) show the
average scores on standardized math and verbal tests for seven high schools in Southern California.
School
Centennial
Norco

Verbal
64
74

Math
84
95

Moreno Valley
Valley View
Ramona
San Bernardino
Canyon Springs
North
a.
b.
c.

27
75
20
38
68
85

45
71
50
27
85
98

Calculate the correlation coefficient r between verbal and math scores. How would you describe the
relationship between verbal and math scores?
Use the applet called Correlation and the Scatterplot to plot the eight data points. What is the correlation
coefficient shown on the applet? Compare with the value you calculated in part a.
Describe the pattern that you see in the scatterplot. Are there any clusters or outliers? If so, how would you
explain them?

Chapter 4:
4.1 The inaugural season of the WNBA in 1997 made professional basketball a reality for
women basketball players in the United States. There are two conferences, each with four
teams, as shown in the table:
Western Conference

Eastern Conference

Phoenix Mercury
New York Liberty
Sacramento Monarchs
Houston Comets
Los Angeles Sparks
Charlotte Sting
Utah Starzz
Cleveland Rockers
Two teams, one from each conference, are randomly selected to play an exhibition game.
a. How many pairs of teams can be chosen?
b. What is the probability that the two teams are Los Angeles and New York?
c. What is the probability that the Western Conference team is from California?
4.2 Voter support for political term limits is strong in many parts of the United States. A poll
conducted by the Field Institute in California showed support for term limits by a 21
margin. The results of this poll of n= 347 registered voters are given in the table:
For (F)
Republican (R)
Democrat (D)
Other (O)
Total

Against (A)
.28
.31
.06
.65

.10
.16
.04
.30

No Opinion
(N)
.02
.03
.00
.05

Total
.40
.50
.10
1.00

If one individual is drawn at random from this group of 347 people, calculate the following
probabilities:
a. P(R)
d. P(F|R)
g. P(D|A)

b. P(F)
e. P(F|D)
h. P(R|F)

c. P(R F)
f. P(F|O)

4.3 Refer to Exercise 4.2. From the conditional probabilities found in that exercise (and any
others you need), does it appear that the 21 ratio is relatively constant for the three
different party affiliations? Explain.
4.4 Do you surf the net? The amount of information available on the Internet has
skyrocketed in the last few years, making access available to more and more segments of
the population. Do you think that the use of the Internet is related to a persons
educational level? The table shows the percentage of people who accessed the net in the
last 30 days, recorded during a specified time period in 1996, along with the percentages
of the U.S. population at various educational levels.

Educational Level

Percent of Population

Not high school graduate


High school graduate
Attended college
College graduate

18%

Accessed Internet

1.4%
34

27
21

3.4
11.9
23.4

a. What is the overall rate of Internet access for the U.S. population?
b. Suppose that a person is randomly selected, and that the person has accessed the
Internet in the last 30 days. What is the probability that the person is a college
graduate?
c. The posterior probabilities for the four groups, such as the one calculated in part b, are determined using the
fact that the person has accessed the Internet in the last 30 days. The prior probabilities for the four
educational groups are given in the second column of the table. How does the information about Internet
access affect the posterior probabilities in comparison to the prior probabilities?
4.5Dell computer owners are very faithful. Despite reporting problems with their current
systems, 91% of Dell owners said they would buy another computer from the company,
based on the service they received. Suppose you randomly select three current Dell
computer users and ask them whether they would buy another Dell computer system.
a. Find the probability distribution for x, the number of Dell users in the sample of three
who would buy another Dell computer.
b. Construct the probability histogram for p(x).
c. What is the probability that exactly one of the three would buy another Dell computer?
d. What is the population mean and standard deviation for the random variable x?
4.6 Americans can be quite suspicious, especially when it comes to government conspiracies.
On the question of whether the U.S. Air Force is withholding proof of the existence of
intelligent life on other planets, the proportions of Americans with varying opinions are
approximately as shown in the table:
Opinion
Very likely
Somewhat likely
Unlikely
Other

Proportion
0.24
0.24
0.40
0.12

Suppose that one person is randomly selected and his or her opinion on this question is
recorded.
a. What are the simple events in the experiment?
b. Are the simple events in part a equally likely? If not, what are the probabilities?
c. What is the probability that the person finds it at least somewhat likely that the Air
Force is withholding information about aliens?
d. What is the probability that the person finds it unlikely that the Air Force is withholding
information?
4.7 The Eyes of the World About a year after the war in Iraq began, Americans worry about how our country is
viewed by people in other nations. When asked In general, how do you think the United States rates in the eyes
of the world? The proportions of Americans with varying opinions are approximately as shown in the following
table:
Opinion
Proportion
Very favorably
.10
Somewhat favorably
.44
Somewhat unfavorably
.34
Very unfavorably
.11
No opinion
.01
Source: Adapted from www.pollingreport.com
Suppose that one person is randomly selected and his or her opinion on this question is recorded.

a.
b.
c.
d.

What are the simple events in the experiment?


Are the simple events in part a equally likely? If not, what are the probabilities?
What is the probability that the person feels our country is rated at least somewhat favorably by people in
other nations?
What is the probability that the person feels the United States is rated very unfavorably by people in other
nations?

4.8 Heaven Do you believe in heaven? In a survey conducted for Time magazine, 81% of the adult Americans
surveyed expressed a belief in heaven, where people live forever with God after they die. Suppose you had
conducted your own telephone survey at the same time. You randomly called people and asked them whether they
believe in heaven. Assume that the percentage given in the Time survey can be taken to approximate the
percentage of all adult Americans who believe in heaven.
a.
b.

Find the probability distribution for x, the number of calls until you find the first person who does not believe
in heaven.
What problems might arise as you randomly call people and ask them to take part in your survey? How
would this affect the reliability of the probabilities calculated in part a?

Chapter 5:
5.1 Voter participation is decreasing in all age groups, but most rapidly among 1824-year-olds. Only one-third of the
eligible voters in this age group voted in the last presidential election compared with 42% a generation ago in
1972. Of ten individuals selected from this age group, four indicated that they voted in the last presidential
election. Exactly five of the ten individuals in this group indicated that they plan to vote in the next presidential
election. Let us assume that voting in the next presidential election is not influenced by whether a person voted in
the last presidential election.
a. What is the probability that exactly four of the five voted in the last presidential
election?
b. What is the probability that at least one of the five voted in the last presidential
election?
c. What is the probability that none of the five voted in the last presidential election?
5.2 It has been reported that approximately 60% of U.S. households have two or more
television sets and that at least half of Americans watch television by themselves.
Suppose that n 15 U.S. households are sampled and x is the number of households that
have two or more television sets.
a.
b.
c.
d.

What is the probability distribution for x?


Find P x 8 .
What is the probability that x exceeds eight?
What is the largest value of c for which P x c .10 ?

5.3 California residents love to complain about smog, freeway traffic, crime, and earthquakes,
and although the quality of life in California may not be what it once was, roughly 75% of
Californians think that it is one of the best places to live or nice, but not outstanding.
Suppose that five randomly selected California residents are interviewed.
a. Find the probability that all five residents think that California is a nice or very nice
place to live.
b. What is the probability that at least one resident does not think that California is a nice
or very nice place to live?
c. What is the probability that exactly one resident does not think that California is a nice
or very nice place to live?
5.4 Pet Peeves Across the board, 22% of car leisure travelers rank traffic and other drivers as their pet peeve while
traveling. Of car leisure travelers in the densely populated Northeast, 33% list this as their pet peeve. A random
sample of n = 8 such travelers in the Northeast were asked to state their pet peeve while traveling. The MINITAB
printout shows the cumulative and individual probabilities.

MINITAB Output for Exercise 5.1


Cumulative Distribution Function
Binomial with n = 8
and p = 0.33
x
P (X x)
0
0.04061
1
0.20061
2
0.47644
3
0.74814
4
0.91543
5
0.98134
6
0.99758
7
0.99986
8
1.00000
a.
b.
c.

Probability Density Function


Binomial with n = 8
and p = 0.033
x
P (X = x)
0
0.040607
1
0.160003
2
0.275826
3
0.271709
4
0.167283
5
0.065915
6
0.016233
7
0.002284
8
0.000141

Use the binomial formula to find the probability that all eight give traffic and other drivers as their pet
peeve.
Confirm the results of part a using the MINITAB printout.
What is the probability that at most seven give traffic and other drivers as their pet peeve?

5.5 College Costs A USA Today snapshot reports that 52% of households with children under 18 have started to save
for college. Of this number, approximately 50% have saved at least $5,000 toward college expense. Suppose that
we randomly select n = 15 families that have begun saving for college, and let x be the number that have saved
$5,000 or more.
a.
b.
c.
d.

What is the probability distribution for x?


What is P(x 8)?
Find the probability that x exceeds 8?
What is the largest value of c for which P(x c) 0.10?

5.6 Vacation Homes As Americans start the 21st century, the number one status symbol is no longer being a top
executive. Approximately 60% of Americans rank owning a vacation home nestled on a beach or near a
mountain resort number one as a status symbol. A sample of n = 400 Americans is randomly selected.
a.
b.
c.
d.

What is the average number in the sample who would rank owning a vacation home number one?
What is the standard deviation of the number in the sample who would rank owning a vacation home number
one?
Within what range would you expect to find the number in the sample who would rank having a vacation
home as the number one status symbol?
If only 200 in a sample of 400 people ranked owning a vacation home as the top status symbol, would you
consider this unusual? Explain. What conclusions might you draw from this sample information?

5.7 Reality TV Reality TV (Survivor, Fear Factor, etc.) is a new phenomenon in television programming, with
contestants escaping to remote locations, taking dares, breaking world records, or racing across the country. Of
those who watch reality TV, 50% say that their favorite reality show involves escaping to remote locations. If 20
reality-TV fans are randomly selected, find the following probabilities:
a.
b.
c.

Exactly 16 say that their favorite reality show involves escaping to remote locations.
From 15 to 18 say that their favorite reality show involves escaping to remote locations.
Five or fewer say that their favorite reality show involves escaping to remote locations. Would this be an
unlikely occurrence?

5.8 Eating Out More American families with two working parents are eating out these days. What do Americans
use as criteria in selecting the places where they will eat? According to USA Today, 56% choose a restaurant
because of its great food, while 22% make their choice because of reasonable prices. Suppose that in a random
sample of n = 5 families who eat out, x is the number who choose a restaurant because of its great food.
a.
b.

c.

What is the probability that all five families say they select a restaurant because of its great food?
What is the probability that exactly three of the five families select a restaurant because of its great food?
What is the probability that at least one family selects a restaurant because of its great food?

5.9 After College What will you do once you graduate? Where will you live? A report in American Demographics
indicates that 60% of all college students plan to move back home after graduation. A group of 12 college
students are randomly selected and asked whether or not they plan to move back home after graduation.
a.
b.
c.

What is the probability that more than six of the students plan to move back home?
What is the probability that fewer than five of the students plan to move back home?
What is the probability that exactly ten of the students plan to move back home?

5.10 After Graduation Most of todays college seniors want to start earning money as soon as they graduate from
college. In fact, only 14% of college seniors say that they are likely to take some time off to travel and relax after
graduation. Suppose that 50 college seniors were randomly selected.
a.
b.
c.

What is the average value of x, the number of college seniors in the group who say they will take some time
off after graduation? What is the standard deviation of x?
Would it be unlikely to find 15 or more in the group who say they will take some time off after graduation?
Use the applet to find the probability of this event.
How many standard deviations from the mean is the value x = 15? Does this confirm your answer in part b?

Chapter 6:
6.1 A Nielsen survey, reported in Automotive News, indicated that one out of four people over age 16 have used the
Internet in the last 6 months. Does this hold true for you and your statistics classmates? Assume that it does and
that your class contains 50 students. What are the approximate probabilities for these events?
a.
b.

More than 20 students have used the Internet in the last 6 months.
Fewer than 15 have used the Internet in the last 6 months.
c. Fewer than 10 have not used the Internet in the last 6 months.
6.2 If you call the Internal Revenue Service with an income tax question, will you get the correct answer? According
to an investigation conducted by the General Accounting Office, the IRS is giving incorrect answers to 39% of
the questions posed by people who call for help with their taxes. Assume 65 taxpayers call the IRS with
questions.
a. What is the probability that more than 30 will receive incorrect advice?
b. What is the probability that fewer than 20 will receive incorrect advice?
6.3 Life Savers What is your favorite Life Saver flavor? The data from a USA Today Snapshot claims that 32% of
Americans prefer cherry LifeSavers.
Does this hold true for you and your statistics classmates? Assume that it does and that your class contains 50
students. What are the approximate probabilities for these events?

Source: USA Today. Copyright 2004. Reprinted with permission.


a.
b.
c.
d.

More than 20 students prefer cherry LifeSavers.


Fewer than 15 students prefer cherry LifeSavers.
Fewer than 28 students do not prefer cherry LifeSavers.
Are you willing to assume that you and your classmates are a representative sample of all Americans when
it comes to this question? How does your answer affect the probabilities in parts a to c?

6.4 Stress A snapshot in USA Today indicates that forty-nine percent of adult Americans list money issues as the top
causes of stress for them in their daily lives. Suppose that 100 adult Americans are randomly selected.
a.
b.

Use the Calculating Binomial Probabilities applet from Chapter 5 to find the exact probability that 60 or
more adults would give money reasons as the top cause of stress.
Use the Normal Approximation to Binomial Probabilities applet to approximate the probability in part a.
Compare your answers.

Chapter 7:
7.1 In 1996 there was a battle in the courts and in the marketplace between Intel and Digital Equipment Corp. about
the technology behind Intels Pentium microprocessing chip. Digital accused Intel of willful infringement on
Digitals patents. Although Digitals Alpha microprocessor chip was the fastest on the market at the time, its
speed fell victim to Intels marketing clout. That same year, Intel shipped 76% of the microprocessor market.
Suppose that a random sample of n= 1000 personal computer (PC) sales is monitored and the type of
be the proportion of personal computers with a Pentium
microprocessor installed is recorded. Let p
microprocessor in the sample.
a.
b.
c.
d.

? How can you approximate the distribution of p


?
What is the distribution of p
What is the probability that the sample proportion of PCs with Pentium chips exceeded 80%?
What is the probability that the proportion of the sample PCs with Pentium micro-processors was between
75% and 80%?
Would a sample proportion of PCs with Pentium microprocessors equal to 70% appear to contradict the
reported figure of 76%?

7.2 Market research surveys depend to a large extent on responses to telephone interviews, door-to-door interviews,
computer-generated telephone surveys, and so on. The techniques used by market researchers require
representative random samples; however, cooperation rates are dropping to the point that the people who are
responding to the surveys may no longer represent the rest of the population. Miami and Fort Lauderdale,
Florida, and Los Angeles have the least cooperative residents, and about one-third of all Americans routinely
refuse to be interviewed. Suppose you took a random sample of 50 people and asked them whether they would be
willing to be interviewed for a market research survey.
a.
b.

, the proportion of people in the sample who refuse to be


Describe the sampling distribution of p
interviewed.
If p =1/3, what is the probability that your sample proportion will lie within .07 of the true value of p?

7.3 The following notice in a recent newspaper refers to one of the many telephone call-in opinion polls conducted
by the major television networks: NBC will conduct a phone vote via 900 numbers (50 cents a call) during its
1:00 P.M. EST baseball telecasts Saturday on whether a designated hitter should be used in both leagues. Bill
Macatee will update the results until 4:00 P.M. from the studio. Many television viewers will interpret the
proportion of telephone calls responding yes to NBCs question as an estimate of the proportion of all baseball
fans who favor using a designated hitter in both leagues. Explain why the respondents do not represent a random
sample of the opinions of all baseball fans. Explain the types of distortions that could creep into a call-in opinion
poll.
7.4 MRIs In a study described in the American Journal of Sports Medicine, Peter D. Franklin and colleagues
reported on the accuracy of using magnetic resonance imaging (MRI) to evaluate ligament sprains and tears on
35 patients. Consecutive patients with acute or chronic knee pain were selected from the clinical practice of one
of the authors and agreed to participate in the study.
a.
b.
c.

Describe the sampling plan used to select study participants.


What chance mechanism was used to select this sample of 35 individuals with knee pain?
Can valid inferences be made using the results of this study? Why or why not?

7.5 Tax Savings An important expectation of a federal income tax reduction is that consumers will reap a
substantial portion of the tax savings. Suppose estimates of the portion of total tax saved, based on a random
sampling of 35 economists, have a mean of 26% and a standard deviation of 12%.
a.
b.

What is the approximate probability that a sample mean, based on a random sample of n = 35 economists,
will lie within 1% of the mean of the population of the estimates of all economists?
Is it necessarily true that the mean of the population of estimates of all economists is equal to the percentage
of tax savings that will actually be achieved? Why?

7.6 Stress and Sweets One of the ways most Americans relieve stress is to reward themselves with sweets.
According to one study in Food Technology, 46% admit to overeating sweet foods when stressed. Suppose that
the 46% figure is correct and that a random sample of n = 100 Americans is selected.
a.
b.
c.
d.

, the sample proportion of Americans who relieve stress by overeating sweet


Does the distribution of p
foods, have an approximately normal distribution? If so, what are its mean and standard deviation?
, exceeds 0.5?
What is the probability that the sample proportion, p
lies within the interval 0.35 to 0.55?
What is the probability that p
What might you conclude if the sample proportion were as small as 30%?

7.7 Surfing the Net Do you use the Internet to gather information for a project? The Press Enterprise reports that
the percentage of students who used the Internet as their major resource for a school project in a recent year was
66%. Suppose that you take a sample of n = 1000 students, and record the number of students who used the
be the proportion of
Internet as their major resource for their school project during the past year. Let p
students surveyed who used the Internet as a major resource in the past year.
a.
b.
c.
d.

? How can you approximate the distribution of p


?
What is the exact distribution of p
exceeds 68%?
What is the probability that the sample proportion p
What is the probability that the sample proportion lies between 64% and 68%?
Would a sample proportion of 70% contradict the reported value of 66%?

7.8 Child Abuse A study of nearly 2000 women included questions dealing with child abuse and its effect on the
womens adult life. The study reported on the likelihood that a woman who was abused as a child would suffer

either physical abuse or physical problems arising from depression, anxiety, low self-esteem, and drug abuse as
an adult.
a.
b.

Is this an observational study or a designed experiment?


What problems might arise because of the sensitive nature of this study? What kinds of biases might occur?

Chapter 8:
8.1 A study of 392 healthy children living in the area of Tours, France, was designed to measure the serum levels of
certain fat-soluble vitaminsnamely, vitamin A, vitamin E, -carotene, and cholesterol. Knowledge of the
reference levels for children living in an industrial country, with normal food availability and feeding habits,
would be important in establishing borderline levels for children living in other, less favorable conditions. Results
of the study are shown in the table.

Retinol (g/dl)
-carotene (g/l)

350

Mean SD (n = 392)

Boys (n= 207)

Girls (n= 185)

42.5 12.0
572 381

43.0 13.0
588 406

41.8 10.7
553

Vitamin E (mg/l)

9.5 2.5

9.6 2.7

9.5

2.2
Cholesterol (g/l)
1.84 .42
Vitamin E/cholesterol (mg/g)
5.26 1.04

1.84 .47
5.26 1.11

1.83 .38
5.26

.25
a.

Describe the sampled population.

b.

Find a point estimate for the mean cholesterol level for the population of boys. Calculate the margin of error.

c.

Find a point estimate for the mean -carotene level for the population of girls. Calculate the margin of error.

d.

Repeat part c for the population of boys.

e.

Display the results of parts c and d graphically, using the form shown in Figure 8.5 (located in the Classic
Exercises folder on the CD-ROM). Use this display to compare the mean -carotene levels for boys and girls.

f.

Can the data collected by the researchers be used to estimate serum levels for all boys and girls in this age
category? Why or why not?

8.2 A survey concerning the health habits of Americans has been taken for over 10 years by Louis Harris &
Associates. Their latest findings have been labeled terribly disturbing by Dr. Todd Davis, the executive vice
president of the American Medical Association. The surveys sample of 1251 Americans showed that only 53% ate
the recommended amounts of fibrous foods and vegetables, 51% avoided fat, and 46% avoided excess salt.
a.
b.

Based on the survey results, construct a 90% confidence interval for the proportion of all Americans who eat
the recommended amounts of fibrous foods and vegetables.
Construct a 99% confidence interval for the proportion of all Americans who avoid fat in their diet.

8.3 How much time do computer users spend on the information superhighway? An Internet server

conducted a survey of 250 of its customers and found that the average amount of time spent online was
10.5 hours per week with a standard deviation of 5.2 hours.
a.
b.

c.
d.

Do you think that the random variable x, the number of hours spent online, has a mound-shaped distribution?
If not, what shape do you expect?
If the distribution of the original measurements is not normal, you can still use the standard normal
distribution to construct a confidence interval for , the average online time for all users of this Internet
server. Why?
Construct a 95% confidence interval for the average online time for all users of this particular Internet server.
If the Internet server claimed that its users averaged 13 hours of use per week, would you agree or disagree?
Explain.

8.4 Americans have varying opinions about the state of the nation, depending not only on the current events in the
country but also on the state of their own personal affairs. A recent survey of 1017 adult Americans was conducted
by Yankelovich Partners Inc. In response to the question, Is the United States going through a period of good
times or bad times right now? 54% of those surveyed answered good times. Construct a 90% confidence
interval for the proportion of all adult Americans who think that the United States is going through a period of
good times.
8.5 One method for solving the electric power shortage involves constructing floating nuclear power plants a few
miles offshore. Because there is concern about the possibility of ships colliding with a floating (but anchored)
plant, an estimate of the density of ship traffic in the area is needed. The number of ships passing within 10 miles
of the proposed power plant location per day, recorded for n= 60 days during July and August, has sample mean
and variance equal to x = 7.2 and s 2 =8.8, respectively.
c.
d.

e.

Find a 95% confidence interval for the mean number of ships passing within 10 miles of the proposed powerplant location during a 1-day period.
The density of ship traffic was expected to decrease during the winter months. A sample of n= 90 daily
recordings of ship sightings for December, January, and February gave a mean and variance of x = 4.7 and

s 2 =4.9. Find a 90% confidence interval for the difference in mean density of ship traffic between the
summer and winter months.
What is the population associated with your estimate in part b? What could be wrong with the sampling
procedure in parts a and b?

8.6 In assessing students leisure interests and stress ratings, Alice Chang and colleagues tested 559 high school
students using the leisure interest checklist (LIC) inventory. The means and standard deviations for each of the
seven LIC factor scales are given in the following table for both boys ( n1 = 252) and girls ( n 2 307):

LIC Factor Scale Activity


Hobbies
Social fun
Sports
Cultural
Trips
Games
Church

Boys
Mean SD
10.0
6.47
6
22.0
5.12
5
13.6
4.82
5
11.4
5.69
8
6.90
3.41
4.95
3.29
2.15
2.10

Girls
Mean SD
13.6
7.46
4
25.9
5.07
6
9.88
4.41
13.2
1
6.49
3.85
3.00

5.31
2.97
2.49
2.26

a. Find a 95% confidence interval for the mean difference between boys and girls for the cultural factor scale
activity. Interpret this interval.
b. Find a 90% confidence interval for the mean difference between boys and girls for the social fun factor scale
activity. Interpret this interval.
c. Do there appear to be significant differences between boys and girls in the two LIC activity scores given in
parts a and b? Explain.
8.7 Refer to Exercise 8.1. The means and standard deviations for the -carotene and serum cholesterol levels (in
micrograms per liter) are given in the table.

-carotene
(g/l)
Cholesterol
(g/l)

Mean SD (n =
392)

Boys (n = 207)

Girls (n = 185)

572 381

588 406

553 350

1.84 .42

1.84 .47

1.83 .38

a. Find a 90% confidence interval for the difference in the mean serum cholesterol levels
of boys versus girls.
b. Find a 95% confidence interval for the difference in the mean -carotene levels of boys versus girls.
c.

Do the intervals in parts a and b contain the value ( 1

2 ) 0 ? Why is this of interest to the researcher?

d. Do the data indicate a difference between boys and girls in average -carotene or serum cholesterol levels?
Explain.
8.8 Age influences not only what you watch on television but also where you watch television. A study has shown
that older Americans are less likely than younger ones to watch television in bed and more likely to watch
television in the dining area. For the data given in the table, assume that the sample size for each age group was
1000.
Area

25 to 44

Living room/family
Room/den
Bedroom
Kitchen
Dining area

45 to 59

60 and
older

95%

95%

93%

58%
12%
10%

57%
20%
10%

45%
20%
19%

a. Find a 95% confidence interval for the difference in the proportions of Americans aged
45 to 59 and those 60 and older who watch television in the dining area.
b. Estimate the difference between the proportions of Americans in the 25 to 59 age
groups and those 60 years and older who watch television in the living room, family
room, or den using a 99% confidence interval. (HINT: The proportion for the 25 to 59 age
groups will be the simple average of the individual proportions based on a sample size
of 2000.)
c. What conclusions can you draw regarding the groups compared in parts a and b?
8.9 Are American students inferior to students from other countries? Test scores in mathematics and science were
recorded for random samples of 100 American and 100 Canadian students in the fourth grade. The results shown
in the table are the sample means plus or minus the standard deviation.

Mathematic
s
Statistics
a.
b.
c.

United
States
545 130

Canada

565 122

549 133

532 126

Construct a 95% confidence interval for the average mathematics score for the American students.
If the average mathematics score for all fourth-graders in Singapore was 625, can you conclude that the
performance of American fourth-graders is significantly lower than that of the Singapore fourth-graders?
Explain.
Construct a 99% confidence interval for the difference in the average science scores for the American and
Canadian students. Can you conclude that the population mean science scores are different for the two
countries? Explain.

8.10 In a controlled pollination study involving Phlox drummondii, a spring-flowering annual plant common along
roadsides and sand fields in central Texas, researchers found that seed survival rates were not affected by water or
nutrient deprivation. In the experiment, flowers on plants were counted as males when donating pollen and as
females when pollinated by donor pollen in crosses involving plants in the three treatment groups: control, low
water, and low nutrient. The following table, which reflects one aspect of the results of the experiment, records the
number of seeds surviving to maturity for each of the three groups for both male parents and female parents.

Treatment

Control
Low water
Low

585
578
568

Male
Number
Surviving
543
522
510

n
632
510
589

Female
Number
Surviving
560
466
546

nutrient
a.
b.
c.
d.
e.

Find the sample estimate of the proportion of surviving seeds and the standard error of the estimate for each
of the six treatmentsex combinations.
Find a 99% confidence interval for the difference in survival proportions for low water versus low nutrients
for the male category.
Find a 99% confidence interval for the difference in survival proportions for low water versus low nutrients
for the female category.
Find a 99% confidence interval for the difference in survival proportions for the male control group
compared to the female control group.
What conclusions can you draw using the confidence intervals in parts b, c, and d?

8.11 Male Teachers Although most school districts do not specifically recruit men to be elementary school teachers,
those men who do choose a career in elementary education are highly valued and find the career very rewarding.
If there were 40 men in a random sample of 250 elementary school teachers, estimate the proportion of male
elementary school teachers in the entire population. Give the margin of error for your estimate.
8.12 Sports Crazy Are you sports crazy? Most Americans love participating in or at least watching a multitude of
sporting events, but many feel that sports have more than just an entertainment value. In a survey of 1000 adults
conducted by KRC Research & Consulting, 78% feel that spectator sports have a positive effect on society.
a.

b.

Find a point estimate for the proportion of American adults who feel that spectator sports have a positive
effect on society. Calculate the margin of error.
The poll reports a margin of error of plus or minus 3.1%. Does this agree with your results in part a? If
not, what value of p produces the margin of error given in the poll?

8.13 Tweens When it comes to advertising, tweens (kids aged 10 to 13) are not ready for the hardline messages
that advertisers often use to reach teenagers. The Geppetto Group study found that 78% of tweens understand
and enjoy ads that are silly in nature. Unlike teenagers, tweens would much rather see dancing (69%) and
boyfriends and girlfriends (63%) than sexy looking people or kissing. Suppose that these results are based
on a sample size n = 1030 tweens.
a.
b.

Construct a 95% confidence interval estimate of the proportion of tweens who understand and enjoy ads that
are silly in nature
Construct a 95% confidence interval for the proportion of tweens who would rather see dancing.

8.14 Movie Ratings Is Americas romance with movies on the wane? In a poll of n = 800 randomly chosen adults,
45% indicated that movies were getting better, while 43% indicated that movies were getting worse. However,
these results vary by age: 69% of the 18-to-29-year-olds say that movies are getting better, 39% of 50-to-64-yearolds and just 15% of those over 60 agree. Construct a 90% confidence interval for the overall proportion of adults
who say that movies are getting better.
8.15 Public Service Jobs Instead of paying to support welfare recipients, many Californians want them to find jobs; if
necessary, they want the state to create public service jobs for those who cannot find jobs in private industry. In a
survey of 500 registered voters reported in The Press-Enterprise 250 Republicans and 250 Democrats 70% of
the Republicans and 86% of the Democrats favored the creation of public service jobs. Use a large-sample
estimation procedure to compare the proportions of Republicans and Democrats in the population of registered
voters in California who favor creating public service jobs. Explain your conclusions.
8.16 Ads in Outer Space? Do you think that we should let Radio Shack film a commercial in outer space? The
commercialism of our space program is a topic of great interest since Dennis Tito paid $20 million to ride along
with the Russians on the space shuttle. In a survey of 500 men and 500 women, 20% of the men and 26% of the
women responded that space should remain commercial-free.
a.
b.
c.

Construct a 98% confidence interval for the difference in the proportions of men and women who think that
space should remain commercial-free.
What does it mean to say that you are 98% confident?
Based on the confidence interval in part a, can you conclude that there is a difference in the proportions of
men and women who think space should remain commercial-free?

8.17 Spectator Sports Exercise 8.12 discussed a research poll conducted for U.S. News & World Report to determine
the publics attitudes concerning the effect of spectator sports on society. Suppose you were designing a poll of
this type.

a.
b.

Explain how you would select your sample. What problems might you encounter in this process?
If you wanted to estimate the percentage of the population who agree with a particular statement in your
survey questionnaire correct to within 1%, with probability 0.95, approximately how many people would
have to be polled?

8.18 Are You Patriotic? In a CNN/USA Today/Gallup poll involving n = 1000 American citizens who were asked
how the term patriotic described them, the following summary was obtained:
Very well
Somewhat well
Not very well
Not well at all
a.
b.
c.

All
53%
31%
10%
6%

18-34
35%
41%
16%
8%

60+
77%
17%
4%
2%

Construct a 95% confidence interval for the proportion of all American citizens who agreed that the term
described them very well.
If the 60+ age group consisted of n = 150 individuals, find a 95% confidence interval for the proportion of
American citizens 60+ years of age who agreed that the term described them somewhat well.
If the 18-34 age group consisted of n = 340 individuals, find a 98% confidence interval for the true difference
in proportions of those aged 60+ and those 18-34 years of age who agreed that the term described them very
well.

8.19 Schools Out! Summer school break adds more than 20 million teenagers to the American workforce, and in fact,
July is the peak time for summer jobs for teenagers. American Demographics reports that amusement and theme
parks average 100 full-time employees, but at peak season average 665 workers per site. A group of n = 40
amusement and theme parks are randomly selected at peak season, and the average number of full-time workers is
found to be 652 with a standard deviation of 32. Find a 99% confidence interval estimate for the true mean
number of full-time workers at amusement and theme parks at peak season. Does this contradict the reported
average of 665 workers per site?
8.20 Case Study How Reliable Is That Poll? In the days and weeks following important or controversial events,
opinion polls saturate the media. In the years 20022004, public attention has been on the war in Iraq and its
consequences to the American public. However, in March 2004 the attention of the nation focused on the fact that
nationwide gas prices were at record highs. A recent CNN/USA Today/Gallup Poll conducted March 2628, 2004
involving n = 1001 adults 18 years or older reported that nearly seven in ten Americans say the cost of gasoline
represents a major crisis/problem for the United States. Do national polls conducted by the Gallup and Harris
organizations, the news media, and so on, really provide accurate estimates of the percentages of people in the
United States who have various opinions? Lets look at some of the results of the poll survey methods that were
used.
Although the level of gas prices is a major concern for the U.S. public, it ranks somewhere below a list of selected
recent events rated as crisis or major problems in Gallup Polls, 20032004 as shown in the next table.
CNN/USA Today/Gallup Poll
Event
Iraqs military capabilities
North Koreas capabilities
Civilian aid for Iraqi civilians
Federal budget deficit
Cost of gasoline

Year
2003
2003
2003
2003
2004

Crisis %
25
23
30
20
13

Major
Problem %
56
58
49
57
56

Crisis/Major
Problem %
81
81
79
77
69

For several questions in the present poll, the sample of 1001 Americans 18 years and older was split into two
groups, one with 510 and the second with 491 adults. The first group answered the following question (1):
Thinking about the cost of gasoline, would you say that the country [rotated: is in a state of crisis, has major
problems, has minor problems, (or) no problems at all]?
Crisis Problem
13%

Major Problem
56%

Minor Problem
26%

No Problem
5%

The second split sample consisting of 491 adults were asked (2) Do you think the current rise in gas prices
represents [rotated: a temporary fluctuation in prices, or a more permanent change in prices]? The latest

survey finds that 55% of Americans believe the recent price increases represent a more permanent change in gas
prices, while 42% say they are only a temporary fluctuation in prices.
Temporary
42%

More Permanent
55%

No Opinion
3%

The following description of the Survey Methods that appears on the Gallup Poll Web site explains how the
results of the poll should be viewed.
The results are based upon telephone interviews with a randomly selected national sample of 1001
adults, aged 18 years or older, conducted March 26-28, 2004. For results based on this sample (the
full sample of 1001 adults 18 years or older), one can say with 95% confidence that the maximum
error attributable to sampling and other random effects is 3 percentage points. In addition to
sampling error, question wording and practical difficulties in conducting surveys can introduce
error or bias into the findings of public opinion polls.
1.
2.
3.
4.
5.
6.

Verify the margin of error of percentage points given by the Gallup Poll for the full sample of 1001 adults.
Find the margin of error for the two split samples of 510 adults answering the first question and of 491 adults
answering the second question.
Do the numbers reported in the tables represent the number of people who fell into those categories? If not,
what do those numbers represent?
When question (1) was asked, the pollster rotated the order of options given to the respondent. Why do you
suppose this technique was used?
Construct 95% confidence intervals for the proportion of Americans who:
a. report that the cost of gasoline in March of 2004 is a major problem.
b. report that it is a crisis problem.
Compare the percentage of people in the first split sample who say that the cost of gasoline is a major
problem with those in the second split sample who say that the prices may be more permanent. Is this a
surprising difference?
If these questions were asked today, would you expect the responses to be similar to those reported here or
would you expect them to differ substantially?

Chapter 9:
9.1 How much time do computer users spend on the information superhighway? An Internet server claimed that its
users averaged 13 hours per week. To determine whether this was an overstatement, a competitor conducted a
survey of 250 customers and found that the average time spent online was 10.5 hours per week with a standard
deviation of 5.2 hours. Does the data provide sufficient evidence to indicate that the average hours of use are less
than that claimed by the first Internet server? Test at the 1% level of significance.
9.2 In comparing high school and college students leisure interests and stress ratings (see Exercise 8.6), Alice Chang
and colleagues tested 559 high school students using the leisure interest checklist (LIC) inventory. The means and
standard deviations for each of the seven LIC factor scales are given in the table for 252 boys and 307 girls.

LIC Factor
Scale Activity
Hobbies
Social Fun
Sports
Cultural
Trips
Games
Church
a.

b.

Mean
10.06
22.05
13.65
11.48
6.90
4.95
2.15

Boys
SD
6.47
5.12
4.82
5.69
3.41
3.29
2.10

Mean
13.64
25.96
9.88
13.21
6.49
3.85
3.00

Girls
SD
7.46
5.07
4.41
5.31
2.97
2.49
2.26

Without looking at the sample data, would you expect that boys would have a greater or lesser interest in
sports as a leisure time activity than girls? Do the data provide sufficient evidence to indicate that the mean
LIC score for sports is greater for boys than the mean for girls? Test using .01.
Do the data provide sufficient evidence to indicate a difference in mean LIC scores for church between boys
and girls? Test using .01.

9.3 In an article entitled A Strategy for Big Bucks, Charles Dickey discusses studies of the habits of white-tailed
deer that indicate that they live and feed within very limited rangesapproximately 150 to 205 acres. To
determine whether there was a difference between the ranges of deer located in two different geographic areas, 40
deer were caught, tagged, and fitted with small radio transmitters. Several months later, the deer were tracked and
identified, and the distance x from the release point was recorded. The mean and standard deviation of the
distances from the release point were as follows:

Sample size
Sample mean
Sample standard deviation
a.
b.
c.

Location 1
40
2980 ft
1140 ft

Location 2
40
3205 ft
963 ft

If you have no preconceived reason for believing one population mean is larger than another, what would you
choose for your alternative hypothesis? Your null hypothesis?
Does your alternative hypothesis in part a imply a one- or a two-tailed test? Explain.
Do the data provide sufficient evidence to indicate that the mean distances differ for the two geographic
locations? Test using .05.

9.4 In a comparison study of homeless and vulnerable meal-program users, Michael Sosin investigated determinants
that account for a transition from having a home (domiciled) but utilizing meal programs to becoming homeless.
The following information provides the study data:

Sample size
Number currently
working
Sample proportion

Homeless Men
112
34

Domiciled Men
260
98

.30

.38

Test for a significant difference in the population proportions who are currently working for homeless versus
domiciled men. Use 0.01.
9.5 Exercise 8.10 presented data from a controlled pollination study involving seed survival rates for plants subjected
to water or nutrient deprivation. The data, representing the number of seeds surviving to maturity, are reproduced
here:

Treatment

Control
Low water
Low
nutrient

585
578
568

a.
b.
c.

Male
Number
Surviving
543
522
510

n
632
510
589

Female
Number
Surviving
560
466
546

Is there a significant difference between the proportions of surviving seeds for low water versus low nutrients
in the male category? Use .01.
Do the data provide sufficient evidence to indicate a difference between survival proportions for low water
versus low nutrients for the female category? Use .01.
Look at the data for the other four treatmentsex combinations. Does it appear that there are any significant
differences due to low levels of water or nutrients? Explain.

9.6 Exercise 8.1 described a study of 392 healthy children who live in the area of Tours, France. The table shows some
results of the study measuring the serum levels of certain fat-soluble vitaminsnamely, vitamin A (retinol) and carotenein both boys and girls.

Retinol ( g/dl)

-carotene (
g/l)

Boys (n = 207)

Girls (n = 185)

x1 43.0
s1 13.0
x1 588

x2 41.8
s 2 10.7
x2 553

s1 406
a.
b.
c.

s2 350

Does the data provide sufficient evidence to indicate a difference between the mean retinol levels for boys
versus girls? Use .01.
Does the data indicate a significant difference between the mean -carotene levels for boys versus girls? Use
.01.
Would the results of parts a and b change if you had used .05?

9.7 Psychology Today reports on a study by environmental psychologist Karen Frank and two colleagues, which was
conducted to determine whether it is more difficult to make friends in a large city than in a small town. Two
groups of graduate students were used for the study: 45 new arrivals at a private university in New York City and
47 at a prestigious university located in a town of 31,000 in an upstate rural area. The number of friends (more
than just acquaintances) made by each student was recorded at the end of 2 months and then again at the end of 7
months. The 7-month sample means are given here:
New York City
Sample Size
Sample Mean (7
months)
Population Mean (7
months)

45
5.1

Upstate Small
Town
47
5.3

a. If your theory is that it is more difficult to make friends in a city than in a small town, state the alternative
hypothesis that you would use for a statistical test. State your null hypothesis.
b. Suppose that 1 2.2 and 2 2.3 for the populations of numbers of friends made by new graduate
students after 7 months at a new location. Do the data provide sufficient evidence to indicate that the mean
numbers of friends acquired after 7 months differ for the two locations? Test using .05.
9.8 According to the National Center for Education Statistics in Washington, DC, approximately
16% of all elementary school teachers are men. A researcher randomly selected 1000
elementary school teachers in California from a statewide computer database and found
that 142 were men. Does this sample provide sufficient evidence that the percentage of
male elementary school teachers in California is different from the national percentage?
9.9 American teenagers are very excited about technology and the impact it will have on their lives. In fact, most
have regular access to computers and the Internet. A survey of 508 teenagers, aged 12-17, was conducted for
Newsweek with the following sample results. Assume that the sample consisted of 272 girls and 236 boys.

Have online experience


Think technology makes a
positive difference in their
lives
a.
b.
c.

Boys
66%

Girls
56%

57%

46%

Is there a difference in the proportions of boys and girls who surf the net? Calculate the test statistic and its
p-value. Use the p-value to evaluate the statistical significance of the results.
Do boys and girls differ in the proportions who think that technology makes a positive difference in their
lives? Use one of the two methods of testing presented in this section, and explain your conclusions.
Are the two tests in parts a and b independent? Explain how this might affect the reliability of your
conclusions.

9.10 Advertising at the Movies Welcome to the new movie pre-show! Before you can see the newly released
movie you have just paid to see, you must sit through a variety of trivia slides, snack bar ads, paid product
advertising, and movie trailers. Although the total barrage of advertising may last up to 20 minutes or more, a
particular theater chain claims that the average length of any one advertisement is no more than 3 minutes. To test
this claim, 50 theater advertisements were randomly selected and found to have an average duration of 3 minutes
15 seconds with a standard deviation of 30 seconds. Do the data provide sufficient evidence to indicate that the

average duration of theater ads is more than that claimed by the theater? Test at the 1% level of significance.
(HINT: Change seconds to fractions of a minute.)
9.11 Princeton Review In Exercise 1.15, we examined an advertising flyer for the Princeton Review, a review course
designed for high school students taking the SAT tests. The flyer claimed that the average score improvements for
students who have taken the Princeton Review course is between 110 and 160 points. Are the claims made by the
Princeton Review advertisers exaggerated? That is, is the average score improvement less than 110, the minimum
claimed in the advertising flyer? A random sample of 100 students who took the Princeton Review course
achieved an average score improvement of 107 points with a standard deviation of 13 points.
a.
b.
c.

Use the p-value approach to test the Princeton Review claim. At which significance levels can you reject H0?
If you were a competitor of the Princeton Review, how would you state your conclusions to put your
company in the best possible light?
If you worked for the Princeton Review, how would you state your conclusions to protect your companys
reputation?

9.12 TV Habits An article in American Demographics reports that approximately 60% of U.S. households have two
or more television sets and that at least half of Americans sometimes watch television alone. Suppose that n = 75
U.S. households are sampled, and of those sampled, 49 had two or more television sets and 35 respondents
sometimes watch television alone.
a.
b.
c.

Two claims can be tested using the sample information. What are the two sets of hypotheses to be tested?
Do the data present sufficient evidence to contradict the claim that at least half of Americans sometimes
watch television alone?
Do the data present sufficient evidence to show that the 60% figure claimed in the magazine article is
incorrect?

9.13 Rocking the Vote What proportion of adults in the United States always votes in the presidential elections? An
article in American Demographics reports this percentage as 67%. To test this claim, a random sample of 300
adults was taken, and 192 of them said that they always vote in presidential elections. Does this sample provide
sufficient evidence to indicate that the percentage of adults who say that they always vote is different from their
reported percentage? Test using = 0.01.
9.14 Clopidogrel A large study was conducted to test the effectiveness of an experimental blood thinner, clopidogrel,
in warding off heart attacks and strokes. The study involved 19,185 patients who had suffered heart attacks,
strokes, or pain from clogged arteries. They were each randomly assigned to take either aspirin or clopidogrel for
a period of 1 to 3 years. Of the patients taking aspirin, 5.3% suffered heart attacks, strokes, or death from
cardiovascular disease; the corresponding percentage in the clopidogrel patients was 5.8%.
a.
b.

The article states that each patient was randomly assigned to one of the two medications. Explain how you
could use the random number table to make these assignments.
Although the article does not give the sample sizes, assume that the randomization in part a results in 9925
aspirin and 9260 clopidogrel assignments. Are the results of the study statistically significant? Use the
appropriate test of hypothesis.

9.15 Commercials in Space The commercialism of our space program was the topic of Exercise 8.16. In a survey of
500 men and 500 women, 20% of the men and 26% of the women responded that space should remain
commercial-free.
a.
b.

Is there a significant difference in the population proportions of men and women who think that space should
remain commercial-free? Use = 0.01.
Can you think of any reason why a statistically significant difference in these population proportions might
be of practical importance to the administrators of the space program? To the advertisers? To the
politicians?

9.16 The Death Penalty A researcher believes that the fraction p1 of Republicans in favor of the death penalty is
greater than the fraction p2 of Democrats in favor of the death penalty. She acquired independent random samples
of 200 Republicans and 200 Democrats and found 136 Republican and 124 Democrats favoring the death penalty.
Do these data support the researchers belief?
a.
b.

Find the p-value for the test.


If you plan to conduct your test using = 0.05, what will be your test conclusions?

9.17 The Death Penalty, continued Refer to Exercise 9.16. Some thought should have been given to designing a test
for which is tolerably low when p1 exceeds p2 by an important amount. For example, find a common sample size
n for a test with = 0.05 and 0.20, when in fact p1 exceeds p2 by 0.1. [HINT: The maximum value of p (1
p) is 0.25.]
9.18 The Cheeseburger Bill In the spring of 2004, the U.S. Congress considered a bill that would prevent
Americans from suing fast-food giants like McDonalds for making them overweight. Although the fast-food
industry may not be to blame, a study by Childrens Hospital in Boston reports that about two-thirds of adult
Americans and about 15% of children and adolescents are overweight. To test this claim, a random sample 100
children is selected, and 13 of them are found to be overweight. Is there sufficient evidence to indicate that the
percentage reported by Childrens Hospital is incorrect? Test at the 5% level of significance.
9.19 Summer Jobs In Exercise 8.19, it was reported that amusement and theme parks average 100 full-time
employees, but at peak season average 665 workers per site. If n = 40 amusement and theme parks are sampled at
peak season and the average number of full-time workers is found to be 652 with a standard deviation of 32, does
this contradict the reported average of 665 workers per site?
a.
b.
c.
d.

What are the null and alternative hypotheses to be tested?


Use the Large-Sample Test of a Population Mean applet to find the observed value of the test statistic.
Use the Large-Sample Test of a Population Mean applet to find the p-value for this test.
Based on your results from part c, what conclusions can you draw about the average of 665 workers reported
in Exercise 8.19?

Chapter 10:
10.1 One family of fish native to tropical Africa can breathe both in air and in water. Thus, the fish can live in water
with relatively low oxygen content by taking a portion of their oxygen from the air. Michael Pettit and Thomas
Beitinger conducted a study to investigate the effects of some variables on oxygen partitioning by reedfishthat
is, the proportions of oxygen taken by reedfish from the surrounding air and from the water in which they lived.
The data in the table represent the oxygen uptake readings from air and water for 11 reedfish in water at 25C. The
first column gives the weights of the fish. The second and third columns give the averages of repeated
measurements of oxygen uptake from air and water, respectively, for each fish. The fourth and fifth columns give
the percentage of total oxygen uptake attributable to air and water, respectively.
Oxygen Uptake (ml
O2/g/h)

Weight
(grams
)
22.1
21.1
13.4
19.5
21.9
19.5
16.5
18.5
18.0
15.4
13.4

Percentage

Air

Water

Air

Water

.043
.034
.026
.048
.009
.028
.032
.045
.061
.028
.024

.045
.066
.084
.045
.056
.072
.081
.051
.041
.060
.133

49
34
24
52
14
28
28
47
60
32
18

51
66
76
48
86
72
72
53
40
68
82

Percentages were calculated by the author based on Pettit and Beitingers data.

Assume that the reedfish included in this experiment represent a random sample of
mature reedfish.
a. Find a 90% confidence interval for the mean weight of mature reedfish. Interpret the interval.
b. Find a 90% confidence interval for the mean oxygen uptake from the air for redfish living in an environment
similar to that used in this experiment. Interpret your interval.
c. Find a 90% confidence interval for the mean percentage of oxygen uptake from air in reedfish living in an
environment similar to that used in this experiment. Interpret your interval.

10.2 Refer to the oxygen partition study in reedfish conducted by Pettit and Beitinger in Exercise10.1. The
accompanying table gives the percentages of oxygen uptake by air (the rest is by water) for reedfish exposed to
two temperature environments: one group of

n 1 11 reedfish in water at 25C and the other group of

n2 12 reedfish in water at 33C. Use the Minitab printout to answer the questions.
25C

49
52
28

34
14
47

24
28
60

33C

32
18

28
41
54

55
27
67

45
44
46

51
48
59

Minitab output for Exercise 10.2


Two Sample T-Test and Confidence Interval
Two sample T for 25 Deg C
N
Mean
25 Deg C
11
35.1
33 Deg C
12
47.1

vs 33 Deg C
StDev SE Mean
14.9
4.5
11.6
3.4

95% CI for mu 25 Deg C - mu 33 Deg C: ( -23.5, -0.5)


T-Test mu 25 Deg C = mu 33 Deg C (vs not =): T = -2.17 P = 0.042 DF = 21
Both use Pooled StDev = 13.3
a. Does the data in the table present sufficient evidence to indicate a difference between the percentages of
oxygen uptake by air for reedfish exposed to water at 25C and those at 33C? Test using .05.
b. What is the p-value for the test?
c. The Minitab analysis uses the pooled estimate of 2 . Is the assumption of equal variances reasonable? Why
or why not?
10.3 Persons submitting computing jobs to a computer center are usually required to estimate the amount of computer
time required to complete the job. This time is measured in CPUs, the amount of time that a job will occupy a
portion of the computers central processing units memory. A computer center decided to compare the estimated
versus actual CPU times for a particular customer. The corresponding times (in minutes) were available for 11
jobs.
Job Number
CPU Time
1
2
3
4
5
6
7
8
9
10
11
Estimated
.50
1.40
.95
.45
.75
1.20
1.60
2.6
1.30
.85
.60
Actual
.46
1.52
.99
.53
.71
1.31
1.49
2.9
1.41
.83
.74
a. Why would you expect these pairs of data to be correlated?

b. Do the data provide sufficient evidence to indicate that, on the average, the customer tends to underestimate
the CPU time required for computing jobs? Test using

= .05.

c. Find the approximate p-value for the test and interpret its value.
d. Find a 90% confidence interval for the difference in mean estimated CPU time versus mean actual CPU time.
10.4 Two computers are often compared by running a collection of different benchmark programs and recording the
difference in CPU times required to complete the same program. Six benchmark programs, run on two computers,
produced the following CPU times (in minutes):
Computer
1
2

1
1.12
1.15

2
1.73
1.72

Benchmark Program
3
4
5
1.04
1.86
1.47
1.10
1.87
1.46

6
2.10
2.15

Does the data provide sufficient evidence to indicate a difference in mean CPU times required for the two
computers to complete a job?

10.5 These data are the occupancy costs in 1996 and 1997 for prime office space in 13 major cities in the United States:
City
1

1997

2
3
4
5
6
7
8
9
10
11
12
13

$18.
36
32.8
2
23.5
8
17.5
2
19.1
2
14.8
5
30.5
0
25.0
6
30.8
9
35.7
4
19.3
3
30.9
2
34.3
0

1996

$18.
41
31.3
4
37.3
6
16.5
8
21.3
5
14.5
9
31.0
0
26.2
1
31.5
2
35.2
1
19.5
5
25.7
5
33.9
1

a. Explain why a paired-difference design is appropriate to determine whether there has been a significant
change in the cost of prime office space as measured in the 13 cities.
b. Can you conclude that the average cost of prime office space has changed significantly over this 1-year
period?
10.6 Refer to the comparison of the estimated versus actual CPU times for computer jobs submitted by the customer in
Exercise 10.3.
Job Number
CPU Time

Estimated

1
.50

Actual

.46

2
1.4
0
1.5
2

3
.95

4
.45

5
.75

.99

.53

.71

6
1.2
0
1.3
1

7
1.6
0
1.4
9

8
2.6
2.9

9
1.3
0
1.4
1

10
.85

11
.60

.83

.74

Find a 90% confidence interval for the variance of the differences between the estimated and actual CPU times.
10.7 Refer to Exercise 10.6. Suppose that the computer center expects customers to estimate their CPU time correct to
within .2 minute most of the time (say, 95% of the time).
a. Do the data provide sufficient evidence to indicate that the customers error in
estimating job CPU time exceeds the computer centers expectation? Test using
.10.
b. Find the approximate p-value for the test and interpret its value.

10.8 In Exercise 10.2, you conducted a test to detect a difference between the percentages of oxygen uptake by air in
reedfish exposed to water at 25C versus those at 33C.
a. What assumption had to be made concerning the population variances so that the test
would be valid?
b. Do the data provide sufficient evidence to indicate that the variances violate the assumption in part a? Test
using 05.

10.9 These data are the price per square foot for existing homes in 13 different areas of Southern California:
Area
1
2
3
4
5
6
7
8
9
10
11
12
13
a.
b.

June-Aug
2001
$83.89
88.72
94.94
72.92
81.21
83.69
146.89
84.80
87.81
90.08
92.09
82.80
80.73

June-Aug
2000
$72.47
72.94
88.86
54.65
72.06
71.08
136.20
75.54
73.43
80.06
83.13
72.05
63.67

Explain why a paired-difference design is appropriate to determine whether there has been a significant
change in the average cost per square foot of existing homes.
Can you conclude that the average cost per square foot of existing homes has changed significantly over this
1-year period?

10.10
In behavior analysis, self-control is most often defined as the choice of a large, delayed reinforcer over a
small, immediate reinforcer. In an experiment by Stephen Flora and fellow researchers to determine the effects of
noise on human self-control, subjects pressed two buttons for points that were exchangeable for money. Pressing
one of the buttons produced 2 points over 4 seconds; pressing the other button, the self-control choice, produced
10 points over 4 seconds after a delay of 16 seconds. Seven subjects were asked to make 30 choices in the
presence of aversive noise and in the absence of noise. Another six subjects were asked to make 30 choices, first
in the absence of noise and then in the presence of aversive noise. The number of correct choices for each subject
for both conditions is given here:
Subject
1
2
3
4
5
6
7
Means

Noise
17
24
11
16
25
12
17
17.43

Quiet
11
27
13
30
27
17
12
19.57

Subject
8
9
10
11
12
13

Quiet
26
26
10
19
20
26

Noise
30
30
25
24
26
30

21.17

27.5

a.

Test for a significant difference in the mean numbers of self-control choices made by the first seven
subjects and the second six subjects under the quiet condition at the .05 level of significance.

b.

Test for a significant difference in the mean numbers of self-control choices made by the first seven
subjects and the second six subjects under the noise condition at the .05 level of significance.

10.11
When should one use a paired-difference analysis in making inferences concerning the difference between
two means?
10.12
Repair Costs Car manufacturers try to design the bumpers of their automobiles to prevent costly damage in
parking-lot type accidents. To compare two models of automobiles, the cars were purposely subject to a series of
four front and rear impacts at 5 mph, and the repair costs were recorded.
Impact Type
Front into barrier
Rear into barrier
Front into angle barrier
Rear into pole

Honda Civic
$403
447
404
227

Hyundai Elantra
$247
0
407
185

Do the data provide sufficient evidence to indicate a difference in the average cost of repair for the Honda Civic
and the Hyundai Elantra? Test using = 0.05.

Chapter 11:
11.1 In Good Hands The cost of automobile insurance varies by location, ages of the drivers and type of coverage.
The following are estimates for a 6-month policy for basic liability coverage for a single male who has been
licensed for 6-8 years with no violations or accidents, and who drives between 12,600 and 15,000 miles per year
provided by the California Department of Insurance for the year 2003 on the Web site
(http://www.insurance.ca.gov).
Insurance Company
21st
Firemans
State
Location
Century
Allstate
AAA
Fund
Farm
Riverside
$736
$745
$668
$1065
$1202
San Bernardino
836
725
618
869
1172
Hollywood
1492
1384
1214
1502
1682
Long Beach
996
884
802
1571
1272
a.
b.
c.

d.

What type of design was used in collecting these data?


Is there sufficient evidence to indicate that insurance premiums for the same type of coverage differ from
company to company?
Is there sufficient evidence to indicate that insurance premiums vary from location to location?
Use Tukeys procedure to determine which insurance companies listed here differ from others in the
premiums they charge for this typical client. Use = 0.05.

11.2 One-stop Shopping Wal-Mart has come to symbolize low costs not only on clothing and household items, but
also for food staples. The Press Enterprise (Riverside, California) surveyed several local food markets to compare
prices on ten different items with those at a Wal-Mart Supercenter in Henderson, Nevada. The results of their
survey are given in the following table
WalStater
Items
Mart
Vons
Ralphs
Bros
Gallon of Milk
2.75
2.89
3.79
3.76
1 lb. Land OLakes butter
3.12
3.99
5.39
3.99
Dozen of AA large eggs
1.65
2.69
2.69
2.39
Cheerios, 15 oz.
2.50
3.99
3.99
3.79
Minute Maid OJ, 12 oz. can
1.34
2.29
2.00
1.50
Oroweat whole-wheat 24 oz. loaf
3.08
3.19
3.19
3.19
Coca-Cola Classic, 2-liter bottle
1.08
1.69
1.39
1.25
Maxwell coffee, 13 oz. can
2.48
3.99
3.19
2.69
Campbells chicken noodle soup, 10.75 oz.
0.50
0.89
0.99
0.79
Hot Dogs, Oscar Mayer, turkey and pork, 1 lb.
2.77
3.69
3.79
3.49
a.
b.
c.

What are the blocks and treatments in this experiment?


Do the data provide evidence to indicate that there are significant differences in prices from store to store?
Support your answer statistically using the ANOVA printout that follows.
Are there significant differences from block to block? Was blocking effective?

Two way ANOVA: Cost versus Items, Stores


Source
Items
Stores
Error
Total
S = 0.3738

DF
9
3
27
39

SS
41.8092
4.9769
3.7729
50.5590

MS
4.64547
1.65895
0.13974

R-Sq = 92.54%

F
33.24
11.87

P
0.000
0.000

R-Sq(adj) = 89.22%

11.3 One-stop Shopping, continued Refer to Exercise 11.2. The printout that follows provides the average costs of
the selected items for the k = 4 stores.
Stores
Ralphs
Stater Brothers
Vons

3.0410
2.6840
2.9300

Wal-Mart
a.
b.
c.

2.1270

What is the appropriate value of q0.05(k, df) for testing for differences among stores?
What is the value of = q0.05(k, df)[(MSE)/b]?
Use Tukeys pairwide comparison test among stores used to determine which stores differ significantly in
average prices of the selected items.

11.4 Crash Tests Recent information on average crash tests conducted by the Insurance Institute for Highway Safety
regarding bumper repair costs for damage sustained in front and rear crashes of vehicles into barriers/pole at 5
mph are given in the following table. These types of crash tests evaluate how well the bumpers prevent costly
damage to vehicles in parking lot-type impacts.
Front
Rear
Front into
Rear
into
into
Angle
into
Autos
Barrier
Barrier
Barrier
Pole
Hyundai Elantra
$ 247
$ 0
$ 407
$ 185
Ford Focus
31
1137
507
939
Honda Civic
403
447
404
227
Dodge Stratus/Chrysler Cirrus
278
174
626
1473
Lexus LS 30
75
395
1526
765
Dodge Grand Caravan
329
822
703
2268
Isuzu Rodeo
1769
924
1932
552
Mitsubishi Montero
1210
2495
2525
2831
Source: From Expensive Bumper Repair: Latest Crash Tests, Consumers Research, April
2001, pp. 20-22. Copyright 2001 Consumers Research. Reprinted by permission.
a.
b.
c.
d.

What is the type of design that was used in these crash tests? If the design used is a randomized block
design, what are the blocks and what are the treatments?
Are there significant differences in the cost of crashes for the vehicles considered here?
Are there significant differences among the four types of crashes?
Use Tukeys pairwise procedure to investigate the differences in average repair costs for the eight vehicles.
Comment on the results found using this procedure. Use = 0.05.

11.5 Professors Salaries In a study of starting salaries of assistant professors, five male assistant professors and five
female assistant professors at each of three types of institutions granting doctoral degrees were polled and their
initial starting salaries were recorded under the condition of anonymity. The results of the survey in $1,000 are
given in the following table.
Gender
Males

Females

a.
b.
c.
d.
e.

Public Universities
49.3
49.9
48.5
68.5
54.0
45.4
54.7
67.0
61.2
63.3

Private-Independent
81.8
71.2
62.9
69.0
69.0
55.1
62.1
59.5
54.8
69.7

Church-Related
66.9
57.3
57.7
46.2
52.2
50.4
52.1
49.8
61.9
51.6

What type of design was used in collecting these data?


Use an analysis of variance to test if there are significant differences in gender, in type of institution, and to
test for a significant interaction of gender type of institution.
Find a 95% confidence interval estimate for the difference in starting salaries for male assistant professors
and female assistant professors. Interpret this interval in terms of a gender difference in starting salaries.
Use Tukeys procedure to investigate differences in assistant professor salaries for the three types of
institutions. Use = 0.01.
Summarize the results of your analysis.

11.6 Professors Salaries II Each year, the American Association of University Professors reports on salaries of
academic professors at universities and colleges in the United States. The following data, adapted from this
report, are based on samples of n = 10 in each of four professorial ranks, for both male and female professors.

Gender
Male

Assistant Professor
$56.9
$57.4
56.9
55.2
57.8
57.2
61.3
57.9
60.5
60.5

Associate Professor
$61.0
$65.4
68.7
68.2
69.1
67.3
67.0
69.8
61.1
64.1

Full Professor
93.4
94.5
95.3
88.4
96.5
90.3
95.6
90.9
102.3
93.9

Female

49.6
52.0
56.4
57.3
95.3
85.9
50.6
51.6
62.9
65.6
82.0
87.8
46.6
54.9
56.9
64.0
76.5
87.0
57.4
55.9
58.9
60.4
88.5
81.7
55.6
52.8
64.6
62.0
80.6
82.8
Source: Based on Average Salary for Men and Women Faculty, by Category,
Affiliation, and Academic Rank 2000-2001, Academe, March-April 2001, p. 39.
Copyright 2001 American Association of University Professors
a.
b.
c.
d.
e.

Identify the design used in this survey.


Use the appropriate analysis of variance for these data.
Do the data indicate that the salary at the different ranks vary by gender?
If there is no interaction, determine whether there are differences in salaries by rank, and whether there are
differences by gender. Discuss your results.
Plot the average salaries using an interaction plot. If the main effect of ranks is significant, use Tukeys
method of pairwise comparisons to determine if there are significant differences among the ranks.

Chapter 12:
12.1 Daily readings of air temperatures and soil moisture percentages were collected by William Reiners and
colleagues at four elevations, 670, 825, 1145, and 1379 meters, in the White Mountains of New Hampshire.
Among the analyses of the data are simple linear regressions relating air temperature y (in C) in a given month to
altitude x (in meters). For example, the regression analysis for July produced these statistics: a = 21.4, b = -.0063,
r 2 = .99, and the estimate of the standard error of b is .0007.
a.
b.
c.

Give the least-squares prediction equation relating air temperature y (in C) to altitude x (in meters) for the
month of July.
How many degrees of freedom are associated with SSE and s 2 in this analysis? (Assume that readings were
taken each day during the month of July at all four elevations.)
Do the data present sufficient evidence to indicate that altitude x contributes information for predicting air
temperature y?

12.2 Another simple linear regression analysis presented in the paper by William Reiners and colleagues involved the
relationship between the number y of degree-days (the sum of days times the number of degrees above 0 C for
those days) per year and the altitude x in meters. The analysis is based on the recorded number of degree-days for
two winters at each of the four altitudes. Statistics presented for this analysis are a= 3393.93, b= - 1.319, and r 2
= .99, and the standard error of b is .091.
a.
b.
c.
d.

How many degrees of freedom are associated with SSE and s 2 in this analysis?
Do the data provide sufficient evidence to indicate that altitude x provides information for the prediction of
degree-days y? Test using =.01.
Give the least-squares prediction equation.
Predict the number of degree-days at an altitude of 825 meters. Do the statistics enable you to construct a
prediction interval for your prediction? Explain.

12.3 Shareholders have filed a record number of class actions suits this year, with half of these suits alleging IPO
(initial public offering) foul play. The numbers of federal securities-fraud class action suits, y, filed from 1996
through 2001 are given in the following table. The years are coded as Year 1995.
y
Time

110
1

178
2

236
3

205
4

211
5

282
6

Use the MINITAB printout to answer the questions that follow.

Regression Analysis: y versus Time


The regression equation is
y = 111 + 26.5 Time
Predictor
Constant
Time
S = 33.0405

Coef
110.87
26.514

SE Coef
30.76
7.898

R-Sq = 73.8%

Analysis of Variance
Source
DF
SS
Regression
1 12303
Residual Error
4
4367
Total
5 16669

T
3.60
3.36

P
0.023
0.028

R-Sq(adj) = 67.3%
MS
12303
1092

F
11.27

P
0.028

Predicted Values for New Observations


New
Obs
Fit SE Fit
95% CI
95% PI
1 296.5
30.8 (211.1, 381.9) (171.1, 421.8)
Values of Predictors for New Observations
New
Obs Time
1 7.00

a.
b.
c.
d.

What is the least-squares line relating the number of class action suits and time?
Is the number of class action suits linearly related to time over the period 1996 to 2001? Using a regression
analysis, test the hypothesis H 0 : 0 , that there is no linear relationship between y and time.
What proportion of the total variation is explained by the regression of y on time?
What is the expected number of fraud actions for 2002? Are there any limitations on your estimate?

12.4 How many weeks can a movie run and still make a reasonable profit? The data that follow show the number of
weeks in release (x) and the average per-site gross (y) for the top 19 movies during the first weekend in September
2001.
Movie
Jeepers Creepers
Rush Hour 2
American Pie 2
The Others
Rat Race
The Princess Diaries
O
Jay and Silent Bob Strike
Back
Summer Catch
Captain Corellis Mandolin
Planet of the Apes
Jurassic Park III
John Carpenters Ghosts of
Mars
The Curse of the Jade
Scorpion
The Deep End
Legally Blonde
Americas Sweethearts
American Outlaws
Bubble Boy
a.

Per-site
Average
$5,378
4,146
3,536
3,742
3,589
2,822
4,823
2,357

Weeks in
Release
1
5
4
4
3
5
1
2

2,104
2,440
2,557
2,030
1,028

2
3
6
7
2

2,191

5,394
1,971
835
1,294
804

4
8
7
3
2

Plot the points in a scatterplot. Does it appear that the relationship between x and y is linear? How would
you describe the direction and strength of the relationship?

b.
c.
d.

Calculate the value of r2. What percentage of the overall variation is explained by using the linear model
rather than y to predict the response variable y?
What is the regression equation? Do the data provide evidence to indicate that x and y are linearly related?
Test using a 5% significance level.
Given the results of parts b and c, is it appropriate to use the regression line for estimation and prediction?
Explain your answer.

12.5 LCD TVs In Exercise 3.6, Consumer Reports gave the price (y) for the top 10 LCD standard definition TVs in
the 14 to 20 inch category along with the screen size (x) in inches:
Brand
Sharp LC-20E1U
Sony KLV-15SR1
Panasonic TC-20LA1
Panasonic TC-17LA1
Gateway GTW-L18M103
Panasonic TC-14LA1
Gateway GTW-L17M103
Toshiba 14VL43U
Toshiba 20VL43U
Sharp LC-15E1U

Price
$1200
800
1050
750
700
500
600
670
1200
650

Size
20
15
20
17
18
14
17
14
20
15

Does the price of an LCD TV depend on the size of the screen? Suppose we assume that the relationship between
x and y is linear, and run a linear regression, resulting in a value r2 = 0.708.
a.
b.

What does the value of r2 tell you about the strength of the relationship between price and screen size?
The residual plot for this data, generated by MINITAB, is shown below. Does this plot suggest that the
relationship between price and screen size might be something other than linear?

c.

Plot the values of x and y using a scatterplot. Does this plot confirm your suspicions in part b? Explain.

12.6 Peyton Manning The number of passes completed and the total number of passing yards recorded for Peyton
Manning, quarterback for the Indianapolis Colts, were recorded for five games randomly selected from the games
played during the 2003 NFL season.
Week

Completions

Total Yards

8
15
2
10
11
a.
b.
c.

22
25
14
28
27

269
290
173
347
401

What is the least-squares line relating the total passing yards to the number of pass completions for Peyton
Manning?
What proportion of the total variation is explained by the regression of total passing yards (y) on the number
of pass completions (x)?
If they are available, examine the diagnostic plots to check the validity of the regression assumptions.

12.7 Peyton Manning, continued Refer to Exercise 12.6.


a.
b.
c.

Estimate the average number of passing yards for games in which Peyton throws 21 completed passes using a
95% confidence interval.
Predict the actual number of passing yards for games in which Peyton throws 21 completed passes using a
95% prediction interval.
Would it be advisable to use the least-squares line from Exercise 12.6 to predict Peytons total number of
passing yards for a game in which he threw only five completed passes? Explain.

12.8 Baseball Stats Does a teams batting average depend in any way on the number of home runs hit by the team?
The data in the table show the number of team home runs and the overall team batting average for eight randomly
selected major league baseball teams in the very early stages of the 2004 baseball season.
Team
Total Home Runs
Anaheim
26
Atlanta
21
Baltimore
21
Boston
26
Chicago
38
Houston
25
Philadelphia
26
San Diego
17
Source: espn.com
a.
b.
c.

Team Batting Average


.279
.271
.282
.257
.276
.290
.245
.271

Plot the points using a scatterplot. Does it appear that there is any relationship between total home runs and
team batting average?
Is there a significant positive correlation between total home runs and team batting average? Test at the 5%
level of significance.
Do you think that the relationship between these two variables would be different if we had looked at an
entire baseball season, rather than just the very early part of the season?

12.9 Movie Reviews How many weeks can a movie run and still make a reasonable profit? The data that follow show
the number of weeks in release (x) and the average per-site gross (y) for the top 20 movies during the last weekend
in January 2004.

1.
2.
3.
4.
5.
6.
7.
8.
9.
10.
11.
12.
13.

Movie
The Butterfly Effect
Along Came Polly
Win a Date with Tad Hamilton
Big Fish
Lord of the Rings: The Return of the King
Cheaper by the Dozen
Cold Mountain
Torque
Somethings Gotta Give
Mystic River
Calendar Girls
The Last Samurai
Monster

Per-site
Average
$6551
5460
2700
2916
2653
2269
1786
1824
1875
2538
2470
1953
6219

Weeks
in Release
1
2
1
7
6
5
5
2
7
16
6
8
5

14. My Babys Daddy


15. Disneys Teachers Pet
16. Peter Pan
17. Paycheck
18. The Cooler
19. Mona Lisa Smile
20. 21 Grams
Source: Entertainment Weekly
a.
b.
c.
d.

1658
704
1090
1062
2238
882
2130

3
2
5
5
9
6
10

Plot the points in a scatterplot. Does it appear that the relationship between x and y is linear? How would
you describe the direction and strength of the relationship?
Calculate the value of r2. What percentage of the overall variation is explained by using the linear model
rather than - to predict the response variable y?
What is the regression equation? Do the data provide evidence to indicate that x and y are linearly related?
Test using a 5% significance level.
Given the results of parts b and c, is it appropriate to use the regression line for estimation and prediction?
Explain your answer.

12.10
Hotel Satisfaction Is your overall satisfaction with a hotel room correlated with the cost of the room?
Satisfaction scores and median room prices were recorded for nine budget hotels with the following results:
Hotel
Median Rate
Score
Sleep Inn
55
80
Microtel
50
75
Super 8 Motel
50
70
Red Roof Inn
50
69
Motel 6
40
68
Days Inn
55
64
Econo Lodge
50
63
Travelodge
55
62
Knights Inn
45
57
Source: Case Study Hotel Room Cost and Satisfaction Copyright 2001 by Consumers Union of U.S., Inc.,
Yonkers, NY 10703-1057, a nonprofit organization. Reprinted with permission from the July 2001 issue of
Consumer Reports for educational purposes only. No commercial use or reproduction permitted.
www.ConsumerReports.org
a.
b.
c.

Calculate the correlation coefficient r between price and overall score. How would you describe the
relationship between price and overall score?
Use the applet called Correlation and the Scatterplot to plot the nine data points. What is the correlation
coefficient shown on the applet? Compare with the value you calculated in part a.
Describe the pattern that you see in the scatterplot. Are there any outliers? If so, how would you explain
them?

Chapter 13:
13.1 LCD TVs In exercises 3.6 and 12.5, Consumer Reports gave the prices (y) for the top ten LCD standard
definition TVs in the 14- to 20-inch category along with the screen size (x) in inches:
Brand
Sharp LC-20E1U
Sony KLV-15SR1
Panasonic TC-20LA1
Panasonic TC-17LA1
Gateway GTW-L18M103
Panasonic TC-14LA1
Gateway GTW-L17M103
Toshiba 14VL43U
Toshiba 20VL43U
Sharp LC-15E1U

Price
$1200
800
1050
750
700
500
600
670
1200
650

Size
20
15
20
17
18
14
17
14
20
15

The diagnostic plots in Exercise 12.5 suggested that a quadratic model might provide a better fit for this data.

Regression Analysis: y versus x, x-sq


The regression equation is
y = 5751 676 x + 22.3 x-sq
Predictor
Constant
x
x-sq

Coef
5751
-676.2
22.270

S = 115.252
Analysis of Variance
Source
Regression
Residual Error
Total
a.
b.
c.

SE Coef
2776
329.4
9.613
R-Sq = 83.5%
DF
2
7
9

SS
469979
92981
562960

T
2.07
-2.05
2.32

P
0.077
0.079
0.054

R-Sq(adj) = 78.8%
MS
234989
13283

F
17.69

P
0.002

Write the equation of the quadratic model.


Does the quadratic term contribute significant information to the prediction of y, in the presence of the linear
term?
The fitted line plots below were generated by MINITAB. Use the values of R2(adj) to compare the linear and
quadratic models. Which one is a better fit to the data?

13.2 Digital Video Recorders Digital video recorders (DVR) are fairly new to the television-viewing public, but they
allow you to record and view programs at whatever time you request. Programs can be paused while you go to the
kitchen for a snack, and commercials can be eliminated. Once the DVR functions become integrated with satellite
and cable television setups, marketers predict that sales will skyrocket, as projected in the following table.
Year
DVRs (in millions)
a.
b.
c.
d.
e.
f.

2000
0.35

2001
0.88

2002
2.50

2003
5.70

2004
11.50

2005
20.20

Plot the predicted number of DVRs (y) as a function of the year (x) using a scatterplot. Describe the nature of
the relationship.
What model would you use to predict y as a function of x? Explain.
Using the model from part b and a computer software package, find the least-squares regression equation for
predicting the DVR market penetration that is, the number of DVRs installed in U.S. homes as a function
of year.
Does the model contribute information for the prediction of y? Test using = 0.01
What is the value of R2? What does this tell you about the fit of the model?
If they are available, examine the residual plots for the analysis. What conclusions can you draw?

Chapter 14:
14.1 One of many polls reported in a local newspaper investigated the opinions of southern Californians about their
quality of life. Residents were asked to rate their quality of life compared to 6 years ago in one of four categories.
The survey involved four groups of 300 people from each of four regions in southern California. The results of the
survey are shown in the table.

Los Angeles County


Orange County
San Diego
Inland Empire

Better

Not Much Change

Worse

No Opinion

63
78
69
78

96
84
93
81

135
132
129
111

6
6
9
30

Use the Minitab printout to determine whether there is a difference in the responses
according to region. Explain the nature of the differences, if any exist.

Minitab output for Exercise 14.1


Chi-Square Test
Chi-Sq = 1.125 + 0.636 + 0.537 + 3.574 +
0.500 + 0.229 + 0.217 + 3.574 +
0.125 + 0.229 + 0.040 + 1.103 +
0.500 + 0.636 + 1.957 + 23.338 = 38.319
DF = 9, P-Value = 0.000
14.2 Voter support for political term limits is strong in many parts of the United States. A poll conducted by the Field
Institute in California showed support for term limits by a 2-to-1 margin. The results of this poll of n= 347
registered voters are given in the table.

Republica
n
Democrat
Other
Total

For

Against

Total

35

No
Opinion
7

97
108
21
226

49
14
98

10
6
23

167
41
347

139

Use the Minitab printout to determine whether there is a difference in the responses
according to political affiliation. Explain the nature of the differences, if any exist.
Minitab output for Exercise 14.2
Chi-Square Test
Chi-Sq = 0.462 + 0.462 + 0.532 +
0.005 + 0.071 + 0.103 +
1.218 + 0.506 + 3.965 = 7.324
DF = 4, P-Value = 0.120
1 cells with expected counts less than 5.0
14.3 Title IX programs mandated that colleges should have equal opportunities for men and women to participate in
college sports programs. The table that follows shows the number of colleges that support womens teams in these
sports.
Years
1986-87
1999-00

Basket
ball
275
317

Track
215
276

Womens Sports
Golf
Soccer
99
192

54
260

Lacross
e
34
75

Rowing
27
71

a. Do these data provide sufficient evidence to conclude that there is a difference in the
distribution of women participating in these sports for the years indicated?
b. If there is a significant difference in part a, describe the nature of these

differences.
14.4 In a telephone poll of n 1,262 New York voters 795 or 63% of voters surveyed said the Port Authority of NY and
NJ should rebuild the World Trade Center in some form. Do these data provide sufficient evidence to indicate that
a majority of New York voters would like to see the New York World Trade Center rebuilt?
a. Use a test of hypotheses based on the z-statistic to answer this question.
b. Use a 2 - test to answer this same question.
c.

How do your answers compare? Check to see that z 2 2 .

14.5 The number of American workers who have Internet access continues to grow as recorded in the following table.
Two independent random samples were taken in two consecutive years.

Last
year
This year

Had
Access
224

Did Not Have


Access
276

Sample Size

328

172

500

500

Is there a significant difference in the proportion of workers who have access to the Internet last year compared to
this year? Use .05 .
14.6 The Generation Gap Is there a generation gap? A sample of adult Americans of three different generations were
asked to agree or disagree with this statement: If I had the chance to start over in life, I would do things
differently. The results, adapted from a Time magazine article, are given in the table. Do the data indicate a
generation gap for this particular question? That is, does a persons opinion change depending on the generation
group from which he or she comes? If so, describe the nature of the differences. Use = 0.05.

Agree
Disagree

GenXers
(born 1965-1976)
118
80

Boomers
(born 1946-1964)
213
87

Matures
(born before 1946)
88
61

14.7 Wealth and Education Levels Does education really make a difference in how much money you will earn?
Researchers randomly selected 100 people from each of three income categories marginally rich, comfortably
rich, and super rich and then recorded their educational attainment, as in the table.
Marginally
Rich
($70-99 K)
No college
Some college
At least an undergraduate degree
Postgraduate study/degree
a.
b.
c.

Comfortably
Rich
($100-249 K)
32
13
43
12

20
16
51
13

Super
Rich
($250 K or more)
23
1
60
16

Describe the multinomial experiments whose proportions are being compared in this experiment.
Do these data indicate that the level of wealth is affected by educational attainment? Test at the 1% level of
significance.
Based on the results of part b, describe the practical nature of the relationship between level of wealth and
educational attainment.

14.8 Politics and Religion A survey reported in the Riverside Press-Enterprise was conducted prior to the 2004
presidential election to explore the relationship between a persons religious fervor and their choice of a political
candidate. Voters were asked how often they attended church and which of the two major presidential candidates
(George W. Bush or his democratic opponent) they would favor in the 2004 election. The results of a similar
survey are shown.
Church Attendance
G.W. Bush
More than once a week
89
Once a week
87
Once or twice a month
93
Once or twice a year
114
Seldom/never
22
Source: Adapted from the Press-Enterprise
a.
b.

Democratic Candidate
53
68
85
134
36

Is there a significant difference in the proportion of voters who intend to vote for George W. Bush depending
on the voters religious fervor? Test using = 0.05
If there is a significant difference in part a, describe the nature of these differences.

14.9 Publish or Perish In the academic world, students and their faculty advisors often collaborate on research
papers, producing works in which publication credit can take several forms. In theory, the first authorship of a
students paper should be given to the student unless the input from the faculty advisor was substantial. In an
attempt to see whether this is, in fact, the case, authorship credit was studied for different levels of faculty input
and two objectives (dissertation versus non-degree research). The frequency of author assignment decisions for
published dissertations is shown in the table as assigned by 60 faculty members and 161 students.

Faculty Respondents
Authorship Assignment
Faculty first author, student mandatory second author
Student first author, faculty mandatory second author
Student first author, faculty courtesy second author
Student sole author

High Input

Student Respondents
Authorship Assignment
Faculty first author, student mandatory second author
Student first author, faculty mandatory second author
Student first author, faculty courtesy second author
Student sole author
a.
b.
c.
d.

4
15
2
2

Medium Input
0
12
7
3

Low Input
0
3
7
5

High Input
19
19
3
0

Medium Input
6
41
7
3

Low Input
2
27
31
3

Is there sufficient evidence to indicate dependence between the authorship assignment and the input of the
faculty advisor as judged by faculty members? Test using = 0.01
Is there sufficient evidence to indicate dependence between the authorship assignment and the input of the
faculty advisor as judged by students? Test using = 0.01
If there is a dependence in the two classifications from parts a and b, does it appear from looking at the data
that students are more likely to assign an higher authorship to their faculty advisors than the advisors
themselves?
Have any of the assumptions necessary for the analysis used in parts a and b been violated? What affect
might this have on the validity of your conclusions?

14.10
Car Colors Although white has long been the most popular car color, trends in fashion and home design
have signaled the emergence of green as the color of choice in recent years. The growth in the popularity of green
hues stems partially from an increased interest in the environment and increased feelings of uncertainty.
According to an article in The Press-Enterprise, green symbolizes harmony and counteracts emotional stress.
The article cites the top five colors and the percentage of the market share for four different classes of cars. These
data are for the truck-van category.
Color
Percent

White
29.72

Burgundy
11.00

Green
9.24

Red
9.08

Black
9.01

In an attempt to verify the accuracy of these figures, we take a random sample of 250 trucks and vans and record
their color. Suppose that the number of vehicles that fall into each of the five categories are 82, 22, 27, 21, and 20
respectively.
a.

Is any category missing in the classification? How many cars and trucks fell into that category?

b. Is there sufficient evidence to indicate that our percentages of trucks and vans differ from those given? Find
the approximate p-value for the test.
Chapter 15:
15.1 One explanation for the bright coloration of male birds versus the dull coloration of females and chicks is the
predator-deflection theory, which suggests that males have developed a bright color to attract predators to
themselves and away from their nesting females and young chicks. To test this theory, G. S. Butcher placed a
stuffed and mounted sharp-shinned hawk (Accipiter striatus) a distance of 12 meters from each of 12 northern
orioles nests. For each nest, he recorded the total number of hits and swoops on the hawk by both the male and
female oriole. He also noted which of the pair made the first attack on the hawk and which was more aggressive.
The data are shown in the accompanying table.
Total Number
of Hits and
Swoops
Presentation
Number

More
Aggressive
Sex (M or F)

1
2
3
4

4
11
16
14

18
47
33
0

F
F
F
M

Sex Seen
First

F
F
F
F

5
27
2
M
F
6
71
59
M
M
7
76
5
M
M
8
61
0
M
M
9
0
25
F
F
10
6
21
F
M
11
35
50
F
F
12
3
59
F
F
For each of the 12 trials, the mounted Accipiter was placed within 12 meters of
the nest for a 10-minute period.
M = male and F = female of the pair who fed offspring in the nest closest to the
mounted Accipiter.
Does the data present sufficient evidence to indicate a difference in the distributions of numbers and hits and
swoops for male versus female members of the nests?
a.
b.

Test using Wilcoxons signed-rank test with = 0.05. State your conclusions.
Give the approximate p-value for the test and interpret it.

15.2 Refer to Exercise 15.1. Does the data present sufficient evidence to indicate that the male member of a nest is
more aggressive than the female? Explain.
15.3 The data in the table are the occupancy costs in 1996 and 1997 for prime office space in 13 major cities in the
United States.
City
1
2
3
4
5
6
7
8
9
10
11
12
13

1997
$18.36

1996
$18.41

32.82
23.58
17.52
19.12
14.85
30.50
25.06
30.89
35.74
19.33
30.92
34.30

31.34
37.36
16.58
21.35
14.59
31.00
26.21
31.52
35.21
19.55
25.75
33.91

a. What nonparametric tests presented in this chapter can be used to decide whether
there has been a significant increase in the costs of prime office space from 1996 to
1997 in major U.S. cities? How do you express the null and alternative hypotheses?
b. Use two nonparametric tests to test whether the distribution of costs of prime office
space in 1997 lies to the right of the distribution of costs of prime office space in 1996.
Let = .05.
c. Compare the conclusions using the two tests in part b with those of Exercise 10.5 for
the paired-sample t test.