Introduction. Stata is a popular alternative to SPSS, especially for more advanced statistical
techniques. This handout shows you how Stata can be used for OLS regression. It assumes
knowledge of the statistical concepts that are presented. Several other Stata commands (e.g.
logit, ologit) often have the same general format and many of the same options.
Rather than specify all options at once, like you do in SPSS, in Stata you often give a series of
commands. In some ways, this is more tedious, but it also gives you flexibility in that you don’t
have to rerun the entire analysis if you think of something else you want. As the Stata 9 User’s
Guide says (p. 43) “The userinterface model is type a little, get a little, etc. so that the user is
always in control.”
For the most part, I find that either Stata or SPSS can give me the results I want, but there are some
tasks that can be done more easily in one program than the other. For example, I personally prefer
to do most of my database manipulation in SPSS and then convert the file to Stata, but that is partly
because I am much more familiar with the SPSS commands than their Stata counterparts.
Conversely, Stata’s statistical commands are generally far more logical and consistent (and
sometimes more powerful) than their SPSS counterparts. Luckily, with the separate Stat Transfer
program, it is very easy to convert SPSS files to Stata and viceversa.
Get the data. First, open the previously saved data set. (Stata, of course, also has means for
entering, editing and otherwise managing data.) You can give the directory and file name, or even
access a file that is on the web. For example,
. use http://www.nd.edu/~rwilliam/stats1/statafiles/reg01.dta, clear
Descriptive statistics. There are various ways to get descriptive statistics in Stata. Since you are
using different commands, you want to be careful that you are analyzing the same data throughout,
e.g. missing data could change the cases that get analyzed. The correlate command below uses
listwise deletion of missing data, which is the same as what the regress command does, i.e. a
case is deleted if it is missing data on any of the variables in the analysis.
. correlate income educ jobexp race, means
(obs=20)
Variable 
Mean
Std. Dev.
Min
Max
+income 
24.415
9.788354
5
48.3
educ 
12.05
4.477723
2
21
jobexp 
12.65
5.460625
1
21
race 
.5
.5129892
0
1

income
educ
jobexp
race
+income 
1.0000
educ 
0.8457
1.0000
jobexp 
0.2677 0.1069
1.0000
race  0.5676 0.7447
0.2161
1.0000
162 19.973396 Residual  281.92019 3 512.981124 . t P>t [95% Conf. Err.5707931 2.8164 4.1129776 1. Err.24589 3.1945 income  Coef.Regression.5940804 +Total  1820.54593 7.845 5. level(99) Source  SS df MS +Model  1538.92019 3 512.8118671 Number of obs F( 3.871949 0.037413 2.162 23.000 1. Specify the DV first followed by the IVs.959128 _cons  7.518362  Confidence Interval.16 0. Interval] +educ  1.8118671 Number of obs F( 3.0000 0.8454 0.5707931 2. Use the regress command for OLS regression (you can abbreviate it as reg).54 0.1811106 3.3231024 6.6419622 .54 0.8164 4.296178 2.369166 1. 16) Prob > F Rsquared Adj Rsquared Root MSE = = = = = = 20 29.863763 5.66607 jobexp  .863763 5. By default.659052 _cons  7.817542 8.505287 16 17. If you want to change the confidence interval. regress income educ jobexp race Using Stata 9 and Higher for OLS Regression – Page 2 .369166 1.1811106 3. Interval] +educ  1.6419622 .2580248 1. Std.845 7.1945 income  Coef.871949 0. regress income educ jobexp race. set level 99 . Stata will report the unstandardized (metric) coefficients. .20 0.13 0.924835 jobexp  . 16) Prob > F Rsquared Adj Rsquared Root MSE = = = = = = 20 29.517466 6.0259 race  . regress income educ jobexp race Source  SS df MS +Model  1538.16 0. you could use the set level command before regress: . t P>t [99% Conf.20 0.46 0.8184  As an alternative.003 .46 0.003 . use the level parameter: .505287 16 17.000 1.973396 Residual  281.42548 19 95.0000 0.8454 0.5940804 +Total  1820.981124 . Std.3231024 6.42548 19 95.170947 race  .13 0.
92019 3 512.426622 educ  2. .003 . Err.8164 4. beta Also. regress income educ jobexp race.g. if the data set is large.46 0.34 0. level(99) .9062733 jobexp  . beta Source  SS df MS +Model  1538. regress income educ jobexp race . add the beta parameter: .3581312 race  .Standardized coefficients.8454 0.20 0.16 0. It uses information Stata has stored internally.6419622 .42548 19 95.0299142 _cons  7. Std.13 0.89 NOTE: vif only works after regress. logit.369166 1.26 0.863763 5. if you just type regress Stata will “replay” (print out again) your earlier results.442403 jobexp  1.0000 0.5940804 +Total  1820.1811106 3. e. which is unfortunate because the information it offers can be useful with many other commands.3231024 6.  NOTE: The listcoef command from Long and Freese’s spost9 package of routines (type findit spost9 from within Stata) provides alternative ways of standardizing coefficients. regress.000 .845 .505287 16 17.5707931 2.871949 0.06 0. . you do not have to repeat the entire command when you change a parameter (indeed. Incidentally. VIF & Tolerances.g. regress.8118671 Number of obs F( 3.973396 Residual  281.54 0.981124 .) The last three regressions could have been executed via the commands . 16) Prob > F Rsquared Adj Rsquared Root MSE = = = = = = 20 29.162 . collin educ jobexp race if !missing(income) Using Stata 9 and Higher for OLS Regression – Page 3 . e. because then Stata will redo all the calculations. t P>t Beta +educ  1. To get the standardized coefficients. Phil Ender’s collin command gives more information and can be used with estimation commands besides regress. You run it AFTER running a regression.946761 +Mean VIF  1. vif Variable  VIF 1/VIF +race  2. you don’t want to repeat the entire command.1945 income  Coef. vif is one of many postestimation commands. Use the vif command to get the variance inflation factors (VIFs) and the tolerances (1/VIF).
add the coef parameter: .424 4.0030 If you want to see what the coefficients of the constrained model are. z P>z [95% Conf.38 0.9852105 . 16) = Prob > F = 27. To test whether the effects of educ and/or jobexp differ from zero (i.80 0.1521516 6.283422 race  6.981712 3.283422 jobexp  .6869989 1.0000 The test command does what is known as a Wald test.287976 0.9852105 . it gives the same result as an incremental F test. coef ( 1) educ . In this case. 16) = Prob > F = 12.692116 1.426358 4.07 0.977921 11.000 . Err. Std. testnl) that can be used to test even more complicated hypotheses. test educ=jobexp ( 1) educ .e.1521516 6. Interval] +educ  .000 .83064  The testparm and cnsreg commands can also be used to achieve the same results. β1 = β2.jobexp = 0 F( 1.21 0. If you want to test whether the effects of educ and jobexp are equal. test educ=jobexp.e.g.jobexp = 0 F( 1.0030 Constrained coefficients  Coef. Using Stata 9 and Higher for OLS Regression – Page 4 . you run the regression first and then do the hypothesis tests. .808031 _cons  3. 16) = Prob > F = 12. Again.5762 2.21 0. i. indeed I think it has some big advantages over SPSS here.001 10.Hypothesis testing. Stata has some very nice hypothesis testing procedures. these are postestimation commands. to test β1 = β2 = 0). use the test command: .48 0. test educ jobexp ( 1) ( 2) educ = 0 jobexp = 0 F( 2.6869989 1. Stata also has other commands (e.48 0.
8454 0.871949 0.32917 35.54 0.000 1.46 0.60934 3.973396 Residual  281.2798 income  Coef. Paul Bern’s hireg command performs a similar function.863763 5.39 0. the nestreg command will compute incremental Ftests. Interval] +race  10.8454 0.2580248 1. 16) Prob > F Rsquared Adj Rsquared Root MSE = = = = = = 20 29.981124 . (If you are stuck with Stata 8.g.517466 6.0090 0.369166 1. t P>t [95% Conf.55 1 18 0. Std. Alternatively.) The nestreg command is particularly handy if you are estimating a series/hierarchy of models and want to see the regression results for each one.3231024 6. To again test whether the effects of educ and/or jobexp differ from zero (i.66607 jobexp  .e.98098 18 68.444492 Residual  1233.003 . Std.Nested models.24589 3.444492 1 586.5940804 +Total  1820.618291 11.845 5. Err.0000 0.42548 19 95.16 0.162 19. the nestreg command would be .42548 19 95.702823 2. to test β1 = β2 = 0). 18) Prob > F Rsquared Adj Rsquared Root MSE = = = = = = 20 8. store) that can be very useful.505287 16 17.5544991 +Total  1820.13 0.296178 2. logit) and has additional options (e.8118671 Number of obs F( 1.000 24.0000 0.92019 3 512.5232  ++ nestreg works with several other commands (e.009 18.20 0.92 0. Using Stata 9 and Higher for OLS Regression – Page 5 . nestreg: Block regress income (race) (educ jobexp) 1: race Source  SS df MS +Model  586.6419622 .8118671 Number of obs F( 3. t P>t [95% Conf. Err.1945 income  Coef.659052 educ  1.518362 ++   Block Residual Change   Block  F df df Pr > F R2 in R2  +  1  8.83 3. lrtable.07 2 16 0.55 0.050657 _cons  29.0090 0. Interval] +race  .g.1811106 3.3221 0.33083 Block 2: educ jobexp Source  SS df MS +Model  1538.83 2.0259 _cons  7.5707931 2.3221   2  27.2845 8.8164 4.
pcorr2 income educ jobexp race (obs=20) Partial and Semipartial correlations of income with Variable  Partial SemiP Partial^2 SemiP^2 Sig.8118671 Number of obs F( 2.3634 0. pcorr2.Partial and SemiPartial Correlations. There is a separate Stata routine.33 0.22521 2 769. The sw prefix lets you do stepwise regression and can be used with many commands besides regress. Here is how to do backwards stepwise regression.413322  Using Stata 9 and Higher for OLS Regression – Page 6 .0500 begin with full model removing race Source  SS df MS +Model  1538. which gives the partial correlations but.845 Stepwise Regression. +educ  0.05): regress income educ jobexp race p = 0.8450 >= 0.0025 0.000 jobexp  0.148322 _cons  7.2099494 9.933393 .77 0.1721589 3.6000156 +Total  1820.000 1.0496 0. Note that SPSS is better if you need more detailed step by step results.8267 4.1214 0. sw. The pcorr command in Stata 11 took some of my code from pcorr2 and now gives very similar output. Err. pcorr.6632 0.541875 jobexp  .003 race  0. pr(. Use the pr (probability for removal) parameter to specify how significant the coefficient must be to avoid removal.002 .0743 income  Coef.067 17. Interval] +educ  1.150409 1. did not give the semipartials.21 0.96 0. I wrote a routine.7015 0. . t P>t [99% Conf. .4399 0. Std.0000 0. which gives both the partial and semipartial correlations.6028 0.0195 0.096855 3. 17) Prob > F Rsquared Adj Rsquared Root MSE = = = = = = 20 46.324911 2.0004 0.42548 19 95.200265 17 16.626412 1.60703 3.8375 0. prior to Stata 11.6493654 .112605 Residual  282.3485 0.8450 0.
119 32. One of the most common ways is the use of the if parameter on commands.40 0.To do forward stepwise instead.413322  Sample Selection.0500 adding educ p = 0.91281 7.060987 9 69.39 0. Std. use the pe (probability for entry) parameter to specify the level of significance for entering the model.7844 income  Coef. pe(. sw. we could type .933393 .8401 0.673443 Number of obs F( 2.8118671 Number of obs F( 2. In SPSS. Err.5314947 .067 17.055683 _cons  13. t P>t [95% Conf.0015 < 0. 17) Prob > F Rsquared Adj Rsquared Root MSE = = = = = = 20 46. reg income educ jobexp if race==0 Source  SS df MS +Model  526.002 .78 0.7944 3.2216794 2. for example.810224 2 263.541875 jobexp  .42219 4.8450 0.596569  Using Stata 9 and Higher for OLS Regression – Page 7 .150409 1. Std.827619 1.096855 3.626412 1.8267 4.002 1.67 0.21445 3. Interval] +educ  2.0073062 1.0000 0.21 0. Err. Interval] +educ  1.5265393 4.77 0.324911 2.048 . . So if.1721589 3. we only wanted to analyze whites.459518 .0000 < 0.250763 7 14.0500 adding jobexp Source  SS df MS +Model  1538.6493654 .96 0.405112 Residual  100.33 0. 7) Prob > F Rsquared Adj Rsquared Root MSE = = = = = = 10 18.704585 jobexp  .05): regress income educ jobexp race begin with empty model p = 0.3215375 +Total  627.148322 _cons  7.22521 2 769. Stata also has a variety of means for handling sample selection.0743 income  Coef.0016 0.42548 19 95.6000156 +Total  1820.200265 17 16.112605 Residual  282.000 1.60703 3.2099494 9. you can use Select If or Filter commands to control which cases get analyzed. t P>t [99% Conf.
by race: reg income educ jobexp _______________________________________________________________________________ > race = white Source  SS df MS +Model  526.4541661 3. Interval] +educ  2.048 .459518 .788485 .7314 0. You use Stata’s corr2data command for this. suppose we wanted to estimate separate models for blacks and whites. t P>t [95% Conf. you could use bysort: .058062 1. SPSS lets you input and analyze these directly.3215375 +Total  627.704585 jobexp  .3237189 2.4355552 Number of obs F( 2.5314947 . In Stata.91281 7. Std. Interval] +educ  1.405112 Residual  100. Or.827619 1.919997 9 67.7145528 2.7074115 . t P>t [95% Conf.21445 3.060987 9 69. we can use the by command (data must be sorted first if they aren’t sorted already): .006 .7844 income  Coef.40 0.67 0. explained in a moment) of the original data.01 0.250763 7 14.Separate Models for Groups. sort race .500947 6. 7) Prob > F Rsquared Adj Rsquared Root MSE = = = = = = 10 9. Sometimes you might want to replicate or modify a published analysis. Err.065 .2216794 2.94 0. bysort race: reg income educ jobexp Analyzing Means. Err.19 0.78 0.002 1.472885 _cons  6. Std.055683 _cons  13.673443 Number of obs F( 2.8401 0.826 income  Coef.810224 2 263. correlations and standard deviations.344 21. Correlations and Standard Deviations in Stata.42219 4.0073062 1.2900768 +Total  606.862417 jobexp  . rather than using separate sort and by commands.0016 0.0100 0. Using Stata 9 and Higher for OLS Regression – Page 8 . but the authors have published their means.596569 _______________________________________________________________________________ > race = black Source  SS df MS +Model  443. 7) Prob > F Rsquared Adj Rsquared Root MSE = = = = = = 10 18.646961  As an alternative.119 32. In Stata. you must first create a pseudoreplication (my term.39 0. In SPSS.53 0.889459 2 221.5265393 4.406053 1.6546 4. You don’t have the original data. we could use the Split File command.64886 8.94473 Residual  163.7944 3.030538 7 23.
matrix input Corr = (1.53. “Ability grouping and contextual determinants of educational expectations in Israel.For example. There are two main ethnic groups in Israel: the Ashkenazim . and correlations .82. sds.1. 0 otherwise X2 . we do the following.46.260 Parental Education Sc ale .72\.” Shavit and Williams examined the effect of ethnicity and other variables on the achievement of Israeli school children.697 X3 .530 .of European birth or extraction . in their classic 1985 paper. Their variables included: X1 .330 .330 Sc holastic Aptitude . and other Mideastern countries..2.. This scale ranges from a low of 4 to a high of 10.and the Sephardim.1..46.530 1.11.44000 .42000 N 10609 10609 10609 10609 Correlations Pears on Correlation Sephardim Parental Education Scale Sc holastic Aptitude Grade Point Average Sephardim 1. * First.1.5.000 . De scriptive Statis tics Sephardim Parental Educ ation Sc ale Sc holastic Aptitude Grade Point Average Mean .00..000 .12) . (I find that entering the data is most easily done via the input matrix by hand submenu of Data).82000 6.A scale which ranges from a low of 0 to a high of 1.12000 Std.42) . corr2data sphrd pared aptd grades.. most of whose families immigrated to Israel during the early fifties from North Africa.46000 7..000 .460 . n(10609) means(Mean) corr(Corr) sds(SD) (obs 10609) Using Stata 9 and Higher for OLS Regression – Page 9 .46.6.33\.Grades (GRADES) .7. * Now use corr2data to create a pseudosimulation of the data ...A composite score based on seven achievement tests.260 . Y .Scholastic Aptitude (APTD) .Parental Education (PARED) .26. Iraq.1.a dummy variable coded 1 if the respondent or both his parents were born in an Asian or North African country.Respondent's gradepoint average during the first trimester of eighth grade.590 .. matrix input SD = (...50000 .59.46.720 Grade Point Average .11000 1. matrix input Mean = (.59.00) .590 1.26\. input the means. Deviation .00.Ethnicity (SPHRD) .53.720 1. .00.46000 2.72.000 To create a pseudoreplication of this data in Stata.44. Shavit and Williams’ published analysis included the following information for students who were ability grouped in their classes.460 .33.
12 1. For example.122911 pared  . e.) This is why I call the data a pseudoreplication of the original.0000 As you can see.82 . and aptd you get results that are very close. Dev.74549  sphrd pared aptd grades +sphrd  1. These data will work fine for correlational/regression analysis where you analyze different sets of variables. corr sphrd pared aptd grades.57268 aptd  6. you may not be able to perfectly replicate published results. Min Max +sphrd  . and other reasons.667998 14. cases examined are not always the same throughout an analysis. subsample analyses. But. but not identical. * Label the variables . More critically. 10274 might be analyzed in another. sphrd.g. Simple rounding in the published results (e.46 2. label variable grades "Grade Point Average" . or you’ve entered the data wrong.46 1..3300 0. then the cases used to compute the correlations may be very different from the cases analyzed in that portion of the paper. If you get results that are very different.) Using Stata 9 and Higher for OLS Regression – Page 10 .11 1.092856 2. you should not try to analyze subsets of the data. or compute new variables. Even if you have entered the data correctly. if you regress grades on pared.04385 grades  7.0000 grades  0. THESE ARE NOT THE REAL DATA!!! The most obvious indication of this is that sphrd is not limited to values of 0 and 1.71.7200 1.4600 0. * Confirm that all is well .5900 1. You can now use the regress command to analyze these data. only reporting correlations to 2 or 3 decimal places) can cause slight differences. means (obs=10609) Variable  Mean Std.49931 11.g. label variable pared "Parental Education Scale" .339484 2.0000 aptd  0. (With SPSS. the new data set has the same means. 10609 cases might be analyzed in one regression. because of missing data.5 1. several cautions should be kept in mind.2600 0. to those reported by Shavit and Williams on p. HOWEVER. you simply input the means. (Either that.44 .5300 1. correlations and standard deviations as the original data. etc.42 2. Table 4 of their paper. recode variables. label variable sphrd "Sephardim" . correlations and standard deviations and it can handle things from there. an advantage of this is that it is more idiotproof than an analysis of data created by corr2data is.0000 pared  0. label variable aptd "Scholastic Aptitude" .
Once you have found a routine that sounds like what you want. however. type findit pcorr2 and then install it on your machine. You can also easily uninstall if you decide you do not want it. it is possible to write your own routines which add to the functionality of the program. I find it helpful to install the StataQuest routines. You might want to doublecheck results against SPSS or published findings. Stata 8 added a menustructure that made it more SPSSlike. Indeed. There are various other options for the regress command and several other postestimation commands that can be useful.) Other Comments. however. findit pcorr2 works. from within Stata. So. With Stata. (A possible complication is that you may find newer and older versions of the same command. I’ve often found that something that SPSS had and Stata did not could be added to Stata. type findit quest7 and install quest1. If it doesn’t have what you want already builtin. if the pcorr2 command was not working. Most of the routines I have installed seem to work fine. To see if such a routine exists. Usually a routine includes a brief description and you can view its help file. I suggest you try something like findit myprog to make sure somebody isn’t already using the name you had in mind. quest2 and quest3. check to make sure you are getting what you think you are getting. Unlike SPSS. income and INCOME would all be different variable names. I generally find it easier just to type the command in directly. Using Stata 9 and Higher for OLS Regression – Page 11 . SPSS is pretty much a closedended program. This can be very handy for commands you are not familiar with. Then. See my notes for graduate statistics II. Sometimes routines are part of a package of related routines and you install the entire package. you are out of luck. the command quest on will activate a Stataquest submenu under the User menu that has various helpful features. Income. A couple of cautions: User written routines are not officially supported by Stata. Stata does NOT have a builtin command for computing semipartial correlations. and you may even find two different commands with the same name. e. If a command works on one machine but not another. It will also give you FAQs and Stata help associated with the term. Stata is picky about case. For commands I know. If not already installed. it is entirely possible that such a routine has bugs or gives incorrect results. Among the things that will pop up is my very own pcorr2. but I have found a few problems. many such routines are publicly available and can be easily installed on your machine. which is an enhanced version of Stata’s pcorr command in that it computes both partial and semipartial correlations (pcorr only does partial correlations). from within Stata you can type findit semipartial The findit command will give you listings of programs that have the keyword semipartial associated with them. Although not as essential as they used to be. Findit pcorr2 does not. you can easily install it.Adding to Stata. it is probably because that command is not installed on both machines. at least under certain conditions. For example. Further.g. For example. If you ever write your own routine.
