0% found this document useful (0 votes)

634 views30 pages

Understanding Multiple Linear Regression

regresi linier berganda Regresi linier sederhana dan berganda [ edit ] https://upload.wikimedia.org/wikipedia/commons/thumb/3/3a/Linear_regress.svg/400px-Linear_regress.svg.png Contoh regresi linier sederhana , yang memiliki satu variabel independen Kasus yang paling sederhana dari variabel prediktor skalar tunggal x dan variabel respons skalar tunggal y dikenal sebagai regresi linier sederhana . Perpanjangan ke variabel prediktor berganda dan / atau bernilai vektor (dilambangkan dengan huru

Uploaded by

Jonesius Eden Manoppo

We take content rights seriously. If you suspect this is your content, claim it here.

Available Formats

Download as DOCX, PDF, TXT or read online on Scribd

0% found this document useful (0 votes)

634 views30 pages

Understanding Multiple Linear Regression

Uploaded by

Jonesius Eden Manoppo

We take content rights seriously. If you suspect this is your content, claim it here.

Available Formats

Download as DOCX, PDF, TXT or read online on Scribd

multiple linear regression

 Simple and multiple linear regression[edit]

Example of simple linear regression, which has one independent variable

The very simplest case of a single scalar predictor variable x and a single scalar response

variable y is known as simple linear regression. The extension to multiple and/or vector-valued

predictor variables (denoted with a capital X) is known as multiple linear regression, also

known as multivariable linear regression. Nearly all real-world regression models involve

multiple predictors, and basic descriptions of linear regression are often phrased in terms of the

multiple regression model. Note, however, that in these cases the response variable y is still a

scalar. Another term, multivariate linear regression, refers to cases where y is a vector, i.e., the

same as general linear regression.

 General linear models[edit]

The general linear model considers the situation when the response variable is not a scalar (for

each observation) but a vector, yi. Conditional linearity of is still assumed, with a

matrix B replacing the vector β of the classical linear regression model. Multivariate analogues

of ordinary least squares (OLS) and generalized least squares (GLS) have been developed.
"General linear models" are also called "multivariate linear models". These are not the same as

multivariable linear models (also called "multiple linear models").

 Heteroscedastic models[edit]

Various models have been created that allow for heteroscedasticity, i.e. the errors for different

response variables may have different variances. For example, weighted least squares is a

method for estimating linear regression models when the response variables may have different

error variances, possibly with correlated errors. (See also Weighted linear least squares,

and Generalized least squares.) Heteroscedasticity-consistent standard errors is an improved

method for use with uncorrelated but potentially heteroscedastic errors.

 Generalized linear models[edit]

Generalized linear models (GLMs) are a framework for modeling response variables that are

bounded or discrete. This is used, for example:

 when modeling positive quantities (e.g. prices or populations) that vary over a large

scale—which are better described using a skewed distribution such as the log-normal

distribution or Poisson distribution (although GLMs are not used for log-normal data,

instead the response variable is simply transformed using the logarithm function);

 when modeling categorical data, such as the choice of a given candidate in an election

(which is better described using a Bernoulli distribution/binomial distribution for

binary choices, or a categorical distribution/multinomial distribution for multi-way

choices), where there are a fixed number of choices that cannot be meaningfully

ordered;

 when modeling ordinal data, e.g. ratings on a scale from 0 to 5, where the different

outcomes can be ordered but where the quantity itself may not have any absolute

meaning (e.g. a rating of 4 may not be "twice as good" in any objective sense as a rating

of 2, but simply indicates that it is better than 2 or 3 but not as good as 5).
Generalized linear models allow for an arbitrary link function, g, that relates the mean of the

response variable(s) to the predictors: . The link function is often related to the distribution

of the response, and in particular it typically has the effect of transforming between

the range of the linear predictor and the range of the response variable.

Some common examples of GLMs are:

 Poisson regression for count data.

 Logistic regression and probit regression for binary data.

 Multinomial logistic regression and multinomial probit regression for categorical data.

 Ordered logit and ordered probit regression for ordinal data.

Single index models[clarification needed]

allow some degree of nonlinearity in the relationship

between x and y, while preserving the central role of the linear predictor β′x as in the classical

linear regression model. Under certain conditions, simply applying OLS to data from a single-

index model will consistently estimate β up to a proportionality constant.[11]

 Hierarchical linear models[edit]

Hierarchical linear models (or multilevel regression) organizes the data into a hierarchy of

regressions, for example where A is regressed on B, and B is regressed on C. It is often used

where the variables of interest have a natural hierarchical structure such as in educational

statistics, where students are nested in classrooms, classrooms are nested in schools, and

schools are nested in some administrative grouping, such as a school district. The response

variable might be a measure of student achievement such as a test score, and different

covariates would be collected at the classroom, school, and school district levels.

 Errors-in-variables[edit]

Errors-in-variables models (or "measurement error models") extend the traditional linear

regression model to allow the predictor variables X to be observed with error. This error causes
standard estimators of β to become biased. Generally, the form of bias is an attenuation,

meaning that the effects are biased toward zero.

 Others[edit]

 In Dempster–Shafer theory, or a linear belief function in particular, a linear regression

model may be represented as a partially swept matrix, which can be combined with

similar matrices representing observations and other assumed normal distributions and

state equations. The combination of swept or unswept matrices provides an alternative

method for estimating linear regression models.

Estimation methods[edit]

A large number of procedures have been developed for parameter estimation and inference in

linear regression. These methods differ in computational simplicity of algorithms, presence of

a closed-form solution, robustness with respect to heavy-tailed distributions, and theoretical

assumptions needed to validate desirable statistical properties such as consistency and

asymptotic efficiency.

Some of the more common estimation techniques for linear regression are summarized below.

 Least-squares estimation and related techniques[edit]

Francis Galton's 1875 illustration of the correlation between the heights of adults and their

parents. The observation that adult children's heights tended to deviate less from the mean

height than their parents suggested the concept of "regression toward the mean", giving

regression its name. The "locus of horizontal tangential points" passing through the leftmost

and rightmost points on the ellipse (which is a level curve of the bivariate normal

distribution estimated from the data) is the OLS estimate of the regression of parents' heights

on children's heights, while the "locus of vertical tangential points" is the OLS estimate of the

regression of children's heights on parent's heights. The major axis of the ellipse is

the TLS estimate.

Main article: Linear least squares

Linear least squares methods include mainly:

 Ordinary least squares

 Weighted least squares

 Generalized least squares

 Maximum-likelihood estimation and related techniques[edit]

 Maximum likelihood estimation can be performed when the distribution of the error

terms is known to belong to a certain parametric family ƒθ of probability

distributions.[12] When fθ is a normal distribution with zero mean and variance θ, the

resulting estimate is identical to the OLS estimate. GLS estimates are maximum

likelihood estimates when ε follows a multivariate normal distribution with a

known covariance matrix.

 Ridge regression[13][14][15] and other forms of penalized estimation, such as Lasso

regression,[5] deliberately introduce bias into the estimation of β in order to reduce

the variability of the estimate. The resulting estimates generally have lower mean

squared error than the OLS estimates, particularly when multicollinearity is present or

when overfitting is a problem. They are generally used when the goal is to predict the

value of the response variable y for values of the predictors x that have not yet been

observed. These methods are not as commonly used when the goal is inference, since

it is difficult to account for the bias.

 Least absolute deviation (LAD) regression is a robust estimation technique in that it

is less sensitive to the presence of outliers than OLS (but is less efficient than OLS

when no outliers are present). It is equivalent to maximum likelihood estimation under

a Laplace distribution model for ε.[16]

 Adaptive estimation. If we assume that error terms are independent of the

regressors, , then the optimal estimator is the 2-step MLE, where the first step is

used to non-parametrically estimate the distribution of the error term.[17]

 Other estimation techniques[edit]

Comparison of the Theil–Sen estimator (black) and simple linear regression (blue) for a set of

points with outliers.

 Bayesian linear regression applies the framework of Bayesian statistics to linear

regression. (See also Bayesian multivariate linear regression.) In particular, the

regression coefficients β are assumed to be random variables with a specified prior

distribution. The prior distribution can bias the solutions for the regression coefficients,

in a way similar to (but more general than) ridge regression or lasso regression. In

addition, the Bayesian estimation process produces not a single point estimate for the

"best" values of the regression coefficients but an entire posterior distribution,

completely describing the uncertainty surrounding the quantity. This can be used to

estimate the "best" coefficients using the mean, mode, median, any quantile

(see quantile regression), or any other function of the posterior distribution.

 Quantile regression focuses on the conditional quantiles of y given X rather than the

conditional mean of y given X. Linear quantile regression models a particular

conditional quantile, for example the conditional median, as a linear function βTx of the

predictors.

 Mixed models are widely used to analyze linear regression relationships involving

dependent data when the dependencies have a known structure. Common applications

of mixed models include analysis of data involving repeated measurements, such as

longitudinal data, or data obtained from cluster sampling. They are generally fit

as parametric models, using maximum likelihood or Bayesian estimation. In the case

where the errors are modeled as normal random variables, there is a close connection

between mixed models and generalized least squares.[18] Fixed effects estimation is an

alternative approach to analyzing this type of data.

 Principal component regression (PCR)[7][8] is used when the number of predictor

variables is large, or when strong correlations exist among the predictor variables. This

two-stage procedure first reduces the predictor variables using principal component

analysis then uses the reduced variables in an OLS regression fit. While it often works

well in practice, there is no general theoretical reason that the most informative linear

function of the predictor variables should lie among the dominant principal components

of the multivariate distribution of the predictor variables. The partial least squares

regression is the extension of the PCR method which does not suffer from the

mentioned deficiency.

 Least-angle regression[6] is an estimation procedure for linear regression models that

was developed to handle high-dimensional covariate vectors, potentially with more

covariates than observations.

 The Theil–Sen estimator is a simple robust estimation technique that chooses the

slope of the fit line to be the median of the slopes of the lines through pairs of sample

points. It has similar statistical efficiency properties to simple linear regression but is

much less sensitive to outliers.[19]

 Other robust estimation techniques, including the α-trimmed mean approach[citation

needed]
, and L-, M-, S-, and R-estimators have been introduced.[citation needed]

Applications[edit]

Main article: Trend estimation

A trend line represents a trend, the long-term movement in time series data after other

components have been accounted for. It tells whether a particular data set (say GDP, oil prices

or stock prices) have increased or decreased over the period of time. A trend line could simply

be drawn by eye through a set of data points, but more properly their position and slope is

calculated using statistical techniques like linear regression. Trend lines typically are straight

lines, although some variations use higher degree polynomials depending on the degree of

curvature desired in the line.

Trend lines are sometimes used in business analytics to show changes in data over time. This

has the advantage of being simple. Trend lines are often used to argue that a particular action

or event (such as training, or an advertising campaign) caused observed changes at a point in

time. This is a simple technique, and does not require a control group, experimental design, or

a sophisticated analysis technique. However, it suffers from a lack of scientific validity in cases

where other potential changes can affect the data.

 Epidemiology[edit]

Early evidence relating tobacco smoking to mortality and morbidity came from observational

studies employing regression analysis. In order to reduce spurious correlations when analyzing

observational data, researchers usually include several variables in their regression models in

addition to the variable of primary interest. For example, in a regression model in which

cigarette smoking is the independent variable of primary interest and the dependent variable is

lifespan measured in years, researchers might include education and income as additional
independent variables, to ensure that any observed effect of smoking on lifespan is not due to

those other socio-economic factors. However, it is never possible to include all

possible confounding variables in an empirical analysis. For example, a hypothetical gene

might increase mortality and also cause people to smoke more. For this reason, randomized

controlled trials are often able to generate more compelling evidence of causal relationships

than can be obtained using regression analyses of observational data. When controlled

experiments are not feasible, variants of regression analysis such as instrumental

variables regression may be used to attempt to estimate causal relationships from observational

data.

 Finance[edit]

The capital asset pricing model uses linear regression as well as the concept of beta for

analyzing and quantifying the systematic risk of an investment. This comes directly from the

beta coefficient of the linear regression model that relates the return on the investment to the

return on all risky assets.

 Economics[edit]

Main article: Econometrics

Linear regression is the predominant empirical tool in economics. For example, it is used to

predict consumption spending,[20] fixed investment spending, inventory investment, purchases

of a country's exports,[21] spending on imports,[21] the demand to hold liquid assets,[22] labor

demand,[23] and labor supply.[23]

 Environmental science[edit]

This section needs

expansion. You can help

by adding to it. (January

2010)
Linear regression finds application in a wide range of environmental science applications. In

Canada, the Environmental Effects Monitoring Program uses statistical analyses on fish

and benthic surveys to measure the effects of pulp mill or metal mine effluent on the aquatic

ecosystem.[24]

 Machine learning[edit]

Linear regression plays an important role in the field of artificial intelligence such as machine

learning. The linear regression algorithm is one of the fundamental supervised machine-

learning algorithms due to its relative simplicity and well-known properties.[25]

History[edit]

Least squares linear regression, as a means of finding a good rough linear fit to a set of points

was performed by Legendre (1805) and Gauss (1809) for the prediction of planetary

movement. Quetelet was responsible for making the procedure well-known and for using it

extensively in the social sciences.[26]

 Censored regression model

 Cross-sectional regression

 Curve fitting

 Empirical Bayes methods

 Errors and residuals

 Lack-of-fit sum of squares

 Line fitting

 Linear classifier
 Linear equation

 Logistic regression

 M-estimator

 Multivariate adaptive regression splines

 Nonlinear regression

 Nonparametric regression

 Normal equations

 Projection pursuit regression

 Segmented linear regression

 Stepwise regression

 Structural break

 Support vector machine

 Truncated regression model

References[edit]

 Citations[edit]

1. ^ David A. Freedman (2009). Statistical Models: Theory and

Practice. Cambridge University Press. p. 26. A simple regression equation has

on the right hand side an intercept and an explanatory variable with a slope

coefficient. A multiple regression equation has two or more explanatory

variables on the right hand side, each with its own slope coefficient

2. ^ Rencher, Alvin C.; Christensen, William F. (2012), "Chapter 10, Multivariate

regression – Section 10.1, Introduction", Methods of Multivariate Analysis,

Wiley Series in Probability and Statistics, 709 (3rd ed.), John Wiley & Sons,

p. 19, ISBN 9781118391679.

3. ^ Hilary L. Seal (1967). "The historical development of the Gauss linear

model". Biometrika. 54 (1/2): 1–24. doi:10.1093/biomet/54.1-

2.1. JSTOR 2333849.

4. ^ Yan, Xin (2009), Linear Regression Analysis: Theory and Computing, World

Scientific, pp. 1–2, ISBN 9789812834119, Regression analysis ... is probably

one of the oldest topics in mathematical statistics dating back to about two

hundred years ago. The earliest form of the linear regression was the least

squares method, which was published by Legendre in 1805, and by Gauss in

1809 ... Legendre and Gauss both applied the method to the problem of

determining, from astronomical observations, the orbits of bodies about the sun.

5. ^ Jump up to:a b Tibshirani, Robert (1996). "Regression Shrinkage and

Selection via the Lasso". Journal of the Royal Statistical Society, Series

B. 58 (1): 267–288. JSTOR 2346178.

6. ^ Jump up to:a b Efron, Bradley; Hastie, Trevor; Johnstone, Iain; Tibshirani,

Robert (2004). "Least Angle Regression". The Annals of Statistics. 32 (2): 407–

451. arXiv:math/0406456. doi:10.1214/009053604000000067. JSTOR 344846

7. ^ Jump up to:a b Hawkins, Douglas M. (1973). "On the Investigation of

Alternative Regressions by Principal Component Analysis". Journal of the

Royal Statistical Society, Series C. 22 (3): 275–286. JSTOR 2346776.

8. ^ Jump up to:a b Jolliffe, Ian T. (1982). "A Note on the Use of Principal

Components in Regression". Journal of the Royal Statistical Society, Series

C. 31 (3): 300–303. JSTOR 2348005.

9. ^ Berk, Richard A. (2007). "Regression Analysis: A Constructive

Critique". Criminal Justice Review. 32 (3): 301–

302. doi:10.1177/0734016807304871.

10. ^ Warne, Russell T. (2011). "Beyond multiple regression: Using commonality

analysis to better understand R2 results". Gifted Child Quarterly. 55 (4): 313–

318. doi:10.1177/0016986211422217.

11. ^ Brillinger, David R. (1977). "The Identification of a Particular Nonlinear

Time Series System". Biometrika. 64 (3): 509–

515. doi:10.1093/biomet/64.3.509. JSTOR 2345326.

12. ^ Lange, Kenneth L.; Little, Roderick J. A.; Taylor, Jeremy M. G.

(1989). "Robust Statistical Modeling Using the t Distribution" (PDF). Journal

of the American Statistical Association. 84 (408): 881–

896. doi:10.2307/2290063. JSTOR 2290063.

13. ^ Swindel, Benee F. (1981). "Geometry of Ridge Regression Illustrated". The

American Statistician. 35 (1): 12–15. doi:10.2307/2683577. JSTOR 2683577.

14. ^ Draper, Norman R.; van Nostrand; R. Craig (1979). "Ridge Regression and

James-Stein Estimation: Review and Comments". Technometrics. 21 (4): 451–

466. doi:10.2307/1268284. JSTOR 1268284.

15. ^ Hoerl, Arthur E.; Kennard, Robert W.; Hoerl, Roger W. (1985). "Practical

Use of Ridge Regression: A Challenge Met". Journal of the Royal Statistical

Society, Series C. 34 (2): 114–120. JSTOR 2347363.

16. ^ Narula, Subhash C.; Wellington, John F. (1982). "The Minimum Sum of

Absolute Errors Regression: A State of the Art Survey". International Statistical

Review. 50(3): 317–326. doi:10.2307/1402501. JSTOR 1402501.

17. ^ Stone, C. J. (1975). "Adaptive maximum likelihood estimators of a location

parameter". The Annals of Statistics. 3 (2): 267–

284. doi:10.1214/aos/1176343056. JSTOR 2958945.

18. ^ Goldstein, H. (1986). "Multilevel Mixed Linear Model Analysis Using

Iterative Generalized Least Squares". Biometrika. 73 (1): 43–

56. doi:10.1093/biomet/73.1.43. JSTOR 2336270.

19. ^ Theil, H. (1950). "A rank-invariant method of linear and polynomial

regression analysis. I, II, III". Nederl. Akad. Wetensch., Proc. 53: 386–392,

521–525, 1397–1412. MR 0036489; Sen, Pranab Kumar (1968). "Estimates of

the regression coefficient based on Kendall's tau". Journal of the American

Statistical Association. 63 (324): 1379–

1389. doi:10.2307/2285891. JSTOR 2285891. MR 0258201.

20. ^ Deaton, Angus (1992). Understanding Consumption. Oxford University

Press. ISBN 978-0-19-828824-4.

21. ^ Jump up to:a b Krugman, Paul R.; Obstfeld, M.; Melitz, Marc

J. (2012). International Economics: Theory and Policy (9th global ed.).

Harlow: Pearson. ISBN 9780273754091.

22. ^ Laidler, David E. W. (1993). The Demand for Money: Theories, Evidence, and

Problems (4th ed.). New York: Harper Collins. ISBN 978-0065010985.

23. ^ Jump up to:a b Ehrenberg; Smith (2008). Modern Labor Economics (10th

international ed.). London: Addison-Wesley. ISBN 9780321538963.

24. ^ EEMP webpage Archived 2011-06-11 at the Wayback Machine

25. ^ "Linear Regression (Machine Learning)" (PDF). University of Pittsburgh.

26. ^ Stigler, Stephen M. (1986). The History of Statistics: The Measurement of

Uncertainty before 1900. Cambridge: Harvard. ISBN 0-674-40340-1.

 Sources[edit]

 Cohen, J., Cohen P., West, S.G., & Aiken, L.S. (2003). Applied multiple

regression/correlation analysis for the behavioral sciences. (2nd ed.) Hillsdale, NJ:

Lawrence Erlbaum Associates

 Charles Darwin. The Variation of Animals and Plants under Domestication.

(1868) (Chapter XIII describes what was known about reversion in Galton's time.

Darwin uses the term "reversion".)

 Draper, N.R.; Smith, H. (1998). Applied Regression Analysis (3rd ed.). John

Wiley. ISBN 978-0-471-17082-2.

 Francis Galton. "Regression Towards Mediocrity in Hereditary Stature," Journal of the

Anthropological Institute, 15:246-263 (1886). (Facsimile at: [1])

 Robert S. Pindyck and Daniel L. Rubinfeld (1998, 4h ed.). Econometric Models and

Economic Forecasts, ch. 1 (Intro, incl. appendices on Σ operators & derivation of

parameter est.) & Appendix 4.3 (mult. regression in matrix form).

 Pedhazur, Elazar J (1982). Multiple regression in behavioral research: Explanation

and prediction (2nd ed.). New York: Holt, Rinehart and Winston. ISBN 978-0-03-

041760-3.

 Mathieu Rouaud, 2013: Probability, Statistics and Estimation Chapter 2: Linear

Regression, Linear Regression with Error Bars and Nonlinear Regression.

 National Physical Laboratory (1961). "Chapter 1: Linear Equations and Matrices:

Direct Methods". Modern Computing Methods. Notes on Applied Science. 16 (2nd

ed.). Her Majesty's Stationery Office

regresi linier berganda

 Regresi linier sederhana dan berganda [ edit ]
Contoh regresi linier sederhana , yang memiliki satu variabel independen

Kasus yang paling sederhana dari variabel prediktor skalar tunggal x dan variabel respons

skalar tunggal y dikenal sebagai regresi linier sederhana . Perpanjangan ke variabel

prediktor berganda dan / atau bernilai vektor (dilambangkan dengan huruf kapital X ) dikenal

sebagai regresi linier berganda , juga dikenal sebagai regresi linier multivariabel . Hampir

semua model regresi dunia nyata melibatkan banyak prediktor, dan deskripsi dasar regresi

linier sering diungkapkan dalam istilah model regresi berganda. Namun, perhatikan bahwa

dalam kasus-kasus ini variabel respons y masih berupa skalar. Istilah lain, regresi

linier multivariat , mengacu pada kasus-kasus di mana y adalah vektor, yaitu, sama

dengan regresi linier umum .

 Model linier umum [ sunting ]

The model linier umum menganggap situasi saat variabel respon tidak skalar (untuk setiap

pengamatan), tetapi vektor, y i . Linearitas kondisional masih diasumsikan, dengan matriks B

yang menggantikan vektor β dari model regresi linier klasik. Analog multivariat dari kuadrat

terkecil biasa (OLS) dan kuadrat terkecil umum (GLS) telah dikembangkan. "Model linear

umum" juga disebut "model linier multivarian". Ini tidak sama dengan model linier

multivariabel (juga disebut "model linier berganda").

 Model heteroskedastik [ sunting ]

Berbagai model telah dibuat yang memungkinkan heteroskedastisitas , yaitu kesalahan untuk

variabel respons yang berbeda mungkin memiliki varian yang berbeda . Sebagai
contoh, kuadrat terkecil tertimbang adalah metode untuk memperkirakan model regresi linier
ketika variabel respon mungkin memiliki varian kesalahan yang berbeda, mungkin dengan

kesalahan yang berkorelasi. (Lihat juga kuadrat terkecil linier tertimbang , dan kuadrat

terkecil umum .) Kesalahan standar yang konsisten heteroskedastisitas adalah metode yang

ditingkatkan untuk digunakan dengan kesalahan tidak berkorelasi tetapi berpotensi

heteroskedastik.
 Model linear umum [ sunting ]

Generalized linear models (GLMs) adalah kerangka kerja untuk memodelkan variabel respon

yang dibatasi atau diskrit. Ini digunakan, misalnya:

 ketika memodelkan jumlah positif (mis. harga atau populasi) yang bervariasi dalam
skala besar — yang lebih baik dideskripsikan dengan menggunakan distribusi
miring seperti distribusi log-normal atau distribusi Poisson (meskipun GLM tidak
digunakan untuk data log-normal, alih-alih responsnya variabel hanya ditransformasikan
menggunakan fungsi logaritma);
 ketika memodelkan data kategorikal , seperti pilihan kandidat yang diberikan dalam
pemilihan (yang lebih baik dijelaskan menggunakan distribusi Bernoulli / distribusi
binomial untuk pilihan biner, atau distribusi kategorikal / distribusi multinomial untuk
pilihan multi-arah), di mana ada sejumlah pilihan tetap yang tidak dapat dipesan secara
bermakna;
 ketika memodelkan data ordinal , misalnya peringkat pada skala dari 0 hingga 5, di
mana hasil yang berbeda dapat dipesan tetapi di mana kuantitas itu sendiri mungkin tidak
memiliki makna absolut (misalnya peringkat 4 mungkin tidak "dua kali lebih baik" dalam
tujuan apa pun) akal sebagai peringkat 2, tetapi hanya menunjukkan bahwa itu lebih baik
dari 2 atau 3 tetapi tidak sebagus 5).

Generalized linear model memungkinkan untuk sewenang-wenang fungsi link , g , yang

berhubungan dengan rata-rata dari variabel respon (s) ke prediktor: . Fungsi tautan sering

dikaitkan dengan distribusi respons, dan khususnya, ia biasanya memiliki efek transformasi

antara rentang prediktor linier dan rentang variabel respons.

Beberapa contoh umum GLM adalah:

 Regresi poisson untuk menghitung data.
 Regresi logistik dan regresi probit untuk data biner.
 Regresi logistik multinomial dan regresi probit multinomial untuk data kategorikal.
 Memerintahkan logit dan memerintahkan regresi probit untuk data ordinal.
Model indeks tunggal [ klarifikasi diperlukan ] memungkinkan beberapa derajat nonlinier dalam hubungan

antara x dan y , sambil mempertahankan peran sentral dari prediktor linier β ′ x seperti pada

model regresi linier klasik. Dalam kondisi tertentu, hanya menerapkan OLS ke data dari model

indeks tunggal akan secara konsisten memperkirakan β hingga konstanta

proporsionalitas. [11]
 Model linear hierarkis [ sunting ]

Model hirarkis linear (atau regresi bertingkat ) mengatur data ke dalam hirarki regresi,

misalnya di mana A adalah kemunduran pada B , dan B adalah kemunduran pada C . Ini sering

digunakan di mana variabel yang diminati memiliki struktur hierarki alami seperti dalam

statistik pendidikan, di mana siswa bersarang di ruang kelas, ruang kelas bersarang di sekolah,

dan sekolah bersarang dalam beberapa pengelompokan administratif, seperti distrik

sekolah. Variabel respon mungkin merupakan ukuran pencapaian siswa seperti skor tes, dan

kovariat yang berbeda akan dikumpulkan di tingkat kelas, sekolah, dan distrik sekolah.
 Kesalahan-dalam-variabel [ sunting ]

Model kesalahan-dalam-variabel (atau "model kesalahan pengukuran") memperluas model

regresi linier tradisional untuk memungkinkan variabel prediktor X diamati dengan

kesalahan. Kesalahan ini menyebabkan penaksir standar β menjadi bias. Secara umum, bentuk

bias adalah pelemahan, yang berarti bahwa efeknya bias ke nol.

 Lainnya [ edit ]
 Dalam teori Dempster-Shafer , atau fungsi kepercayaan linier khususnya, model
regresi linier dapat direpresentasikan sebagai matriks sapuan sebagian, yang dapat
dikombinasikan dengan matriks serupa yang mewakili pengamatan dan distribusi normal
yang diasumsikan lainnya serta persamaan keadaan. Kombinasi matriks tersapu atau tidak
tersapu menyediakan metode alternatif untuk memperkirakan model regresi linier.

Metode estimasi [ sunting ]

Sejumlah besar prosedur telah dikembangkan untuk estimasi parameter dan inferensi dalam

regresi linier. Metode-metode ini berbeda dalam kesederhanaan komputasi algoritma,

keberadaan solusi bentuk tertutup, ketahanan terhadap distribusi berekor berat, dan asumsi

teoritis yang diperlukan untuk memvalidasi sifat statistik yang diinginkan

seperti konsistensi dan efisiensi asimptotik .
Beberapa teknik estimasi yang lebih umum untuk regresi linier dirangkum di bawah ini.
 Estimasi kuadrat terkecil dan teknik terkait [ edit ]

Ilustrasi Francis Galton tahun 1875 tentang korelasi antara ketinggian orang dewasa dan orang

tua mereka. Pengamatan bahwa ketinggian anak-anak dewasa cenderung menyimpang kurang

dari tinggi rata-rata daripada orang tua mereka menyarankan konsep " regresi menuju rata-

rata ", memberikan regresi namanya. "Lokus titik tangensial horizontal" melewati titik paling

kiri dan paling kanan pada elips (yang merupakan kurva level dari distribusi normal bivariat

yang diperkirakan dari data) adalah estimasi OLS dari regresi ketinggian orang tua pada

ketinggian anak-anak, sementara "lokus titik tangensial vertikal" adalah perkiraan OLS dari

regresi ketinggian anak-anak pada ketinggian orangtua. Sumbu utama elips

adalah estimasi TLS .

Artikel utama: Linear least square

Metode linear kuadrat terkecil meliputi:

 Kuadrat terkecil biasa
 Kuadrat terkecil tertimbang
 Kuadrat terkecil umum
 Estimasi kemungkinan maksimum dan teknik terkait [ edit ]
 Estimasi kemungkinan maksimum dapat dilakukan ketika distribusi istilah kesalahan
diketahui milik parametrik keluarga tertentu ƒ θ dari distribusi
probabilitas . [12] Ketika f θ adalah distribusi normal dengan nol mean dan varians θ,
estimasi yang dihasilkan identik dengan estimasi OLS. Estimasi GLS adalah estimasi
kemungkinan maksimum ketika ε mengikuti distribusi normal multivariat dengan matriks
kovarian yang diketahui .
 Regresi punggungan [13] [14] [15] dan bentuk-bentuk lain dari estimasi hukuman,
seperti regresi Lasso , [5] sengaja memasukkan bias ke dalam estimasi β untuk
mengurangi variabilitas estimasi. Perkiraan yang dihasilkan umumnya
memiliki kesalahan kuadrat rata-rata lebih rendah dari perkiraan OLS, terutama
ketika multikolinieritas hadir atau ketika overfitting adalah masalah. Mereka umumnya
digunakan ketika tujuannya adalah untuk memprediksi nilai variabel respon y untuk nilai-
nilai prediktor x yang belum diamati. Metode-metode ini tidak seperti yang biasa
digunakan ketika tujuannya adalah inferensi, karena sulit untuk memperhitungkan
bias.
 Regresi minimum absolut (LAD) adalah teknik estimasi kuat karena kurang sensitif
terhadap kehadiran outlier daripada OLS (tetapi kurang efisien daripada OLS ketika tidak
ada outlier yang hadir). Ini sama dengan estimasi kemungkinan maksimum di
bawah model distribusi Laplace untuk ε . [16]
 Estimasi adaptif . Jika kita mengasumsikan bahwa istilah kesalahan
tidak tergantung pada regressor,, maka penaksir optimal adalah MLE 2 langkah, di mana
langkah pertama digunakan untuk secara non-parametrik memperkirakan distribusi istilah

kesalahan. [17]
 Teknik estimasi lain [ sunting ]
Perbandingan estimator Theil-Sen (hitam) dan regresi linier sederhana (biru) untuk satu set

poin dengan outlier.

 Regresi linier Bayesian menerapkan kerangka kerja statistik Bayesian terhadap
regresi linier. (Lihat juga regresi linear multivariat Bayesian .) Secara khusus, koefisien
regresi β diasumsikan sebagai variabel acak dengan distribusi sebelumnya
yang ditentukan . Distribusi sebelumnya dapat bias solusi untuk koefisien regresi, dengan
cara yang mirip dengan (tetapi lebih umum daripada) regresi ridge atau regresi
laso . Selain itu, proses estimasi Bayesian menghasilkan bukan hanya satu titik estimasi
untuk nilai "terbaik" dari koefisien regresi tetapi seluruh distribusi posterior , sepenuhnya
menggambarkan ketidakpastian seputar kuantitas. Ini dapat digunakan untuk
memperkirakan koefisien "terbaik" menggunakan mean, mode, median, semua kuantil
(lihat regresi kuantil ), atau fungsi lain dari distribusi posterior.
 Regresi kuantil berfokus pada quantiles bersyarat dari y diberikan X daripada mean
bersyarat dari y diberikan X . Linier model regresi kuantil suatu kuantil tertentu bersyarat,
misalnya median bersyarat, sebagai fungsi β linear T x prediktor.
 Model campuran banyak digunakan untuk menganalisis hubungan regresi linier yang
melibatkan data dependen ketika dependensi memiliki struktur yang diketahui. Aplikasi
umum dari model campuran meliputi analisis data yang melibatkan pengukuran berulang,
seperti data longitudinal, atau data yang diperoleh dari cluster sampling. Mereka
umumnya cocok sebagai model parametrik , menggunakan kemungkinan maksimum
atau estimasi Bayesian. Dalam kasus di mana kesalahan dimodelkan sebagai variabel
acak normal , ada hubungan dekat antara model campuran dan kuadrat terkecil
umum. [18] Estimasi efek tetap adalah pendekatan alternatif untuk menganalisis jenis data
ini.
 Regresi komponen utama (PCR) [7] [8] digunakan ketika jumlah variabel prediktor
besar, atau ketika korelasi kuat ada di antara variabel prediktor. Prosedur dua tahap ini
pertama-tama mengurangi variabel prediktor menggunakan analisis komponen
utama kemudian menggunakan variabel yang dikurangi dalam kecocokan regresi
OLS. Meskipun sering bekerja dengan baik dalam praktiknya, tidak ada alasan teoritis
umum bahwa fungsi linear paling informatif dari variabel prediktor harus terletak di
antara komponen utama yang dominan dari distribusi multivariat dari variabel
prediktor. The regresi kuadrat terkecil parsial merupakan perpanjangan dari metode PCR
yang tidak menderita dari kekurangan yang disebutkan.
 Regresi sudut terkecil [6] adalah prosedur estimasi untuk model regresi linier yang
dikembangkan untuk menangani vektor kovariat dimensi tinggi, berpotensi dengan lebih
banyak kovariat daripada pengamatan.
 The Theil-Sen estimator adalah sederhana estimasi kuat teknik yang memilih
kemiringan garis cocok menjadi median dari lereng garis melalui pasang titik sampel. Ini
memiliki sifat efisiensi statistik yang mirip dengan regresi linier sederhana tetapi jauh
lebih sensitif terhadap pencilan . [19]
 Teknik estimasi kuat lainnya, termasuk pendekatan rata-rata terpangkas [ rujukan? ] ,
Dan estimator L-, M-, S-, dan R- telah diperkenalkan. [ rujukan? ]

Aplikasi [ sunting ]

Lihat juga: Linear least square § Aplikasi

Regresi linier banyak digunakan dalam ilmu biologi, perilaku dan sosial untuk menggambarkan

kemungkinan hubungan antar variabel. Ini peringkat sebagai salah satu alat paling penting yang

digunakan dalam disiplin ilmu ini.

 Garis tren [ edit ]

Artikel utama: Estimasi tren

Garis tren mewakili tren, pergerakan jangka panjang dalam data deret waktu setelah

komponen lainnya diperhitungkan. Ini memberitahu apakah kumpulan data tertentu

(katakanlah PDB, harga minyak atau harga saham) telah meningkat atau menurun selama

periode waktu tertentu. Garis tren dapat dengan mudah ditarik melalui serangkaian titik data,

tetapi lebih tepat posisi dan kemiringannya dihitung menggunakan teknik statistik seperti

regresi linier. Garis tren biasanya adalah garis lurus, meskipun beberapa variasi menggunakan

polinomial tingkat tinggi tergantung pada tingkat kelengkungan yang diinginkan dalam garis

tersebut.

Garis tren terkadang digunakan dalam analitik bisnis untuk menunjukkan perubahan data

seiring waktu. Ini memiliki keuntungan karena sederhana. Garis tren sering digunakan untuk

menyatakan bahwa tindakan atau peristiwa tertentu (seperti pelatihan, atau kampanye iklan)

menyebabkan perubahan yang diamati pada suatu titik waktu. Ini adalah teknik sederhana, dan

tidak memerlukan kelompok kontrol, desain eksperimental, atau teknik analisis yang
canggih. Namun, itu menderita dari kurangnya validitas ilmiah dalam kasus di mana perubahan

potensial lainnya dapat mempengaruhi data.

 Epidemiologi [ sunting ]

Bukti awal terkait merokok tembakau dengan mortalitas dan morbiditas berasal dari studi

observasional menggunakan analisis regresi. Untuk mengurangi korelasi palsu ketika

menganalisis data pengamatan, peneliti biasanya memasukkan beberapa variabel dalam model

regresi mereka di samping variabel minat utama. Misalnya, dalam model regresi di mana

merokok merupakan variabel independen yang menjadi perhatian utama dan variabel dependen

adalah umur yang diukur dalam tahun, para peneliti dapat memasukkan pendidikan dan

pendapatan sebagai variabel independen tambahan , untuk memastikan bahwa setiap efek yang

diamati dari merokok pada umur adalah bukan karena faktor-faktor sosial ekonomi

lainnya . Namun, tidak pernah mungkin untuk memasukkan semua variabel pengganggu

yang mungkin dalam analisis empiris. Misalnya, gen hipotetis dapat meningkatkan angka

kematian dan juga menyebabkan orang lebih banyak merokok. Untuk alasan ini, uji coba

terkontrol secara acak seringkali dapat menghasilkan bukti yang lebih kuat dari hubungan

kausal daripada yang bisa diperoleh dengan menggunakan analisis regresi data

pengamatan. Ketika eksperimen terkontrol tidak layak, varian analisis regresi seperti variabel

instrumental dapat digunakan untuk mencoba memperkirakan hubungan sebab akibat dari data

pengamatan.
 Keuangan [ edit ]

The capital asset pricing model menggunakan regresi linear serta konsep beta untuk

menganalisis dan mengukur risiko sistematis investasi. Ini datang langsung dari koefisien beta

dari model regresi linier yang menghubungkan pengembalian investasi dengan pengembalian

semua aset berisiko.

 Ekonomi [ sunting ]

Artikel utama: Ekonometrika

Regresi linier adalah alat empiris yang dominan dalam bidang ekonomi . Misalnya, digunakan
untuk memprediksi pengeluaran konsumsi , [20] investasi tetap belanja, investasi persediaan ,
pembelian suatu negara ekspor , [21] pengeluaran impor , [21] yang permintaan untuk memegang

aset likuid , [22] permintaan tenaga kerja , [23] dan pasokan tenaga kerja . [23]
 Ilmu lingkungan [ sunting ]

Bagian ini membutuhkan

ekspansi . Anda dapat membantu

dengan menambahkannya . (Januari

2010)

Regresi linier menemukan aplikasi dalam berbagai aplikasi ilmu lingkungan. Di Kanada,

Program Pemantauan Efek Lingkungan menggunakan analisis statistik pada ikan

dan survei bentik untuk mengukur efek dari pabrik pulp atau limbah tambang logam pada

ekosistem perairan. [24]

 Pembelajaran mesin [ edit ]

Regresi linier memainkan peran penting dalam bidang kecerdasan buatan seperti pembelajaran

mesin . Algoritma regresi linier adalah salah satu algoritma pembelajaran mesin yang

diawasi karena kesederhanaannya dan sifat-sifatnya yang terkenal. [25]

Sejarah [ sunting ]

Regresi linear kuadrat terkecil, sebagai sarana untuk menemukan kesesuaian linear kasar yang

baik dengan sekumpulan poin dilakukan oleh Legendre (1805) dan Gauss (1809) untuk

prediksi pergerakan planet. Quetelet bertanggung jawab untuk membuat prosedur ini terkenal

dan menggunakannya secara luas dalam ilmu sosial. [26]

Lihat juga [ edit ]

 Portal statistik
 Analisis varian
 Dekomposisi Blinder – Oaxaca
 Model regresi yang disensor
 Regresi cross-sectional
 Kurva pas
 Metode empiris Bayes
 Kesalahan dan residu
 Jumlah kotak yang tidak sesuai
 Pemasangan garis
 Klasifikasi linier
 Persamaan linier
 Regresi logistik
 M-estimator
 Splines regresi adaptif multivariat
 Regresi nonlinier
 Regresi nonparametrik
 Persamaan normal
 Regresi pengejaran proyeksi
 Regresi linier tersegmentasi
 Regresi bertahap
 Istirahat struktural
 Mesin dukungan vektor
 Model regresi terpotong

Referensi [ sunting ]
 Kutipan [ sunting ]

1. ^ David A. Freedman (2009). Model Statistik: Teori dan

Praktek. Cambridge University Press . hal. 26. Persamaan regresi sederhana

memiliki di sisi kanan sebuah intersep dan variabel penjelas dengan koefisien

kemiringan. Persamaan regresi berganda memiliki dua atau lebih variabel

penjelas di sisi kanan, masing-masing dengan koefisien kemiringan sendiri

2. ^ Rencher, Alvin C.; Christensen, William F. (2012), "Bab 10, Regresi

multivariat - Bagian 10.1, Pendahuluan", Metode Analisis Multivariat , Seri

Wiley dalam Probabilitas dan Statistik, 709 (edisi ketiga), John Wiley & Sons,

p. 19, ISBN 9781118391679 .

3. ^ Hilary L. Seal (1967). "Perkembangan historis model linear

Gauss". Biometrika. 54 (1/2): 1–24. doi : 10.1093 / biomet / 54.1-

2.1 . JSTOR 2333849 .

4. ^ Yan, Xin (2009), Analisis Regresi Linier: Teori dan Komputasi ,

World Scientific, hlm. 1–2, ISBN 9789812834119 , Analisis regresi ... mungkin
salah satu topik tertua dalam statistik matematika sejak sekitar dua ratus tahun

yang lalu. lalu. Bentuk paling awal dari regresi linier adalah metode kuadrat

terkecil, yang diterbitkan oleh Legendre pada tahun 1805, dan oleh Gauss pada

tahun 1809 ... Legendre dan Gauss keduanya menerapkan metode ini pada

masalah penentuan, dari pengamatan astronomi, orbit tubuh. tentang

matahari.

5. ^ Melompat ke: a b Tibshirani, Robert (1996). "Penyusutan dan

Pemilihan Regresi melalui Lasso". Jurnal Masyarakat Statistik Kerajaan, Seri

B. 58 (1): 267–288. JSTOR 2346178 .

6. ^ Melompat ke: a b Efron, Bradley; Hastie, Trevor; Johnstone,

Iain; Tibshirani, Robert (2004). "Regresi Sudut Paling Sedikit". The Annals of

Statistics. 32 (2): 407–451. arXiv : matematika / 0406456 . doi : 10.1214 /

009053604000000067 . JSTOR 3448465 .

7. ^ Melompat ke: a b Hawkins, Douglas M. (1973). "Tentang Investigasi

Regresi Alternatif oleh Analisis Komponen Utama". Jurnal Masyarakat Statistik

Kerajaan, Seri C. 22 (3): 275–286. JSTOR 2346776 .

8. ^ Melompat ke: a b Jolliffe, Ian T. (1982). "Catatan tentang Penggunaan

Komponen Utama dalam Regresi". Jurnal Masyarakat Statistik Kerajaan, Seri

C. 31 (3): 300–303. JSTOR 2348005 .

9. ^ Berk, Richard A. (2007). "Analisis Regresi: Kritik

Konstruktif". Tinjauan Keadilan Pidana. 32 (3): 301–302. doi : 10.1177 /

0734016807304871 .

10. ^ Warne, Russell T. (2011). "Melampaui regresi berganda:

Menggunakan analisis kesamaan untuk lebih memahami hasil R2". Triwulan

Anak Berbakat. 55 (4): 313–318. doi : 10.1177 / 0016986211422217 .

11. ^ Brillinger, David R. (1977). "Identifikasi Sistem Seri Waktu Nonlinier

Tertentu". Biometrika. 64 (3): 509–515. doi : 10.1093 / biomet /

64.3.509 . JSTOR 2345326 .

12. ^ Lange, Kenneth L .; Sedikit, Roderick JA; Taylor, Jeremy MG

(1989). "Pemodelan Statistik yang Kuat Menggunakan t

Distribusi" (PDF). Jurnal Asosiasi Statistik Amerika. 84 (408): 881–

896. doi : 10.2307 / 2290063 . JSTOR 2290063 .

13. ^ Swindel, Benee F. (1981). "Geometry of Ridge Regression

Illustrated". Ahli Statistik Amerika. 35 (1): 12–15. doi : 10.2307 /

2683577 . JSTOR 2683577 .

14. ^ Draper, Norman R .; van Nostrand; R. Craig (1979). "Regresi Ridge

dan Estimasi James-Stein: Ulasan dan Komentar". Technometrics. 21 (4):

451–466. doi : 10.2307 / 1268284 . JSTOR 1268284 .

15. ^ Hoerl, Arthur E .; Kennard, Robert W .; Hoerl, Roger W.

(1985). "Penggunaan Praktis Regresi Punggung: Sebuah Tantangan

Ditemui". Jurnal Masyarakat Statistik Kerajaan, Seri C. 34 (2): 114–

120. JSTOR 2347363 .

16. ^ Narula, Subhash C .; Wellington, John F. (1982). "Jumlah Minimum

Regresi Kesalahan Mutlak: Survei Keadaan Mutakhir". Tinjauan Statistik

Internasional. 50 (3): 317–326. doi : 10.2307 /

1402501 . JSTOR 1402501 .

17. ^ Stone, CJ (1975). + Msgstr "Penduga kemungkinan maksimum

adaptif dari parameter lokasi". The Annals of Statistics. 3 (2): 267–

284. doi : 10.1214 / aos / 1176343056 . JSTOR 2958945 .

18. ^ Goldstein, H. (1986). "Analisis Multilevel Mixed Linear Model

Menggunakan Iterative Generalized Least Squares". Biometrika. 73 (1): 43–

56. doi : 10.1093 / biomet / 73.1.43 . JSTOR 2336270 .

19. ^ Theil, H. (1950). "Metode analisis regresi linear dan polinomial

peringkat-invarian. I, II, III". Nederl. Akad. Wetensch., Proc. 53 : 386–392,

521–525, 1397–1412. MR 0036489 ; Sen, Pranab Kumar (1968). "Perkiraan

koefisien regresi berdasarkan pada Kendall's tau". Jurnal Asosiasi Statistik

Amerika . 63 (324): 1379–1389. doi : 10.2307 /

2285891 . JSTOR 2285891 . MR 0258201 .

20. ^ Deaton, Angus (1992). Memahami Konsumsi. Oxford University

Press. ISBN 978-0-19-828824-4 .

21. ^ Melompat ke: a b Krugman, Paul R .; Obstfeld, M .; Melitz, Marc

J. (2012). Ekonomi Internasional: Teori dan Kebijakan (edisi ke-9

global). Harlow: Pearson. ISBN 9780273754091 .

22. ^ Laidler, David EW (1993). Permintaan Uang: Teori, Bukti, dan

Masalah (edisi ke-4). New York: Harper Collins. ISBN 978-0065010985 .

23. ^ Melompat ke: a b Ehrenberg; Smith (2008). Ekonomi Perburuhan

Modern (edisi internasional ke-10). London: Addison-

Wesley. ISBN 9780321538963 .

24. ^ Laman web EEMP Diarsipkan 2011-06-11 di Wayback Machine

25. ^ "Regresi Linier (Pembelajaran Mesin)" (PDF). Universitas

Pittsburgh.

26. ^ Stigler, Stephen M. (1986). Sejarah Statistik: Pengukuran

Ketidakpastian sebelum 1900 . Cambridge: Harvard. ISBN 0-674-40340-1 .

 Sumber [ sunting ]
 Cohen, J., Cohen P., West, SG, & Aiken, LS (2003). Analisis regresi / korelasi
berganda diterapkan untuk ilmu perilaku . (2nd ed.) Hillsdale, NJ: Lawrence Erlbaum
Associates
 Charles Darwin . Variasi Hewan dan Tumbuhan di Bawah Domestikasi . (1868) (Bab
XIII menjelaskan apa yang diketahui tentang pembalikan pada zaman Galton. Darwin
menggunakan istilah "pembalikan".)
 Draper, NR; Smith, H. (1998). Analisis Regresi Terapan (edisi ke-3). John
Wiley. ISBN 978-0-471-17082-2 .
 Francis Galton. "Regresi Menuju Mediokritas dalam Perawakan Turunan," Jurnal
Institut Antropologi , 15: 246-263 (1886). (Faksimili di: [1] )
 Robert S. Pindyck dan Daniel L. Rubinfeld (1998, edisi 4 jam). Model Ekonometrik
dan Prakiraan Ekonomi , ch. 1 (Pendahuluan, termasuk lampiran pada & operator &
derivasi parameter est.) & Lampiran 4.3 (mult. Regresi dalam bentuk matriks).
Bacaan lebih lanjut [ sunting ]
 Pedhazur, Elazar J (1982). Regresi berganda dalam penelitian perilaku: Penjelasan
dan prediksi (edisi kedua). New York: Holt, Rinehart dan Winston. ISBN 978-0-03-
041760-3 .
 Mathieu Rouaud, 2013: Probabilitas, Statistik dan Estimasi Bab 2: Regresi Linier,
Regresi Linier dengan Bar Kesalahan dan Regresi Nonlinear.
 Laboratorium Fisik Nasional (1961). "Bab 1: Persamaan Linear dan Matriks: Metode
Langsung". Metode Komputasi Modern. Catatan tentang Sains Terapan. 16 (2nd
ed.). Kantor Alat Tulis Yang Mulia

(https://en.wikipedia.org/wiki/File:Linear_regression.svg)multiple linear regression

Simple and multiple linear regressi

"General linear models" are also called "multivariate linear models". These are not the same as
multivariable linear models

Generalized linear models allow for an arbitrary link function, g, that relates the mean (https://en.wikipedia.org/wiki/Mean

standard estimators of β to become biased. Generally, the form of bias is an attenuation,
meaning that the effects are biase

(https://en.wikipedia.org/wiki/File:Galton's_correlation_diagram_1875.jpg)
Francis Galton's (https://en.wikipedia.org/wiki


Maximum likelihood estimation (https://en.wikipedia.org/wiki/Maximum_likelihood_estimation)can be performed when the dist

(https://en.wikipedia.org/wiki/File:Thiel-Sen_estimator.svg)
Comparison of the Theil–Sen estimator (black) and simple linea

longitudinal data, or data obtained from cluster sampling. They are generally fit
as parametric (https://en.wikipedia.org/w

Linear regression is widely used in biological, behavioral and social sciences to describe
possible relationships between va

(https://en.wikipedia.org/wiki/File:Wiki_letter_w_cropped.svg)independent variables, to ensure that any observed effect of s

General Linear Model Overview
No ratings yet
General Linear Model Overview
5 pages
Minitab 15 by Bs 3rd Year Statistics
100% (1)
Minitab 15 by Bs 3rd Year Statistics
138 pages
Panel Data Analysis in R
No ratings yet
Panel Data Analysis in R
1 page
Stat Fiai 2 Sampel
No ratings yet
Stat Fiai 2 Sampel
6 pages
PDF Latihan Soal Ujian Masuk Apoteker Uin Jakarta Compress
No ratings yet
PDF Latihan Soal Ujian Masuk Apoteker Uin Jakarta Compress
20 pages
The Burden of Antimicrobial Resistance (AMR) in Indonesia
No ratings yet
The Burden of Antimicrobial Resistance (AMR) in Indonesia
4 pages
Summary of Econometrics Chapters 3-5
No ratings yet
Summary of Econometrics Chapters 3-5
64 pages
Moving Average and Smoothing Techniques
No ratings yet
Moving Average and Smoothing Techniques
17 pages
QMT634 Buku Online
No ratings yet
QMT634 Buku Online
71 pages
Santander Customer Transaction Prediction Using R - PDF
No ratings yet
Santander Customer Transaction Prediction Using R - PDF
171 pages
(Springer Series in Statistics) Wolfgang Härdle (Auth.) - Smoothing Techniques - With Implementation in S-Springer-Verlag New York (1991)
No ratings yet
(Springer Series in Statistics) Wolfgang Härdle (Auth.) - Smoothing Techniques - With Implementation in S-Springer-Verlag New York (1991)
266 pages
Statistical Analysis of 24-1 Factorial Designs
100% (1)
Statistical Analysis of 24-1 Factorial Designs
10 pages
Statistika untuk Bisnis dan Manajemen
No ratings yet
Statistika untuk Bisnis dan Manajemen
11 pages
Phytochemicals in Chemotaxonomy Explained
No ratings yet
Phytochemicals in Chemotaxonomy Explained
11 pages
Cost Measurement and Estimation Guide
No ratings yet
Cost Measurement and Estimation Guide
28 pages
Advanced Statistics for Analysts
No ratings yet
Advanced Statistics for Analysts
4 pages
Lazy Learning (Or Learning From Your Neighbors)
No ratings yet
Lazy Learning (Or Learning From Your Neighbors)
3 pages
Document Overview and Analysis
No ratings yet
Document Overview and Analysis
28 pages
PCNE Classification For Drug-Related Problems V9.1 - Page 1
No ratings yet
PCNE Classification For Drug-Related Problems V9.1 - Page 1
10 pages
Lampiran Tabel R Product Moment
No ratings yet
Lampiran Tabel R Product Moment
1 page
ANOVA and PCA Analysis of Salary Data
No ratings yet
ANOVA and PCA Analysis of Salary Data
35 pages
Lampiran Syntax R Sarima
No ratings yet
Lampiran Syntax R Sarima
4 pages
Nardl Package: Cointegration Bounds Test Dynamic Multipliers Plot
No ratings yet
Nardl Package: Cointegration Bounds Test Dynamic Multipliers Plot
1 page
B.Tech Probability & Statistics Question Bank
100% (1)
B.Tech Probability & Statistics Question Bank
7 pages
Medicinal Chemistry Analysis
No ratings yet
Medicinal Chemistry Analysis
92 pages
NP and P Chart Construction Guide
No ratings yet
NP and P Chart Construction Guide
28 pages
Statistics for Economics Syllabus
No ratings yet
Statistics for Economics Syllabus
6 pages
Simple Linear Regression Overview
No ratings yet
Simple Linear Regression Overview
9 pages
1 - Linear Models
No ratings yet
1 - Linear Models
22 pages
Generalised Linear Model
No ratings yet
Generalised Linear Model
4 pages
Lecture 3
No ratings yet
Lecture 3
18 pages
Lecture3 221109 035214
No ratings yet
Lecture3 221109 035214
87 pages
Regression III: Advanced Methods: William G. Jacoby Department of Political Science
No ratings yet
Regression III: Advanced Methods: William G. Jacoby Department of Political Science
21 pages
Simple Linear Regression Overview
No ratings yet
Simple Linear Regression Overview
27 pages
Understanding Linear Regression Models
No ratings yet
Understanding Linear Regression Models
23 pages
Unit III Regression
No ratings yet
Unit III Regression
24 pages
Gauss-Markov Assumptions Explained
No ratings yet
Gauss-Markov Assumptions Explained
11 pages
Generalized Linear Model
No ratings yet
Generalized Linear Model
9 pages
Mungadze Linear
No ratings yet
Mungadze Linear
21 pages
ML - LAB - BE CSE (DS) Final
No ratings yet
ML - LAB - BE CSE (DS) Final
110 pages
Understanding Linear Regression Basics
No ratings yet
Understanding Linear Regression Basics
34 pages
FML Unit2
No ratings yet
FML Unit2
13 pages
Machine Learning Unit2
No ratings yet
Machine Learning Unit2
31 pages
Unit 3 Notes
100% (2)
Unit 3 Notes
32 pages
Understanding Regression Models Basics
No ratings yet
Understanding Regression Models Basics
15 pages
Simple Linear Regression Assumptions
No ratings yet
Simple Linear Regression Assumptions
20 pages
Overview of Generalized Linear Models
No ratings yet
Overview of Generalized Linear Models
3 pages
Unit III
No ratings yet
Unit III
18 pages
Regression Analysis in Statistics
No ratings yet
Regression Analysis in Statistics
19 pages
Understanding Linear Regression Basics
No ratings yet
Understanding Linear Regression Basics
7 pages
Unit III
No ratings yet
Unit III
24 pages
Data Analytics Iii Unit
No ratings yet
Data Analytics Iii Unit
8 pages
Understanding Multiple Regression Analysis
100% (2)
Understanding Multiple Regression Analysis
23 pages
Simple Linear and Logistic Regression
No ratings yet
Simple Linear and Logistic Regression
81 pages
Daunit 3
No ratings yet
Daunit 3
32 pages
Linear Regression Models Guide
No ratings yet
Linear Regression Models Guide
42 pages
Understanding Simple Linear Regression
No ratings yet
Understanding Simple Linear Regression
42 pages
Understanding Linear & Polynomial Regression
No ratings yet
Understanding Linear & Polynomial Regression
56 pages
Understanding Linear Regression Basics
No ratings yet
Understanding Linear Regression Basics
53 pages
Descriptive Epidemiology Data Analysis
No ratings yet
Descriptive Epidemiology Data Analysis
3 pages
Epidemiology: Understanding Confounders
No ratings yet
Epidemiology: Understanding Confounders
19 pages
Practical Guidance On The Use of Premix Insulin - Final
100% (1)
Practical Guidance On The Use of Premix Insulin - Final
14 pages
Cross-Sectional Study Guide
No ratings yet
Cross-Sectional Study Guide
10 pages
ABPA Diagnostic Criteria Overview
No ratings yet
ABPA Diagnostic Criteria Overview
1 page
Opportunities For Investments in Nutrition in Low-Income Asia
No ratings yet
Opportunities For Investments in Nutrition in Low-Income Asia
28 pages
Historical Insights into Social Epidemiology
No ratings yet
Historical Insights into Social Epidemiology
10 pages
Daftar Pustaka
No ratings yet
Daftar Pustaka
7 pages
Importance of Statistics for Teachers
No ratings yet
Importance of Statistics for Teachers
9 pages
Surface Roughness Tester
No ratings yet
Surface Roughness Tester
17 pages
Masonry NC II (Superseded)
No ratings yet
Masonry NC II (Superseded)
67 pages
Reservation Skills for Students
No ratings yet
Reservation Skills for Students
2 pages
CFD Analysis of Flow Over A Winglet Using K-Omega SST Turbulence Model
No ratings yet
CFD Analysis of Flow Over A Winglet Using K-Omega SST Turbulence Model
28 pages
DLL July 7
No ratings yet
DLL July 7
4 pages
The Alchemist: A Journey to Destiny
No ratings yet
The Alchemist: A Journey to Destiny
10 pages
Neural Network Control Methods Guide
No ratings yet
Neural Network Control Methods Guide
11 pages
Sabina Portfolio 2018
No ratings yet
Sabina Portfolio 2018
21 pages
QRG
No ratings yet
QRG
1,894 pages
Understanding Random Variables and Distributions
No ratings yet
Understanding Random Variables and Distributions
6 pages
37 Useful Words and Phrases For Business Negotiations in English
No ratings yet
37 Useful Words and Phrases For Business Negotiations in English
6 pages
Class Xi Physics PWT-1 Question
No ratings yet
Class Xi Physics PWT-1 Question
5 pages
Chapter 6 Communication
No ratings yet
Chapter 6 Communication
8 pages
Hetronic Bms-2 (GB)
92% (13)
Hetronic Bms-2 (GB)
20 pages
Oracle 11i E-Business Essentials Exam Guide
No ratings yet
Oracle 11i E-Business Essentials Exam Guide
21 pages
Poetry Analysis for Students
No ratings yet
Poetry Analysis for Students
2 pages
Mathematics in The Environment (Assignment)
No ratings yet
Mathematics in The Environment (Assignment)
2 pages
Kant, Math Physics & Transcendental Philosophy
No ratings yet
Kant, Math Physics & Transcendental Philosophy
15 pages
How To Write Formal Letters
100% (1)
How To Write Formal Letters
5 pages
Som Lab Manual
No ratings yet
Som Lab Manual
34 pages
Types of Counselling and Counselling Skills 2
No ratings yet
Types of Counselling and Counselling Skills 2
12 pages
Ayurvedic Insights for All Ages
No ratings yet
Ayurvedic Insights for All Ages
27 pages
Allama Iqbal Open University, Islamabad: (Department of Business Administration)
No ratings yet
Allama Iqbal Open University, Islamabad: (Department of Business Administration)
7 pages
FLSA Exempt vs Non-Exempt Guide
No ratings yet
FLSA Exempt vs Non-Exempt Guide
1 page
English Language
0% (2)
English Language
55 pages
Understanding English Idioms in Context
No ratings yet
Understanding English Idioms in Context
12 pages
Yarooms Alorica Room Reservation Guidelines
No ratings yet
Yarooms Alorica Room Reservation Guidelines
11 pages
1 s2.0 S0263786301000515 Main
No ratings yet
1 s2.0 S0263786301000515 Main
8 pages

Understanding Multiple Linear Regression

Uploaded by

Understanding Multiple Linear Regression

Uploaded by

multiple linear regression

 Simple and multiple linear regression[edit]

Example of simple linear regression, which has one independent variable

same as general linear regression.

 General linear models[edit]

multivariable linear models (also called "multiple linear models").

and Generalized least squares.) Heteroscedasticity-consistent standard errors is an improved

method for use with uncorrelated but potentially heteroscedastic errors.

 Generalized linear models[edit]

bounded or discrete. This is used, for example:

(which is better described using a Bernoulli distribution/binomial distribution for

binary choices, or a categorical distribution/multinomial distribution for multi-way

Some common examples of GLMs are:

 Poisson regression for count data.

 Logistic regression and probit regression for binary data.

 Ordered logit and ordered probit regression for ordinal data.

Single index models[clarification needed]

index model will consistently estimate β up to a proportionality constant.[11]

 Hierarchical linear models[edit]

regressions, for example where A is regressed on B, and B is regressed on C. It is often used

meaning that the effects are biased toward zero.

 In Dempster–Shafer theory, or a linear belief function in particular, a linear regression

state equations. The combination of swept or unswept matrices provides an alternative

method for estimating linear regression models.

linear regression. These methods differ in computational simplicity of algorithms, presence of

a closed-form solution, robustness with respect to heavy-tailed distributions, and theoretical

assumptions needed to validate desirable statistical properties such as consistency and

 Least-squares estimation and related techniques[edit]

the TLS estimate.

Main article: Linear least squares

Linear least squares methods include mainly:

 Ordinary least squares

 Weighted least squares

 Generalized least squares

 Maximum-likelihood estimation and related techniques[edit]

terms is known to belong to a certain parametric family ƒθ of probability

likelihood estimates when ε follows a multivariate normal distribution with a

known covariance matrix.

 Ridge regression[13][14][15] and other forms of penalized estimation, such as Lasso

regression,[5] deliberately introduce bias into the estimation of β in order to reduce

it is difficult to account for the bias.

 Least absolute deviation (LAD) regression is a robust estimation technique in that it

when no outliers are present). It is equivalent to maximum likelihood estimation under

a Laplace distribution model for ε.[16]

 Adaptive estimation. If we assume that error terms are independent of the

used to non-parametrically estimate the distribution of the error term.[17]

 Other estimation techniques[edit]

points with outliers.

 Bayesian linear regression applies the framework of Bayesian statistics to linear

regression. (See also Bayesian multivariate linear regression.) In particular, the

regression coefficients β are assumed to be random variables with a specified prior

"best" values of the regression coefficients but an entire posterior distribution,

(see quantile regression), or any other function of the posterior distribution.

conditional mean of y given X. Linear quantile regression models a particular

of mixed models include analysis of data involving repeated measurements, such as

as parametric models, using maximum likelihood or Bayesian estimation. In the case

alternative approach to analyzing this type of data.

 Principal component regression (PCR)[7][8] is used when the number of predictor

 Least-angle regression[6] is an estimation procedure for linear regression models that

was developed to handle high-dimensional covariate vectors, potentially with more

covariates than observations.

much less sensitive to outliers.[19]

 Other robust estimation techniques, including the α-trimmed mean approach[citation

See also: Linear least squares § Applications

Main article: Trend estimation

curvature desired in the line.

or event (such as training, or an advertising campaign) caused observed changes at a point in

where other potential changes can affect the data.

those other socio-economic factors. However, it is never possible to include all

possible confounding variables in an empirical analysis. For example, a hypothetical gene

experiments are not feasible, variants of regression analysis such as instrumental

return on all risky assets.

Main article: Econometrics

predict consumption spending,[20] fixed investment spending, inventory investment, purchases

demand,[23] and labor supply.[23]

This section needs