Data Science Cheatsheet

Uploaded by

yane.data

We take content rights seriously. If you suspect this is your content, claim it here.

Available Formats

Download as PDF or read online on Scribd

100% found this document useful (1 vote)

317 views15 pages

Data Science Cheatsheet

Uploaded by

yane.data

We take content rights seriously. If you suspect this is your content, claim it here.

Available Formats

Download as PDF or read online on Scribd

Cheat Sheet — Bias-Variance Tradeoff What is Bias? bias = E[f'(x)] — f(e) + Error between average model prediction and ground truth + The bias of the estimated function tells us the capacity of the underlying model to predict the values What is Variance? variance = E|(f"(z) -Elf'(e)])”| + Average variability in the model prediction for the given dataset * The variance of the estimated function tells you how much the function can adjust to the change in the dataset High Bias © ———> Overly-simplified Model ——> Under-fitting > High error on both test and train data High Variance Overly-complex Model ——> Over-fitting ———> Low error on train data and high on test ——— Starts modelling the noise in the input High Bias Low Bias Low Variance High Variance Low Bias High Bias g 5 Minimum Error 5 > Bias e Variance 8 Error 8 g Z a > 2 ee x __Under-fitting Just Right Over-fitting Preferred if size Preferred if size of dataset is small of dataset is large Bias variance Trade-off + Increasing bias (not always) reduces variance and vice-versa + Error = bias? + variance +irreducible error + The best model is where the error is reduced + Compromise between bias and variance Source: https: //www.cheatsheets.ageel-anwar.com Tutorial: Click hereCheat Sheet — Imbalanced Data in Classification SRAGOCASEE: r=: @ Comet Frodtins | MOO nnn ' Accurae ¥ ~ “Total Predictions Classifier that always predicts label blue yields prediction accuracy of 90% Accuracy doesn’t always give the correct insight about your trained model Accuracy: %age correct prediction Correct prediction over total predictions One value for entire network Precision: Exactness of model From the detected cats, how many were Each class/label has a value actually cats Recall: Completeness of model Correctly detected cats over total cats Each class/label has a value F1 Score: Combines Precision/Recall Harmonic mean of Precision and Recall Each class/label has a value Performance metrics associated with Class 1 (Is your prediction correct?) (What did you predict) ‘Actual Labels Te Negative + (Your prediction is correct) (You predicted 9) (Prec x Rec) F1 score = 2x —————— (Prec + Rec) Predicted Labels Ce Possible solutions 1. Data Replication: Replicate the available data until the Blue: Label] *(@ number of samples are comparable Green: Label 0 2. Synthetic Data: Images: Rotate, dilate, crop, add noise to Blue: Label #0 existing input images and create new data Green: LabelO 1: @ Bh M@OOAOEO 3. Modified Loss: Modify the loss to reflect greater error when misclassifying smaller sample set 4, Change the algorithm: Increase the model/algorithm complexity so that the two classes are perfectly separable (Con: Overfitting) loss = a* lossgreen tb 0SSpiue a >b Increase model complexity ‘No straight line (y=ax) passing through origin can perfectly Straight line (y=ax+b) can perfectly separate data, separate data, Best solution: line y=0. predict all labels blue Green class will no longer be predicted as blue Source: https://www.cheatsheets.aqeel-anwar.com Tutorial: Click here ogpagécCheat Sheet - PCA Dimensionality Reduction What is PCA? + Based on the dataset find a new set of orthogonal feature vectors in such a way that the data spread is maximum in the direction of the feature vector (or dimension) + Rates the feature vector in the decreasing order of data spread (or variance) + The datapoints have maximum variance in the first feature vector, and minimum variance in the last feature vector + The variance of the datapoints in the direction of feature vector can be termed as a measure of information in that direction, Steps - X = mean(X) 1. Standardize the datapoints pe std(X) 2. Find the covariance matrix from the given datapoints Cli, j] = cov(ei, x) 3. Carry out eigen-value decomposition of the covariance matrix c=vsv-! 4, Sort the eigenvalues and eigenvectors Short = sort(S) Vione = sort(V; Seon) Dimensionality Reduction with PCA + Keep the first m out of n feature vectors rated by PCA. These m vectors will be the best m vectors preserving the maximum information that could have been preserved with m vectors on the given dataset Steps: 1. Carry out steps 1-4 from above D0) Keep frat mal fenturelvectors|from\ the sorted eigenvector mater Vreducea = V[+,0: m] 3. Transform the data for the new basis (feature vectors) Xreduced = Xnew X Vreduced 4. The importance of the feature vector is proportional to the magnitude of the eigen value Figure 1 E 2 Ea g 3 i a F2 FI Feature #2 (F2) r2 FI Figure 1: Datapoints with feature vectors as x and y-axis Figure 2: The cartesian coordinate system is rotated to maximize the standard deviation along any one axis (new feature # 2) Figure 3: Remove the feature vector with minimum standard deviation of datapoints F2 (new feature # 1) and project the data on new feature # 2 Source: https: //www.cheatsheets.aqeel-anwar.comCheat Sheet — Bayes Theorem and Classifier What is Bayes’ Theorem? + Describes the probability of an event, based on prior knowledge of conditions that might be related to the event. P(B\A) clihood) x P(A) P(B)(evidence) P(AIB) = + How the probability of an event changes when we have knowledge of another event P(A) —> P(AIB) L, Unualy, a better estimate than P(A) Example + Probability of fire P(F) = 1% + Probability of smoke P(S) = 10% + Prob of smoke given there is a fire P(SIF) = 90% Ex + What is the probability that there is a fire given POBIA = P(B) we see a smoke P(F|S)? oy _ P(S|F)x P(F) _ 0.90.01 _ ,, p(ris) = FO seen = 9% Maximum Aposteriori Probability (MAP) Estimation ‘The MAP estimate of the random variable y, given that we have observed iid (x1, x2, x5, «+. ), is given by. We try to accommodate our prior knowledge when estimating. g=argmax, Ply P(xily y that maximizes the product of Hz argmax, P(y)T] Plaily) pale ee Maximum Likelihood Estimation (MLE) ‘The MAP estimate of the random variable y, given that we have observed iid (x1, x2, x5, «+. )s i8 given by. We assume we don't have any prior knowledge of the quantity being estimated. = argmary T[P@il) y that maximizes only the likelihood MLE is a special case of MAP where our prior is uniform (all values are equally likely) Naive Bayes’ Classifier (Instantiation of MAP as classifier) Suppose we have two classes, y=y; and y=ys, Say we have more than one evidence/features (x: Xo, Xg, «- ), using Bayes’ theorem Pla ly) x Pw) +) ) are iid. ie P(ai,a Ply) = IP@w Peumnran) Plyie ) ) Plyo|er, 22,23, Plyler 2,25 Plo Naive Bayes’ theorem assumes the features (3x1, x2, -.. 3,---|y) = [J Plaidy) Plyler, 22,2: ©2,03, gamit >1 else J=y2 Source: https: //www.cheatsheets.aqeel-anwar.comCheat Sheet — Regression Analysis ession Analysis? Fitting a function {(,) to datapoints y;=f(x;) under some error function, Based on the estimated function and error, we have the following types of regression 1. Linear Regression: ming lw — 4" eI? Fits a line minimizing the sum of mean-squared error 7 for each datapoint. SBME) = Bo + Pa 2. Polynomial Regression: ming lly — Sel? Fits a polynomial of order k (k-+1 unknowns) minimizing = the sum of mean-squared error for each datapoint. £8" (ae) = By + Hiei + Bar? +... + Berk 3. Bayesian Regression: For each datapoint, fits a eanssian distribution by ming Y llw —N (fa(wi),07) |? minimizing the mean-squared error, As the number of Fale) = f(a) o H8°°"(as) data points x; increases, it converges to point estimates ie. n + 00,07 +0 A (y.02) -+ Gaussian with mean x and variance o? 4, Ridge Regression: m . Can fit either a line, or polynomial minimizing the sum ming x Ive — fa(@s)|? + Le of mean-squared error for each datapoint and the . . weighted L2 norm of the function parameters beta. 5. LASSO Regression: Can fit either a line, or polynomial minimizing the the sum of mean-squared error for each datapoint and the w 1L1 nom of the function parameters beta. 6. Logistic Regression: sming > —wlog (0 (f(2s)))~ (1 wlog (1 ~ a (F560) Can fit either a line, or polynomial with sigmoid activation minimizing the binary cross-entropy loss for fale) = Ras) or fine" (ay) each datapoint. The labels y are binary class labels. ott) = Visual Representation: Linear Regression Polynomial Regression ‘Bayesian Linear Regression What does it fit? Estimated function Error Function Linear A line in n dimensions: St (as) + Bias ym Sole Polynomial A polynomial of order Fy as) = f+ Bas + Bat + Ete nor Bayesian Linear Gaussian distribution for cach point Usted?) El x Gaede Ridge Linear/polynomial Sp 0) on F(a) LASSO Linear/polynomial spa) or ta) Logistic Linear polynomial with sigmoid aXfolen) ing wl afte) ~ win ee) Source: https: //www.cheatsheets.ageel-anwar.com Tutorial: Click hereerr if sie referred if size ntact is stall of dataset is large Figure 1. Overfitting Cheat Sheet — Regularization in ML ‘What is Regularization in ML? + Regularization is an approach to address over-fitting in ML. 4 * Overfitted model fails to generalize estimations on test data * When the underlying model to be learned is low bias/high variance, or when we have small amount of data, the Underfitting Just Right Over-fitting estimated model is prone to over-fitting. + Regularization reduces the variance of the model ‘Types of Regularizatio 1. Modify the loss function: + L2 Regularization: Prevents the weights from getting too large (defined by L2 norm). Larger the weights, more complex the model is, more chances of overfitting. loss = error(y,f)+A})?| AZO, Ax model bias, Xo i eel ‘model variance + L1 Regularization: Prevents the weights from getting too large (defined by L1 norm). Larger the weights, more complex the model is, more chances of overfitting. L1 regularization introduces sparsity in the weights. It forces more weights to be zero, than reducing the the average magnitude of all weights loss = error(y,)|+ A> |B;|| 20, Ax model bias, do model variance + Entropy: Used for the models that output probability. Forces the probability distribution towards uniform distribution. loss = error(p,p)|~ 2 pilog(p,)) > 0, Axx model bias, do 1 model variance 2. Modify data sampling: + Data augmentation: Create more data from available data by randomly cropping, dilating, rotating, adding small amount of noise etc. + K-fold Cross-validation: Divide the data into k groups. Train on (k-L) groups and test on 1 group. Try all k possible combinations. 3. Change training approach: + Injecting noise: Add random noise to the weights when they are being learned. It pushes the model to be relatively insensitive to small variations in the weights, hence regularization + Dropout: Generally used for neural networks. Connections between consecutive layers are randomly dropped based on a dropout-ratio and the remaining network is trained in the current iteration. In the next iteration, another set of random connections are dropped. Sefold cross-validation Original Network Dropout-ratio = 30% (QUUUNUNETSGRI FtsE | comnections = 16 Active = 11 (705) Active = 11 (70%) Figure 2. K-fold CV Figure 3. Drop-out Source: https://www.cheatsheets.ageel-anwar.com Tutorial: Click hereCheat Sheet —- Convolutional Neural Network Convolutional Neural Network: Dataset: ( 1 month > 3 months Fig. 1 ~ Preparati + Review Data structures and Complexities: The following 7 data structures are necessary for the interview, and their time/space complexity + List/Arrays, Linked List, Hash Table/dictionary, Tree, Graph, Heap, Queue * Click here for tutorial. + Practice coding questions: + Multiple online resources such as LectCode.com, InterviewBit.com, HackerRank.com ete. + Pick one online resource and aim for easy and medium coding questions (approx. 100-150). + Beginners start preparing 2-3 months before the interview, and intermediates about 1 month. + Note: + From my personal experience, paid subscription of LeetCode.com was worth it. + Facebook, Uber, Google and Microsoft tagged question of LeetCode covered almost 90% of the questions asked Part 2 — How to answer a coding question?* + Listen to the question ‘The interviewer will explain the question with an example. Note down the important points. + Talk about your understanding of the question Repeat the question and confirm your understanding. Ask clarifying questions such as 1. Input/Output data type limitations 2. Input size/length limitations 3. Special/Comer cases + Discuss your approach Walk through how would you approach the problem and ask the interviewer if he agrees with it. Talk about the data structure you prefer and why. Discuss the solution with the bigger picture. + Start coding Ask the interviewer if you could start coding. Define useful functions and explain as you write. ‘Think out loud so the interviewer can evaluate your thought process. + Discuss the time and space complexity Discuss the time and space complexity in terms of Big O for your coded approach. + Optimize the approach If your approach is not the most optimized one, the interviewer will hint you a few improvements. Pay attention to hints and try to optimize your code. ‘Timeline for Coding Interviews Pen aoe plexity "*Diocaimer: The ccomnendatons ae based on petonal expec ofthe autor. The motioned approsch and resources might work reat for sme, but ot so mc for thers Fig, 2— How to answer a coding question? Source: https: //www.cheatsheets.aqeel-anwar.comHow to prepare for i 4 behavioral interview? 3 ‘ollect stories, assign keywords, prac the STAR format Keywords 222rerneis easy Negotiation COmBFEMET? Creativity __‘Flexibity Convincing Handling Challenging Working with Another team Adjust to a E priorities not colleague Take Stand Crisis Situation difficult people ae ad Handling -ve Coworker Working with ay. strength Your Influence feedback view of you deadline a weakness Others \ Handling Converting, Decision ‘ , Handling ynexpected challenge to without enough ,Couflict_ Mentorship/ Resolution Leader situation opportunity. data Stories 1. List all the organizations you have been a part of. For example 1. Academia: BSc, MSc, PhD 2. Industry: Jobs, Internship 3. Societies: Cultural, Technical, Sports 2. Think of stories from step 1 that can fall into one of the keywords categories. The more stories the better. You should have at least 10-15 stories. 3. Create a summary table by assigning multiple keywords to each stories. This will help you filter out the stories when the question asked in the interview. An example can be seen below Story 1 [Convincing] [Take Stand] [influence other Story 2: [Mentorship] [Leadership] Story 8: [Conflict resolution) [Negotiation] Story 4 [decision-without-enough-data STAR Format Write down the stories in the STAR format as explained in the 2/4 part of this cheat sheet. This will help you practice the organization of story in a meaningful way. Source: https: //www.cheatsheets.ageel-anwar.com Tutorial: Click hereHow to prepare for o 4 behavioral interview? Direct*, meaningful*, personalized*, logical* "(Respective colors are used to identify these characteristics inthe example) Example: “Tell us about a time when you had to convince senior executives” Situation Explain the situation and provide necessary context for your story. Task Explain the task and your responsibility in the situation Action Walk through the steps and actions you took to address the issue Result State the outcome of the result of your actions Source: https://www.cheatsheets.aqeel-anwar.com Tutorial: ile A R “I worked as an intern in XYZ company in the summer of 2019. The project details provided to me was elaborative. After some initial brainstorming, and research I realized that the project approach can be modified to make it more efficient in term: decided to talk to my manage he underlying KP bout it "Lhad an hour-long call with my manager and expla approach and how it could improve the KPIs. 1 was able to convince him. H d him in detail the proposed ked me if I will be able to preser posed appro working out of the ABC(city) office and the executives need to fly in from XYZ(city) office.” “I did a quick background check on the executives to know better about their area of expertise so that T can convince them accordingly. I prepared an elaborative 15 slide presentation starting with explainin ir approach, mo ny prop pproach and finally comparing them on “After some active discussion we were able to establish that the proposed approach was better than the initial one. The ‘o my approach and really appreciated my ‘tand, At the end of my intemship, 1 was selected among the 3 out of 68 intems who got to meet the senior vice president of the company over lunch” fon Source: www Click here ocee ScHow to answer a a i 4 behavioral question? Understand, Extract, Map, Select and Apply =“ Example: “Tell us about a time when you had to convince senior executives” Understand the question Example: A story where I was able to convince Understand _ ay ceniors. Maybe they had something in mind, and I had a better approach and tried to convince them Extract keywords and tags Extract useful keywords that encapsulates the Extract gist of the question Example: [Convincing], [Creative], [Leadership] Map the keyword to your stories Shortlist all the stories that fall under the Map keywords extracted from previous step Example: Storyl, Story2, Story3, Storyd, ... , Story N See the hee Hoey From the shortlisted stories, pick the one that Select best describes the question and has not been used so far in the interview Example: Story3 Apply the STAR method Apply the STAR method on the selected story to Apply answer the question Example: See Cheat Sheet 2/3 for details Source: https: //www.cheatsheets.ageel-anwar.com Tutorial: Click hereBehavioral Interview - 4 /4 Cheat Sheet Summarizing the behavioral interview =“ Gather important topics as keywords ‘Understand and collect all the important topics commonly asked in the interview How to prepare Practice stories in STAR format Practice ach the STAR format. You will have for the pagel hp bony commensal interview Create a summary table ‘Create a summary table mapping stories to their associated ‘keywords, This will be used during the behavioral question ‘Understand the question ‘Understand the question and clavify any confusions that you have How to answer a 7 Map the keywords to stories question ‘Based on the keywords extracted, find the stories using the during summary table created during preparation (Step 4) interview Apply the START format Once the story has been shortlisted, apply STAR format on the story to answer the question. lo Source: www: Aletcon cam Source: https://www.cheatsheets.aqeel-anwar.com Tutorial: Click here oggagéc

The Python Bible
97% (34)
The Python Bible
506 pages
Object Oriented Python Tutorial
100% (21)
Object Oriented Python Tutorial
111 pages
The Complete Guide To Modern Javascript
100% (12)
The Complete Guide To Modern Javascript
295 pages
Data Structure and Algorithms With Python
100% (17)
Data Structure and Algorithms With Python
369 pages
Python Programming for Beginners_ From Basics to AI Integrations. 5-Minute Illustrated Tutorials, Coding Hacks, Hands-On Exercises & Case Studies to Master Python in 7 Days and Get Paid More by Prince
100% (16)
Python Programming for Beginners_ From Basics to AI Integrations. 5-Minute Illustrated Tutorials, Coding Hacks, Hands-On Exercises & Case Studies to Master Python in 7 Days and Get Paid More by Prince
244 pages
Full Course of Machine Learning
100% (18)
Full Course of Machine Learning
660 pages
PYTHON Learn Python Programming in 90 Minutes or Less Python Learning Python Python Programming Python Tutorial Python Programming For Beginners Python For Dummies Book 1 PDF
93% (15)
PYTHON Learn Python Programming in 90 Minutes or Less Python Learning Python Python Programming Python Tutorial Python Programming For Beginners Python For Dummies Book 1 PDF
161 pages
Data Structure and Algorithmic Thinking With Python Data Structure and Algorithmic Puzzles PDF
96% (24)
Data Structure and Algorithmic Thinking With Python Data Structure and Algorithmic Puzzles PDF
471 pages
Essential Linux Commands Handbook
100% (17)
Essential Linux Commands Handbook
135 pages
Front-End Developer Handbook 2019 PDF
97% (32)
Front-End Developer Handbook 2019 PDF
145 pages
Git Essentials
88% (8)
Git Essentials
120 pages
System Design Interview - An Insider's Guide
90% (10)
System Design Interview - An Insider's Guide
103 pages
Learning The Pandas Library Python Tools For Data Munging Analysis and Visual PDF
100% (19)
Learning The Pandas Library Python Tools For Data Munging Analysis and Visual PDF
208 pages
Practical Projects
100% (32)
Practical Projects
478 pages
Jenkins Book PDF
No ratings yet
Jenkins Book PDF
224 pages
Artificial Intelligence With Python (Machine Learning Foundations, Methodologies, and Applications) (Teik Toe Teoh, Zheng Rong)
95% (19)
Artificial Intelligence With Python (Machine Learning Foundations, Methodologies, and Applications) (Teik Toe Teoh, Zheng Rong)
334 pages
SQL Commands Cheat Sheet
82% (11)
SQL Commands Cheat Sheet
1 page
Python Notes For Professionals
100% (18)
Python Notes For Professionals
814 pages
The Python Manual
94% (33)
The Python Manual
196 pages
Linux Essentials For Cybersecurity
96% (25)
Linux Essentials For Cybersecurity
1,966 pages
Learning React Modern Patterns For Developing React Apps 2nbsped Compress
100% (7)
Learning React Modern Patterns For Developing React Apps 2nbsped Compress
488 pages
Hacking - Hacking With Python - The Complete Beginner'Srcises (Hacking Series Book 1)
100% (10)
Hacking - Hacking With Python - The Complete Beginner'Srcises (Hacking Series Book 1)
150 pages
Python Programming. A Step-by-Step Guide For Absolute Beginners
92% (49)
Python Programming. A Step-by-Step Guide For Absolute Beginners
181 pages
Applied Generative AI For Beginners Practical Knowledge 1703207445
95% (19)
Applied Generative AI For Beginners Practical Knowledge 1703207445
221 pages
System Design Interview Fundamentals
100% (6)
System Design Interview Fundamentals
412 pages
Competitive Programming in Python 128 Algorithms To Develop Your Coding Skills
100% (10)
Competitive Programming in Python 128 Algorithms To Develop Your Coding Skills
267 pages
Hackers Guide To Machine Learning With Python PDF
100% (16)
Hackers Guide To Machine Learning With Python PDF
272 pages
The JavaScript Beginner's Handbook
91% (11)
The JavaScript Beginner's Handbook
76 pages
Hacking The Art of Exploitation 2nd Edition Jon Erickson
100% (21)
Hacking The Art of Exploitation 2nd Edition Jon Erickson
492 pages
The Road To React
92% (13)
The Road To React
250 pages
Adjudicación Simplificada 024-2023 Perú Compras
No ratings yet
Adjudicación Simplificada 024-2023 Perú Compras
55 pages
Metodos Estadisticos Con R
No ratings yet
Metodos Estadisticos Con R
157 pages
Ggplot 2
No ratings yet
Ggplot 2
19 pages
Términos de Referencia
No ratings yet
Términos de Referencia
5 pages
Estadística para Ciencia de Datos
No ratings yet
Estadística para Ciencia de Datos
1 page

Data Science Cheatsheet

Uploaded by

Data Science Cheatsheet

Uploaded by

You might also like