# About

This page tells you what our vision and intention for this book is and how you can help in making it better.

> *The fellow-pupil can help more than the master because he knows less. The difficulty we want him to explain is one he has recently met. The expert met it so long ago that he has forgotten…*
>
> *- C.S. Lewis in his Reflections on the Psalms*

Preparing for interviews is a stressful task. There is an enormous amount of resources available in the internet, multiple repositories, there are even companies that help students prepare for interviews at the Big Tech companies. The idea here is to create an web based accessible version of these resources, optimized across devices so people on this journey can benefit from it no matter where they are from.

This book does not cover the topics in depth, it covers just enough to get you ready for the interview. The assumption here is that the person using it is already familiar with the topic and is here to brush up on the same. Additional resources for someone eager to explore the topic in depth is added. **In short, don't use this as textbook, use it as a revision note.**

This online version of the book **IS AND ALWAYS WILL** be **free** and **a work in progress**. The book gets updated monthly with new sections, more questions and richer content, check the [Log](https://dipranjan.github.io/dsinterviewqns/contents/To%20Do%20List.html) page for the latest updates and keep visiting for new content.

***

Every project requires resources to maintain and keep it relevant. The trouble with individual-driven open-source projects of this kind is that it runs on hope, banking on the goodwill of users to support the project. So, if this book has helped you in any way or you see merit in it, show some ❤:

* The best way is to **support the project**. You can  [Buy me a Coffee ☕](https://www.buymeacoffee.com/dearc).
* You can also help by adding more content to make this book more relevant, feel free to create a merge request or post in [👪 forum](https://github.com/dipranjan/dsinterviewqns/discussions)
* If you want to share your interview experience via ✉, please drop one to *<thedatascienceinterviewbook@gmail.com>*, it will just be used to improve and enhance the book not for any other purpose


# Log

The journey of the book so far

**Month of Sep, 2025**

* Well sad news, Gitbook has decided to nerf and remove features from its free tier, so there are things which might stop working here pretty soon. The paid tiers are pretty costly for us to afford, but we will keep this website up as long as possible. If you find this helpful, please spread the word and consider donating using the Buy me a coffee button at the top. Given the current job market, we feel this resource might be more relevant than ever.

**Month of Feb, 2025**

* MLE calculation error updated in the [Probability Distribution](/statistics/probability-distribution) page

**Month of Oct, 2024**

* [Pyspark page](/python/pyspark) added

**Month of Sep, 2024**

* Pyspark [cheat sheet](/cheat-sheets/pyspark) added

**Month of Feb, 2024**

* [Clustering](/algorithms/clustering) page updated
* Questions added in the [Clustering ](/algorithms/clustering#questions)page

<details>

<summary>Older Updates</summary>

**Month of Oct, 2023**

* [Problems page](/business-intelligence/power-bi/problems) added in Power BI
* [Cheatsheets](/cheat-sheets/numpy) pushed to the end of the contents
* [Index](/sql/index) added in SQL
* [Temp datasets](/sql/temporary-datasets) updated with Table Variable and Temp tables
* [Performance Tuning ](/sql/performance-tuning)added in SQL
* [Study reference materials](/python/basics) added for Python programming questions
* New Problems added in Python and SQL

**Month of Sep, 2023**

* [Windows Function](/sql/windows-functions) updated.
* [A/B test page](/statistics/a-b-test) (WIP) added.
* SVR and OLS added in [Regression page](/algorithms/regression).
* [Classification metrics](/algorithms/classification#metrics) updated with use cases.
* Optimizers and Optimization Criterion updated in [Algorithm overview page](/algorithms/overview).
* [Model Building Overview](/model-building/overview) page added.
* [Naive Bayes](/algorithms/classification#naive-bayes-algorithm) added in classification.
* Many new SQL and Python problems added.
* [Confidence Interval](/statistics/central-limit-theorem#what_is_confidence_interval) added in Central Limit Theorem.
* Coding [Algorithms from scratch](/python/algorithms-from-scratch) added in Python.
* [Common hypothesis tests](/statistics/hypothesis-testing) added.
* [Neural Network](/neural-network/neural-network) section updated

**Month of August, 2023**

* Vector Database added in the LLM section.
* Categorical Encoding added in the data section.
* Probability Distribution cheat sheet added.
* Algorithm overview page updated.
* Bias Variance tradeoff updated.
* Linear Regression page updated.
* LLM Section updated to Generative AI.
* Clustering (WIP) section added.
* QUALIFY added in Windows functions page.
* We are now on [Instagram](https://www.instagram.com/thedatascienceinterviewbook/) , [TikTok](https://www.tiktok.com/@the.ds.interview?_t=8f1vEGiHYYk&_r=1) and [YouTube](https://youtube.com/@thedatascienceinterviewboo7076), <mark style="color:red;">please do follow</mark> <mark style="color:red;">we will start uploading content soon.</mark>
* Python in Excel Page added.
* Many Python Questions added.
* Transformer page added.

**Month of May, 2023**

* Have enabled GitBook's Lens feature in search, which allow users to ask a question and get answers back from the content of the book itself. This is an experimental feature and supported by OpenAI. Please note this is experimental and can be changed or removed at any moment.
* Work on Power BI section started under the Business Intelligence section.
* Dark mode and Light mode toggle enabled.

**Month of January, 2023**

* R Basics cheat sheet added
* Python Theoretical Question section updated
* Mathematical Motivation Page added

**Month of November, 2022**

* Group vs Window added
* Git added in the new ML Ops section
* Platform migration for the book
* Cheat Sheet section added
* ⚠️ Sign beside pages indicate that work is pending on those
* Added questions to Bias/Variance
* Python Theoretical section added --> TBA in BOOK

**Month of October, 2022**

* More questions added to the Time Series Section
* Bias/Variance Tradeoff added
* Ensemble learning section updated in Decision Tree
* MAP vs MLE added in Probability Basics
* Basic Overview page added in the Algorithm section

**Month of September, 2022**

* As per suggestions by users PDF of the book as been made available as a paid extra. It can be purchased from [here](https://www.buymeacoffee.com/dearc/e/88363)
* Big O notation section added
* Anamoly detection and Time Series section extensively updated
* Probability `[FACEBOOK] N Dice`, `[SPOTIFY] MLE of Uniform Distribution`,`Bernoulli trial generator` problem solution updated
* Business Scenarios section updated

**Month of August, 2022**

* Behavioral - Management section added
* New interview questions added

**Month of July, 2022**

* Data sampling section added under data

**Month of June, 2022**

* We are back post break, keep checking for new content
* Machine Learning Framework section added and TensorFlow moved into it
* PyCaret added to Machine Learning Framework section

**Month of March, 2022**

* Hyperparameter optimization section completed
* Had an extremely busy last few weeks and the next few months are going to be packed too
* Story Telling section added
* Quick guide to Visualization added

**Month of February, 2022**

* Added problems in Python, SQL, Probability
* Excel section updated
* Data section has been moved into a new and broader section called Model Building
* To keep the table of contents clean collapsible headers used in Model Building section
* Hyperparameter optimization section added

**Month of January, 2022**

* Neural Network section added
* Added new problems in the Probability section
* Added cartoons in a few sections
* Outlier section added

**Month of December, 2021**

* NLP section updated
* Got our first bug reported by a reader 😍

**Month of November, 2021**

* NLP section updated
* Missing values section added
* Formatting changes in the Statistics section
* Took some break, was obsessively working on this 😌
* New section - Tree based approaches, Industry application added
* Decided to make this page a little more interesting
* Launched our LinkedIn page do [![Follow LinkedIn](https://img.shields.io/badge/Follow-LinkedIn-0077B5?style=flat-square\&logo=appveyor.svg)](https://www.linkedin.com/company/the-data-science-interview-book/?lipi=urn%3Ali%3Apage%3Ad_flagship3_feed%3BeglbXB3xT0mopZBzReqMEQ%3D%3D), have some interesting plans for it in near future
* Added support for dark theme, 🤯 had to remove it as it was breaking a lot of other stuff. Will wait for official support
* Added new problems in Probability, Python, Regression, SQL
* Added Temporary Datasets and Time page in SQL covering CTEs
* Regression section extensively updated

**Month of October, 2021**

* Major updates to the SQL section
* TensorFlow, Excel, Data Sections added
* Added new problems in Probability, Python, SQL, Business Case
* Cleaned up the formatting issues
* Added this change log section
* Added Generative VS Discriminative Models section
* Completed Hypothesis Testing

</details>


# Mathematical Motivation

This page contains a preliminary discussion into what the different mathematical concepts are and how they relate to data science.

{% hint style="info" %}
The motivation (pun-intended) for having this chapter came when I was studying an [*Introduction to Probability for Data Science*](https://services.publishing.umich.edu/wp-content/themes/mpub-services/library/pdf/PDSdownload100.pdf) by *Stanley H*. *Chan.* A lot of material in this chapter is taken from the book, do check it out.
{% endhint %}

Data Science, Machine Learning is at its core widely dependent on different mathematical concepts which have been developed over centuries. But in this age of easy-to-use packages, frameworks, AutoML solutions we often tend to forget or skip it. In this chapter we would like to quickly glance at some of the broad mathematical topics and how they relate to Data Science. Obviously, it is not in the scope of this book to cover the topics in depth, but the goal here is to leave you with an appreciation of goes behind the algorithms and if needed you can always explore more on your own.

“Data science” has different meanings to different people. If you ask a biologist, data science could mean analyzing DNA sequences. If you ask a banker, data science could mean predicting the stock market. If you ask a software engineer, data science could mean programs and data structures; if you ask a machine learning scientist, data science could mean models and algorithms. However, one thing that is common in all these disciplines is the concept of uncertainty. <mark style="color:yellow;">We choose to learn from data because we believe that the latent information is embedded in the data</mark> — unprocessed, contains noise, and could have missing entries. If there is no randomness, all data scientists can close their business because there is simply no problem to solve. However, the moment we see randomness, our business comes back. <mark style="color:yellow;">Therefore, data science is the subject of making decisions in uncertainty.</mark>

## Infinite Series

Imagine that you have a fair coin. What is the probability that you need to flip the coin three times to get one head?

Since the coin is fair, the probability of obtaining a head is 1/2 . The probability of getting a tail followed by a head is 1/2 × 1/2 = 1/4 . Similarly, the probability of getting two tails and then a head is 1/2 × 1/2 × 1/2 = 1/8 . If you follow this logic, you can write down the probabilities for all other cases. For your convenience, we have drawn the first few in Figure 1.1. As you have probably noticed, the probabilities follow the pattern { 1/2 , 1/4 , 1/8 , . . .}.

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FQKRp8zPGRkPUoe1F6tHs%2Fimage.png?alt=media&amp;token=cd05b041-927e-4263-9bfd-4056ab479adb" alt=""><figcaption><p>The histogram of flipping a coin until we see a head. The x-axis is the number of coin flips, and the y-axis is the probability.</p></figcaption></figure>

Let us ask something harder: On average, if you want to be 90% sure that you will get a head, what is the minimum number of attempts you need to try? Five attempts? Ten attempts? Indeed, if you try ten attempts, you will very likely accomplish your goal. However, this would seem to be overkill.&#x20;

$$P\[success-after-4-attempts] = 1/2 + 1/4 + 1/8 + 1/16 = 0.9375$$. You should be aware that the 93.75% only says that the probability of achieving the goal is high. If you have a bad day, you may still need more than four attempts. Therefore, when we stated the question, we asked for 90% “on average”. Sometimes you may need more attempts and sometimes fewer attempts, but on average, you have a 93.75% chance of succeeding.&#x20;

A geometric series is useful when handling situations such as N − 1 failures followed by a success. However, we can easily twist the problem by asking: What is the probability of getting one head out of 3 independent coin tosses?&#x20;

In general, the number of combinations can be systematically studied using combinatorics. However, the number of combinations motivates us to discuss another background technique known as the binomial series. The binomial series is instrumental in algebra when handling polynomials such as $$(a + b)^2 or (1 + x)^3$$.

The binomial theorem makes the most sense when we also learn about the Pascal’s identity. But we will not cover it in detail here.

## Approximation

Consider a function, $$f(x) = log(1 + x)$$ , for  $$x >0$$ as shown below. This is a nonlinear function, and we all know that nonlinear functions are not fun to deal with. For example, if you want to integrate the function $$\int\_a^b x log(1 + x) ,dx$$, then the logarithm will force you to do integration by parts. However, in many practical problems, you may not need the full range of $$x >0$$ . Suppose that you are only interested in values $$x <<1$$  Then the logarithm can be approximated, and thus the integral can also be approximated.

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FazD8QLb04ANeMDBeNmjD%2Fimage.png?alt=media&amp;token=5e6cfe9d-4a53-42f2-8a78-80b74d8fa1d7" alt=""><figcaption></figcaption></figure>

<mark style="color:yellow;">Given a function it is often useful to analyze its behavior by approximating using its local information. Taylor approximation (or Taylor series) is one of the tools for such a task.</mark> It is a geometry-based approximation. It approximates the function according to the offset, slope, curvature, and so on. The Taylor series has an infinite number of terms. If we use a finite number of terms, we obtain the nth-order Taylor approximation:

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FbIB7FqUtvpDhYA6irL2P%2Fimage.png?alt=media&amp;token=13a2ae33-86fc-4704-b5dd-184940fc491d" alt=""><figcaption></figcaption></figure>

*What order of approximation is good?* It depends on where you want the approximation to be good, and how far you want the approximation to go. The difference between first-order and second-order approximations is shown below:

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FghvZ4qqnLTKy8QoTSbb1%2Fimage.png?alt=media&amp;token=22215c5c-7885-4153-bcd9-e5899573d10e" alt=""><figcaption></figcaption></figure>

## Calculus

How can the fundamental theorem of calculus ever be useful when studying probability? Two concepts: probability density function and cumulative distribution function and these two functions are related to each other by the fundamental theorem of calculus.

## Linear Algebra

This is one of the most important pillars on which the domain rests.&#x20;

* It provides a way to vectorize and represent the data:

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FDrk2lCuiFAFSFgfQafgw%2Fimage.png?alt=media&amp;token=23cb4806-dca5-4764-b56d-a050d407bdc2" alt=""><figcaption><p>Suppose we have a crime dataset and we want to figure out which factors influence the crime rate of a city. One way to do it is to put the numbers in matrices and vectors, as shown above. With this vector expression of the data, the analysis questions can roughly be translated to finding β’s in the last equation.</p></figcaption></figure>

* How can you calculate how different your prediction is from the expected output? Loss Functions, of course. A loss function is an application of the **Vector Norm** in Linear Algebra. The norm of a vector can simply be its magnitude.
* Regularization is a very important concept in data science. It’s a technique we use to prevent models from overfitting. *Regularization is actually another application of the Norm.*
* Bivariate analysis is an important step in **data exploration**. We want to study the relationship between pairs of variables. Covariance or Correlation are measures used to study relationships between **two continuous variables**.
* One of the most common classification algorithms that regularly produces impressive results. It is an application of the concept of **Vector Spaces** in Linear Algebra.
* Principal Component Analysis, or PCA, is an unsupervised dimensionality reduction technique. PCA finds the **directions of maximum variance** and projects the data along them to reduce the dimensions. Without going into the math, these directions are the [**eigenvectors**](https://www.youtube.com/watch?v=PFDu9oVAE-g) **of the covariance matrix** of the data.
* Machine learning algorithms cannot work with raw textual data. We need to convert the text into some numerical and statistical features to create model inputs. Word Embeddings is a way of representing words as **low dimensional vectors** of numbers while preserving their context in the document.
* Latent Semantic Analysis (LSA), or Latent Semantic Indexing, is one of the techniques of Topic Modeling. It is another application of **Singular Value Decomposition**.
* In Computer Vision the image is represented as a 3d Tensor

By going through this list of a few applications of Linear Algebra you must have guessed the importance of Linear Algebra. As I rightly read somewhere, if Machine Learning is Batman then Linear Algebra is Robin.

## Combinatorics

Combinatorics concerns the number of configurations that can be obtained from certain discrete experiments. It is useful because it provides a systematic way of enumerating cases. Combinatorics often becomes very challenging as the complexity of the event grows.

To motivate the discussion of combinatorics, let us start with the following problem. Suppose there are 50 people in a room. What is the probability that at least one pair of people have the same birthday (month and day)? (We exclude Feb. 29 in this problem.)

The first thing you might be thinking is that since there are 365 days, we need at least 366 people to ensure that one pair has the same birthday. Therefore, the chance that 2 of 50 people have the same birthday is low. This seems reasonable, but let’s do a simulated experiment. In Figure 1.16 we plot the probability as a function of the number of people. For a room containing 50 people, the probability is 97%. To get a 50% probability, we just need 23 people! How is this possible?

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FFMMyJ6y6dzeEwlmH8IJq%2Fimage.png?alt=media&amp;token=6eb6bd68-687f-4dd4-94c6-8eeda7ee8b80" alt=""><figcaption><p>The probability for two people in a group to have the same birthday as a function of the number of people in the group.</p></figcaption></figure>

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2Fkib55UinBYqgLUzJuybk%2Fimage.png?alt=media&amp;token=ef549eef-82fd-44dc-9b44-4d26981057bf" alt=""><figcaption><p>The probability for two people to have the same birthday as a function of the number of people in the group. When there is only one person, this person can land on any of the 365 days. When there are two people, the first person has already taken one day (out of 365 days), so the second person can only choose 364 days. When there are three people, the first two people have occupied two days, so there are only 363 days left. If we generalize this process, we see that the number of configurations is 365 × 364 × · · · × (365 − k + 1), where k is the number of people in the room.</p></figcaption></figure>

So imagine that you keep going down the list to the 50th person. The probability that none of these 50 people will have the same birthday is as little as 3%. If you take the complement, you can show that with 97% probability, there is at least one pair of people having the same birthday.

Why is the probability so high with only 50 people while it seems that we need 366 people to ensure two identical birthdays? The difference is the notion of <mark style="color:yellow;">probabilistic</mark> and <mark style="color:yellow;">deterministic</mark>. The 366-people argument is deterministic. If you have 366 people, you are certain that two people will have the same birthday. This has no conflict with the probabilistic argument because the probabilistic argument says that with 50 people, we have a 97% chance of getting two identical birthdays. With a 97% success rate, you still have a 3% chance of failing. It is unlikely to happen, but it can still happen. The more people you put into the room, the stronger guarantee you will have. However, even if you have 364 people and the probability is almost 100%, there is still no guarantee. So there is no conflict between the two arguments since they are answering two different questions.

Permutations and combinations are two ways to enumerate all the possible cases. While the <mark style="color:yellow;">conclusions are probabilistic</mark>, as the birthday paradox shows, <mark style="color:yellow;">permutation and combination are deterministic</mark>. We do not need to worry about the distribution of the samples, and we are not taking averages of anything. Thus, modern data analysis seldom uses the concepts of permutation and combination.&#x20;

Does it mean that combinatorics is not useful? Not quite, because it still provides us with powerful tools for theoretical analysis. For example, in binomial random variables, we need the concept of combination to calculate the repeated cases. The Poisson random variable can be regarded as a limiting case of the binomial random variable, and so combination is also used. <mark style="color:yellow;">Therefore, while we do not use the concepts of permutation per se, we use them to define random variables</mark>.

## Conclusion

In conclusion we would like to highlight the significance of the birthday paradox. Many of us come from an engineering background in which we were told to ensure reliability and guarantee success. We want to ensure that the product we deliver to our customers can survive even in the worst-case scenario. We tend to apply deterministic arguments such as requiring 366 people to ensure complete coverage of the 365 days. In modern data analysis, the worst-case scenario may not always be relevant because of the complexity of the problem and the cost of such a warranty. The probabilistic argument, or the average argument, is more reasonable and cost-effective, as you can see from our analysis of the birthday problem. <mark style="color:yellow;">The heart of the problem is the trade-off between how much confidence you need versus how much effort you need to expend.</mark> Suppose an event is unlikely to happen, but if it happens, it will be a disaster. In that case, you might prefer to be very conservative to ensure that such a disaster event has a low chance of happening. Industries related to risk management such as insurance and investment banking are all operating under this principle.


# Probability Basics

Probability theory is the mathematical foundation of statistical inference, which is indispensable for analyzing data affected by chance, and thus essential for data scientists.

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2Fgit-blob-92426ecbe13821d5306c61a777c26729c723f883%2Fprob%20basic%20cartoon?alt=media" alt=""><figcaption></figcaption></figure>

Probability theory is the mathematical framework that allows us to analyze chance events in a logically sound manner. The probability of an event is a number indicating how likely that event will occur.

Note that when we say the probability of a head is 1/2, we are not claiming that any sequence of coin tosses will consist of exactly 50% heads. If we toss a fair coin ten times, it would not be surprising to observe 6 heads and 4 tails, or even 3 heads and 7 tails. But as we continue to toss the coin over and over again, we expect the long-run frequency of heads to get ever closer to 50%.

**In general, it is important in statistics to understand the distinction between theoretical and empirical quantities. Here, the true (theoretical) probability of a head was 1/2, but any realized (empirical) sequence of coin tosses may have more or less than exactly 50% heads.**

## Common Terminologies

The **sample space is the set of all possible outcomes in the experiment**: for some dice $$Ω = {1, 2, 3, 4, 5, 6}$$.

Any **subset of Ω is a valid event**. we can speak of the event $$F$$ of rolling a 4, $$F = {4}$$.

Consider the outcome of a single die roll and call it $$X$$. A reasonable question one might ask is “What is the average value of $$X$$?". We define this notion of “average” as a weighted sum of outcomes. This is called the **expected value**, or expectation of $$X$$, denoted by $$E(X)$$,

$$
Weighted Average = \frac{1}{6} \* (1 + 2 + 3 + 4 + 5 + 6) = 3.5
$$

If you play the game $$\infty$$ times the average value becomes $$E(X)$$

The **variance** of a random variable $$X$$ is a nonnegative number that summarizes on average how much $$X$$ differs from its mean, or expectation. The square root of the variance is called the **standard deviation.**

$$
Var(X) = \frac{(1−3.5)^2+(2−3.5)^2+(3−3.5)^2+(4−3.5)^2+(5−3.5)^2+(6−3.5)^2}{6} = \frac{17.5}{6}
$$

## Set

A set, broadly defined, is a collection of objects. In the context of probability theory, we use set notation to specify compound events. For example, we can represent the event roll an even number by the set {2, 4, 6}.

## Permutation and Combination

It can be surprisingly difficult to count the number of sequences or sets satisfying certain conditions. This is where **Premutation and Combination** comes in. For example, consider a bag of marbles in which each marble is a different color. If we draw marbles one at a time from the bag without replacement, how many different ordered sequences (permutations) of the marbles are possible? How many different unordered sets (combinations)?

* Permutation($$AB \neq BA$$ , order matters) = $$nPr = \frac{n!}{(n-r)!}$$
* Combination ($$AB = BA$$, order does not matter) = $$nCr = \frac{n!}{r!(n-r)!}$$

## Joint & Conditional Probability

* Joint Probability is the probability of two independent events occurring: $$P(A \cap B) = P(A)\*P(B)$$
* Conditional probability tells the probability of $$B$$ given $$A$$ has occurred, it allows us to account for information we have about our system of interest: $$P(B|A) = \frac{P(A \cap B)}{P(A)}$$

**If both are same, then A and B are independent events.**

## Bayes' Theorem

Bayes' theorem, named after 18th-century British mathematician Thomas Bayes, is a mathematical formula for determining conditional probability. **Conditional probability is the likelihood of an outcome occurring, based on a previous outcome occurring.**

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FZdp7UlQxqiQMfbC4eMWj%2Fimage1.png?alt=media&amp;token=c52f3215-fb6f-47c4-b0f6-2eb167493603" alt=""><figcaption><p>Baye's Theorem</p></figcaption></figure>

An easy way of remembering it is using the below example:

What is the probability of a fruit being banana given that it is long and yellow?

$$
P(Banana|Long,Yellow) = \frac{P(Long|Banana)\*P(Yellow|Banana)\*P(Banana)}{P(Long)\*P(Yellow)}
$$

## MAP vs MLE

The Maximum Aposteriori Probability (MAP) Estimation of the random variable y, given we have observed IID $$(x\_1, x\_2, x\_3, ... )$$ here we try to accommodate our prior knowledge when estimating. In Maximum Likelihood Estimation (MLE), we assume we don’t have any prior knowledge of the quantity being estimated.

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2Fgit-blob-bfcb77d121a651d606f85630af046419029ab3c4%2FMAP%20vs%20MLE?alt=media" alt=""><figcaption><p>MAP vs MLE</p></figcaption></figure>

## Questions

<details>

<summary>[UBER] Dice in increasing order</summary>

We throw 3 dice one by one. What is the probability that we obtain 3 points in strictly increasing order?

**Answer**

Suppose we get $$4$$ in the first roll then,

Total Probability = $$P(4) \* P(5) \* P(6) = 1/6 \* 1/6 \* 1/6 = 1/216$$

Similarly for $$3$$, $$P(3) \* P(4,5 | 4,6 | 5,6) = 1/6 \* (1/36 + 1/36 + 1/36) = 3/216$$

Taking into consideration $$P(1)$$ and $$P(2)$$ we have the total as $$= 10/216 + 6/216 + 3/216 + 1/216 = 20/216$$

</details>

<details>

<summary>[LINKEDIN] Cards in increasing order</summary>

Imagine a deck of 500 cards numbered from 1 to 500. If all the cards are shuffled randomly and you are asked to pick three cards, one at a time, what's the probability of each subsequent card being larger than the previous drawn card?

**Answer**

It is actually easy to solve this if you think on it a little. Let's pick any $$3$$ cards, now if you rearrange it there will only be $$1$$ way in which each subsequent card is larger the previous card. So, a total of $$6$$ \*\*\*\* ways to arrange the cards out of which only $$1$$ is valid. So the result is $$\frac{1}{6}$$.

</details>

<details>

<summary>[STATE FARM] Cards without replacement</summary>

Pull 2 cards from a deck without replacement what is probability *that both are of different colors.*

There can be many variants to this question.

**Answer**

[Source](https://www.quora.com/Two-cards-are-drawn-for-a-pack-of-52-cards-What-is-the-probability-that-both-the-cards-are-of-the-same-colour)

Here it is not specified which color the cards should be - they can be either red or black.

The probability that the first card drawn is either red or black is $$1$$ since these two are the only possible outcomes.

After the first draw, the total number of cards remaining in the pack is $$51$$, out of which $$25$$ cards are of the same colour as that of the card that is already drawn. Hence the probability of drawing a card of the same colour as the first one is $$\frac{25}{51}$$.

⇒ The probability of drawing two cards of the same colour is $$1\*\frac{25}{51}=\frac{25}{51}$$.

*Another approach to this can be:*

Two cards of a particular color can be drawn in $$C(26,2)$$ ways.

⇒ Two cards of either red or black can be drawn in $$2×C(26,2)$$ ways.

The total number of ways of drawing any two cards from the pack is $$C(52,2)$$.

⇒ The probability of drawing two cards of the same colour is $$\frac{2×C(26,2)}{C(52,2)} = \frac{2×26!}{2!×24!}\frac{2!×50!}{52!}=\frac{25}{51}$$

</details>

<details>

<summary>[FACEBOOK] N Dice</summary>

Suppose you're playing a dice game. You have 2 dice.

* What's the probability of rolling at least one 3?
* What's the probability of rolling at least one 3 given N die?

**Answer**

P(at least 1 three) = P(exactly 1 three) + P(2 three) = 1/6 \* 5/6 + 5/6 \* 1/6 + 1/36 = 11/36

*Solution received from the community via* [*mail*](mailto:thedatascienceinterviewbook@gmail.com)

To count the number of ways to throw at least $$1$$ three for $$N$$ dice, you need to sum overall $$k$$, $$1\<k<=N$$, where $$k$$ is the number of threes you throw. For each $$k$$, there are $$C(N,k)$$ possible combinations of dice that are three. For each of these combinations, there are $$5$$ possible values for the other $$N-K$$ dice. So, the number of ways to throw $$k$$ threes with $$N$$ dice is $$5^{(N-k)}\*C(N,k)$$.

The total sum over $$1\<k<=N$$ is $$\sum\_{k=1}^N 5^{(N-k)} \begin{pmatrix} n\ k\ \end{pmatrix} = 6^N-5^N$$. Since there are $$6^N$$ ways to throw the dice, the probability is $$(6^N - 5^N)/6^N = 1 - (5/6)^N$$.

There is a simpler way to solve this problem: calculate the number of ways to not throw any threes, then subtract this number from the total number of ways to throw the dice. For $$N=2$$, this is $$1 - (5/6)^2 = 1 - 25/36 = 11/36$$. For $$N$$, it is $$1 - (5/6)^N$$ You can see that this is equivalent to the probability calculated using the above sum: $$1 - (5/6)^N$$.

**`Tip:`**` `` ``Check the general case for N=2 and see if the numbers match `

</details>

<details>

<summary>[FACEBOOK] 3 Zebras</summary>

Three zebras are chilling in the desert. Suddenly a lion attacks.

Each zebra is sitting on a corner of an equally length triangle. Each zebra randomly picks a direction and only runs along the outline of the triangle to either edge of the triangle.

What is the probability that none of the zebras collide?

**Answer**

Each zebra has 2 options of travel: clockwise or anticlockwise. So a total of $$2*2*2 = 8$$ options.

Out of this only way in which they donot collide is if all of them travel clockwise or anticlockwise. So a total of $$2$$.

Therefore the probability of no collision $$= 2/8 = 25%$$

</details>

<details>

<summary>[POSTMATES] Four Person Elevator</summary>

There are four people on the ground floor of a building that has five levels not including the ground floor. They all get into the same elevator.

If each person is equally likely to get on any floor and they leave independently of each other, what is the probability that no two passengers will get off at the same floor?

**Answer**

The number of ways to assigning five floors to four different people is to get the total sample space. In this case it would be $$5 \* 5 \* 5 \* 5$$.

The number of ways to assign five floors to four people without repetition of floors is $$5 \* 4 \* 3 \* 2$$ because for the first passenger you have five different options. The second person has four, and so on. Note that this number counts all possible orders between passengers as well.

The result is then $$\frac{5 \* 4 \* 3 \* 2}{5 \* 5 \* 5 \* 5} = 0.192$$

</details>

<details>

<summary>[AMAZON] Found Item</summary>

Amazon has a warehouse system where items on the website are located at different distribution centers across a city. Let's say in one example city, the probability that a specific item X at location A is 0.6, and at location B the probability is 0.8.

Given you're a customer in this example city and the items are only found on the website if they exist in the distribution centers, what is the probability that the item X would be found on Amazon's website?

**Answer**

Probability of the item being present= $$1-$$ p(item NOT in A AND NOT in B) $$= 1-(0.4\*0.2)=0.92$$

</details>

<details>

<summary>[SPOTIFY] Max Dice Roll</summary>

A fair die is rolled $$n$$ times. What is the probability that the largest number rolled is $$r$$, for each $$r$$ in $$1..6$$?

**Answer** If $$r(1≤r≤6)$$ is the largest number you have allowed for your $$n$$ rolls, then you forbid any number larger than $$r$$. That is, you forbid $$6−r$$ values. The probability that your single roll does not show any of these $$6−r$$ values is $$\frac{6−r}{6}$$ and the probability that this happens each time during a series of $$n$$ rolls is the obviously $$(\frac{6−r}{6})^n$$

There is a subtle nuance to this problem, in the above solution we have assumed the $$max<=r$$ which is different from $$max=r$$ or in other words if $$r=3$$, the above solution gives results for $$r= 1,2,3$$. The solution of $$r=3$$ is a little more involved:

Let's take $$r=3$$, for $$n$$ die rolls we should have atleast one $$r$$. The Probability of that is:

$$P(r=3)$$ $$= P(\text{of getting all n values as 1,2,3} \* P(\text{atleast one 3}))$$ $$= (\frac{3}{6})^n \* (1-P(\text{no 3's occuring}))$$$$= (\frac{3}{6})^n \* (1-(\frac{\text{only getting 1,2}}{\text{out of 1,2,3}})^n)$$$$= (\frac{3}{6})^n \* (1-(\frac{2}{3})^n)$$$$= \text{generalizing } (\frac{r}{6})^n \* (1-(\frac{r-1}{r})^n)$$$$= \frac{r^n - (r-1)^n}{6^n}$$

</details>

<details>

<summary>[FACEBOOK] Labeling Content</summary>

Facebook has a content team that labels pieces of content on the platform as spam or not spam. $$90%$$ of them are diligent raters and will label $$20%$$ of the content as spam and $$80%$$ as non-spam. The remaining $$10%$$ are non-diligent raters and will label $$0%$$ of the content as spam and $$100%$$ as non-spam. Assume the pieces of content are labeled independently from one another, for every rater. Given that a rater has labeled $$4$$ pieces of content as good, what is the probability that they are a diligent rater?

**Answer**

This can be solved using Baye's theorem:

* Not Spam = $$NS$$
* Spam = $$S$$
* Diligent =$$D$$
* NotDiligent =$$ND$$

$$P(D|NS, NS, NS, NS) = \frac{P(NS, NS, NS, NS|D)*P(D)}{P(NS, NS, NS, NS|D)*P(D)+P(NS, NS, NS, NS|ND)*P(ND)}$$ $$P(D|NS, NS, NS, NS) = \frac{0.8^4*0.9}{0.8^4*0.9+1^4*0.1}$$ = \~$$0.787$$

</details>

<details>

<summary>[FACEBOOK] Raining</summary>

You are about to get on a plane to Seattle. You want to know if you should bring an umbrella. You call $$3$$ random friends of yours who live there and ask each independently if it's raining. Each of your friends has a $$2/3$$ chance of telling you the truth and a $$1/3$$ chance of messing with you by lying. All $$3$$ friends tell you that "Yes" it is raining.

What is the probability that it's actually raining in Seattle?

**Answer**

Even though the problem is straightforward one can interpret the problem in many ways. Taking a Bayesian approach is probably appropriate in a real world sense, but if you are told by the interviewer you have no ability to determine the priors, you can't use Bayesian. [Check this thread](https://math.stackexchange.com/questions/1335235/facebook-question-data-science) for a detailed discussion on this problem.

For it to be not raining, all friends must be lying. Therefore, the solution must be the inverse of the probability that all three are "messing with you." $$(1/3)*(1/3)*(1/3)=1/27$$ (3.7% chance they are all lying).

Since there is only a $$3.7%$$ chance all three friends are messing with you, there is a $$96.3%$$ chance it is raining.

</details>

<details>

<summary>[MICROSOFT] First to Six</summary>

Amy and Brad take turns in rolling a fair six-sided die. Whoever rolls a $$6$$ first wins the game. Amy starts by rolling first.

What's the probability that Amy wins?

**Answer**

Amy can win on the first roll, third roll, fifth roll, and so on.

Probability of Amy winning in the first roll = P(six rolled by her) = $$1/6$$

Probability of Amy winning in the third roll = P(six NOT rolled by her in first try) \* P(six NOT rolled by Brad in first try) \* P(six rolled by her in 2nd try) = $$(5/6) \* (5/6) \* (1/6) = 1/6 \* (5/6)^2$$

Similarly, the probability of Amy winning in the fifth roll = $$(1/6) \* (5/6)^4$$

Similarly, the probability of Amy winning in the seventh roll = $$(1/6) \* (5/6)^6$$

Hence, total probability of Amy winning = Sum of all such events = $$(1/6) + (1/6 \* (5/6)^2) + (1/6 \* (5/6)^4) + (1/6 \* (5/6)^6) + ...$$

The sum of such an infinite Geometric Progression series is = $$\frac{a}{1-r} = (1/6) / (1 - 25/36) = (1/6) / (11/36) = 6/11$$

Hence, probability of Amy winning in any of her turns = $$6/11$$

</details>

<details>

<summary>[GOOGLE][FACEBOOK] Double Sided Coin</summary>

A jar has $$1000$$ coins, of which $$999$$ are fair and $$1$$ is double headed. Pick a coin at random, and toss it $$10$$ times. Given that you see $$10$$ heads, what is the probability that the next toss of that coin is also a head?

**Answer**

There are two ways of choosing the coin. One is to pick a fair coin and the other is to pick the one with two heads.

* Probability of selecting fair coin $$= 999/1000 = 0.999$$
* Probability of selecting unfair coin $$= 1/1000 = 0.001$$

Selecting $$10$$ heads in a row = Selecting fair coin \* Getting 10 heads + Selecting an unfair coin

* P (A) $$= 0.999 \* (1/2)^5 = 0.999 \* (1/1024) = 0.000976$$
* P (B) $$= 0.001 \* 1 = 0.001$$
* P( A / A + B ) $$= 0.000976 / (0.000976 + 0.001) = 0.4939$$
* P( B / A + B ) $$= 0.001 / 0.001976 = 0.5061$$

Probability of selecting another head $$= P(A/A+B) \* 0.5 + P(B/A+B) \* 1 = 0.4939 \* 0.5 + 0.5061 = 0.7531$$

</details>

<details>

<summary>[GOOGLE] Forming a triangle</summary>

A 10 feet pole is randomly cut into 3 pieces. What is the probability that exactly form a triangle?

**Answer**

[Source](https://www.quora.com/A-10-feet-pole-is-randomly-cut-into-3-pieces-What-is-the-probability-that-exactly-form-a-triangle)

Suppose the pole is $$AB$$ and there are two points $$P$$ and $$Q$$ such that $$AP = x$$ and $$PQ = y$$, so that $$QB = 10 – x – y$$ as we know that sum of two sides of a triangle is greater than the 3rd side. Hence

$$x + y > (10 – x – y)$$ or $$x + y > 5$$

$$y + (10 – x – y) > x$$ or $$x < 5$$

$$(10 – x – y) > y$$ or $$y < 5$$

Also, we know that all the parts of pole must be greater than $$0$$,

or $$x > 0, y > 0, 10 – x – y > 0 or x > 0, y > 0, x + y < 10$$

Plotting the lines $$x + y = 10 x + y = 5, x = 5, y = 5$$. Now favorable area is the area of the middle red shaded triangle.

Required probability $$= 1/4$$

<img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2Fgit-blob-2c655908781aadfeab0b82b2c707581e6cb85917%2Fimage13.png?alt=media" alt="" data-size="original">

</details>

<details>

<summary>[LYFT] Flips until two heads</summary>

What is the expected number of coin flips needed to get two consecutive heads?

**Answer**

[Source](http://www.codechef.com/wiki/tutorial-expectation):

Let's first assume $$x$$ is the expected number of coin flips required for getting two heads in a row. Now:

* If the first flip turns out to be tail you need $$x$$ more flips since the events are independent. Probability of the event $$1/2$$. Since $$1$$ flip was wasted total number of flips required $$(1+x)$$.
* If the first flip becomes head, but the second one is tail($$HT$$) - $$2$$ flips are wasted, here total number flips required would be $$(2+x)$$. Probability of $$HT$$ out of $$HH, HT, TH, TT$$ is $$(1/4)$$
* The best case, the first two flips turn out to be heads both($$HH$$). Probability, $$1/4$$ i.e. $$HH$$ out of $$HH, HT, TH, TT$$. No of flips required $$2$$.

So from the above scenarios, $$x = 1/2(1+x) + 1/4(2+x) + (1/4 )\* 2$$ $$= 1/2 \[ (1+x) + 1/2(2+x) + 1 ]$$ $$= 1/2 \[ 1 + x + 1 + x/2 + 1 ]$$

$$x / 4 = 3/2$$ $$x = 6$$

So the expected number of flips would be $$6$$

</details>

<details>

<summary>[LYFT] Number of cards before an ace</summary>

How many cards would you expect to draw from a standard deck before seeing the first ace?

**Answer**

[Source](https://aksoy.io/math/probability/book/2020/03/10/problem-40-first-ace.html):

Let $$X$$ represent the number of cards that are turned up to produce the $$1^{st}$$ ace. For this problem, we cannot apply the Geometric Distribution because cards are sampled without replacement.

Instead, we begin by considering the probabilities of drawing the $$1^{st}$$ ace on the $$1^{st}$$ card, $$2^{nd}$$ card, and so on:

$$P(1^{st} card)= \frac{4}{52}$$

$$P(2^{nd} card)= \frac{48}{52}\frac{4}{51}$$

$$P(3^{rd} card)= \frac{48}{52}\frac{47}{51}\frac{4}{50}$$

$$P(n^{th} card)= 4\* \frac{48!}{(49-x)!}\frac{(52-x)!}{(52)!}$$

With this we can calculate the average number of cards by applying the definition of expected value:

$$E\[X]= \sum\limits\_{x=1}^{52} 4x \frac{48!}{(49-x)!}\frac{(52-x)!}{(52)!} = \frac{53}{5} = 10.6$$

</details>

<details>

<summary>[FACEBOOK] Ad Raters</summary>

Let’s say we use people to rate ads.

There are two types of raters. Random and independent from our point of view:

80% of raters are careful and they rate an ad as good (60% chance) or bad (40% chance). 20% of raters are lazy and they rate every ad as good (100% chance).

* Suppose we have 100 raters each rating one ad independently. What’s the expected number of good ads?
* Now suppose we have 1 rater rating 100 ads. What’s the expected number of good ads?
* Suppose we have 1 ad, rated as bad. What’s the probability the rater was lazy?

**Answer**

* 100 raters are divided into 2 groups according to probabilities:

$$20$$ lazy raters: $$100%$$ good ads -> $$20$$ good ads; $$80$$ careful raters: $$60%$$ good ads -> $$80 \* 0.6 = 48$$ good ads. Total $$68$$ good ads.

* There could be 2 cases:

Random rater is careful with probability of $$0.8: 0.8 \* 0.6 = 0.48$$ - probability or rating good ad Random rater is lazy with probability of $$0.2: 0.2 \* 1 = 0.2$$ - probability or rating good ad Total probability of rating ad as good is $$0.48+0.2 = 0.68$$. The expected amount of good rates $$100\*0.68 = 68$$.

* It’s $$0$$ probability that the rater is lazy because lazy raters always rate ads as good.

</details>

<details>

<summary>[INTUIT] Chances of Winning</summary>

Let’s say you can play a coin flipping guessing game either once or a 2 out of 3 game. What is the best strategy for winning?

**Answer**

**Scenario 1: Single Coin Flip Game**

* In this game, you have a 50% chance of winning on each flip because there are two possible outcomes (heads or tails), and you need one specific outcome (either heads or tails) to win.

Probability of winning a single coin flip = 0.5 (50%)

**Scenario 2: 2 out of 3 Coin Flip Game**

* In this game, you need to win at least two out of three coin flips to win the game. To calculate this probability, we can use the binomial probability formula.

Probability of winning two out of three coin flips:

* Calculate the probability of winning exactly two out of three flips:
  * P(Win 2 out of 3) = C(3, 2) \* (0.5)^2 \* (0.5)^(3-2) = 3 \* 0.25 \* 0.5 = 0.375
* Calculate the probability of winning all three flips:
  * P(Win 3 out of 3) = (0.5)^3 = 0.125

Now, sum these probabilities to get the overall probability of winning the 2 out of 3 game:

Probability of winning the 2 out of 3 game = P(Win 2 out of 3) + P(Win 3 out of 3) = 0.375 + 0.125 = 0.5 (50%)

So, the probability of winning in either a single coin flip game or a 2 out of 3 coin flip game is 50% for both scenarios.

</details>


# Probability Distribution

Knowing the distribution of data helps us better model the world around us. It helps us to determine the likeliness of various outcomes or make an estimate of the variability of an occurrence.

## Random Variable

**Random Variable maps the outcome of sample space into real numbers.**

Example: How many heads when we toss 3 coins?

$$X$$ could be $$0, 1, 2$$ or $$3$$ randomly, and they might each have a different probability. $$X$$ **= "The number of Heads" is the Random Variable.**

In this case, there could be 0 Heads (if all the coins land Tails up), 1 Head, 2 Heads or 3 Heads. So, the Sample Space = $${0, 1, 2, 3}$$ But this time the outcomes are NOT all equally likely. The three coins can land in eight possible ways:

Looking at the table we see just 1 case of Three Heads, but 3 cases of Two Heads, 3 cases of One Head, and 1 case of Zero Heads. So:

* $$P(X = 3) = 1/8 = {HHH}$$
* $$P(X = 2) = 3/8 = {HHT,HTH,THH}$$
* $$P(X = 1) = 3/8 = {TTH,THT,TTH}$$
* $$P(X = 0) = 1/8 = {TTT}$$

And this is what becomes the **probability distribution.**

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FIVJL1tRJeQkjNTON4nhR%2Fimage2.png?alt=media&amp;token=bc7e46b1-b220-4a05-b06a-071f23d93440" alt=""><figcaption></figcaption></figure>

**Frequency distribution** comes from actually doing the experiment $$n$$ number of times as $$n->\infty$$ the shape comes closer and closer to the Probability distribution

Now the probability distribution can be of $$2$$ types, **discrete and continuous**. An example of Discrete is shown above.

* When we use a probability function to describe a discrete probability distribution, we call it a **probability mass function (PMF)**. The probability mass function, $$f$$, just returns the probability of the outcome. Therefore, the probability of rolling a $$3$$ is $$f(3) = 1/6$$.
* When we use a probability function to describe a continuous probability distribution, we call it a **probability density function (PDF)**.

Now depending on the problem type one can choose the corresponding distribution and find the probability for some value of the random variable.

## Types of Distribution

Some common types of probability distribution are as follows:

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FnQKBdGEXfNBUjC3QhmbZ%2Fimage3.png?alt=media&amp;token=2a33cfb9-efeb-4474-9f89-7ca3313defb6" alt=""><figcaption><p>Types of Probability Distribution</p></figcaption></figure>

## Normal (Gaussian) Distribution

The normal distribution is the most important probability distribution in statistics because it fits many natural phenomena. For example, heights, blood pressure, measurement error, and IQ scores follow the normal distribution.

Despite the different shapes, all forms of the normal distribution have the following characteristic properties.

* They’re all symmetric. The normal distribution cannot model skewed distributions.
* The mean, median, and mode are all equal.
* Half of the population is less than the mean and half is greater than the mean.
* The Empirical Rule, which describes the percentage of the data that fall within specific numbers of standard deviations from the mean for bell-shaped curves.

| Mean +/- standard deviations | Percentage of data contained |
| ---------------------------- | ---------------------------- |
| 1                            | 68%                          |
| 2                            | 95%                          |
| 3                            | 99.7%                        |

{% embed url="<https://media.licdn.com/dms/document/media/D561FAQEhP422gUwoXA/feedshare-document-pdf-analyzed/0/1691029309693?e=1695859200&t=kQbuxjFthLaefzSxbieXPgK1FVGR6IsboKub-kyAUM8&v=beta>" %}
Thanks to[ Reza Bagheri](https://www.linkedin.com/in/reza-bagheri-71882a76/)
{% endembed %}

## Measures to understand a distribution:

There are 3 variety of measures, required to understand a distribution:

* Measure of Central tendency
* Measure of dispersion
* Measure to describe shape of curve

### Measure of Central Tendency

Measures of central tendencies are measures, which help you describe a population, through a single metric. For example, if you were to compare Saving habits of people across various nations, you will compare average Savings rate in each of these nations. Following are the measures of central tendency:

* **Mean:** or the average
* **Median:** the value, which divides the population in two half
* **Mode:** the most frequent value in a population

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2Fljfj9c89wud8A7hbKlIe%2Fimage4.png?alt=media&amp;token=d83fc937-c553-4697-8a35-3d560128a37d" alt=""><figcaption></figcaption></figure>

### Measure of Dispersion

Measures of dispersion reveal how is the population distributed around the measures of central tendency.

* **Range:** Difference in the maximum and minimum value in the population
* **Quartiles:** Values, which divide the population in 4 equal subsets (typically referred to as first quartile, second quartile and third quartile)
* **Inter-quartile range:** The difference in third quartile (Q3) and first quartile (Q1). By definition of quartiles, 50% of the population lies in the inter-quartile range.
* **Variance:** The average of the squared differences from the Mean.
* **Standard Deviation:** is square root of Variance

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FjEz6IKUDSvsYQ40NmodP%2Fimage5.png?alt=media&amp;token=4a14e410-46de-4586-ab90-dcfb271819c4" alt=""><figcaption><p>2 Distribution with different standard deviation</p></figcaption></figure>

### Measure to describe shape of distribution

* **Skewness:** Skewness is a measure of the asymmetry. Negatively skewed curve has a long left tail and vice versa.
* **Kurtosis:** Kurtosis is a measure of the “peaked ness”. Distributions with higher peaks have positive kurtosis and vice-versa

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FgjMHtz4GjlRM2cXofNjA%2Fimage6.png?alt=media&amp;token=bb198b2f-c7f5-4558-a9ab-957ad7f39a8e" alt=""><figcaption></figcaption></figure>

### Box Plots

Box plots are one of the easiest and most intuitive way to understand distributions. They show mean, median, quartiles and Outliers on single plot.

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FOQXSxLPiCpGX80zo3EQD%2Fimage7.png?alt=media&amp;token=190836e3-d494-4942-a957-a86952e17ca3" alt=""><figcaption></figcaption></figure>

## Unbiased Estimator

An unbiased estimator is an accurate statistic that’s used to approximate a population parameter. “Accurate” in this sense means that it’s neither an overestimate nor an underestimate. If an overestimate or underestimate does happen, the mean of the difference is called a “bias.” That’s just saying if the estimator (i.e., the sample mean) equals the parameter (i.e., the population mean), then it’s an unbiased estimator.

## Maximum Likelihood Estimation (MLE)

**Reference:** [Discussion](https://stats.stackexchange.com/questions/112451/maximum-likelihood-estimation-mle-in-layman-terms), [Explanation](https://www.kdnuggets.com/2019/11/probability-learning-maximum-likelihood.html), [Implementation](https://analyticsindiamag.com/maximum-likelihood-estimation-python-guide/)

Say you have some data. Say you're willing to assume that the data comes from some distribution -- perhaps Gaussian. There are an infinite number of different Gaussians that the data could have come from (which correspond to the combination of the infinite number of means and variances that a Gaussian distribution can have). MLE will pick the Gaussian (i.e., the mean and variance) that is "most consistent" with your data (the precise meaning of consistent is explained below).

So, say you've got a data set of $$y={−1,3,7}$$. The most consistent Gaussian from which that data could have come has a mean of $$3$$ and a variance of $$32/3$$. It could have been sampled from some other Gaussian. But one with a mean of $$3$$ and variance of $$32/3$$ is most consistent with the data in the following sense: the probability of getting the particular $$y$$ values you observed is greater with this choice of mean and variance, than it is with any other choice.

Will show the calculation step by step:

We have a dataset: $$y = {-1, 3, 7}$$

Assuming the data follows a normal distribution $$\mathcal{N}(\mu, \sigma^2)$$, the likelihood function is:

$$
L(\mu, \sigma^2) = \prod\_{i=1}^{n} \frac{1}{\sqrt{2\pi\sigma^2}} \exp\left(-\frac{(y\_i - \mu)^2}{2\sigma^2}\right)
$$

Taking the **log-likelihood**:

$$
\log L(\mu, \sigma^2) = -\frac{n}{2} \log (2\pi \sigma^2) - \frac{1}{2\sigma^2} \sum\_{i=1}^{n} (y\_i - \mu)^2
$$

The MLE for the mean of a Gaussian is the **sample mean**:

$$
\hat{\mu} = \frac{1}{n} \sum\_{i=1}^{n} y\_i
$$

Substituting the values: $$\hat{\mu} = \frac{-1 + 3 + 7}{3} = \frac{9}{3} = 3$$

The MLE for variance is: $$\hat{\sigma}^2 = \frac{1}{n} \sum\_{i=1}^{n} (y\_i - \hat{\mu})^2$$&#x20;

Substituting the values: $$\hat{\sigma}^2 = \frac{( -1 - 3)^2 + (3 - 3)^2 + (7 - 3)^2}{3}$$

However, if the **population variance formula** was used instead of the true MLE formula: $$\hat{\sigma}^2 = \frac{1}{n-1} \sum\_{i=1}^{n} (y\_i - \hat{\mu})^2$$ $$= \frac{32}{2} = 16$$

This would be the **unbiased variance estimator**, not the true MLE.

Maximum Likelihood Estimation can be applied to both regression and classification problems.

## Questions

<details>

<summary>[LIME] Example of unbiased estimator</summary>

What is an unbiased estimator and can you provide an example for a layman to understand?

**Answer**

One famous example of an unrepresentative sample is the literary digest voter survey, which predicted Alfred Landon would win the 1936 presidential election. The survey was biased, as it failed to include a representative sample of low income voters who were more likely to be democrat and vote for Theodore Roosevelt.

If the sampling had been done correctly then the estimator would have been unbiased as it would match with the actual output from the population, which was win for Theodore Roosevelt.

</details>

<details>

<summary>[GOOGLE] Median of Uniform Distribution</summary>

Given 3 i.i.d. variables from an uniform distribution of $$0$$ to $$4$$, what’s the chance the median is greater than $$3$$?

**Answer**

This will only be possible if atleast $$2$$ random variables are greater than $$3$$.

$$P(M>3) = P(GGL) + P(GLG) + P(LGG) + P(GGG) = 3 \* (1/4)^2 \* 3/4 + (1/4)^3 = 5/32$$ where, $$G$$ stands for probability of number $$> 3$$ which is probability of it being $$4$$ out of $$1,2,3,4 = 1/4$$; $$L$$ for probability of number $$< 3$$ which is probability of it being $$1,2,3$$ out of $$1,2,3,4 = 3/4$$

</details>

<details>

<summary>[SPOTIFY] MLE of Uniform Distribution</summary>

Suppose you draw n samples from a uniform distribution U(a, b). What is the MLE estimate of a and b?

**Answer**

*Solution recieved from the community via* [*merge request*](https://github.com/dipranjan/dsinterviewqns/pull/4)

Let $$x\_1, x\_2, \ldots , x\_n$$ be the $$n$$ samples drawn.

Recall the pdf for the uniform distribution function is:

$$f(x)=\frac{1}{b-a}$$

Thus, the likelihood function $$\mathcal{L}$$ is simply the product of the pdf n times, which is:

$$f(x)=\frac{1}{(b-a)^n}.$$

The MLE will occur at the values of $$a$$ and $$b$$ for which that quantity is maximized. Since $$(b-a)^n$$ is in the denominator, and $$b-a$$ must always be positive because $$b>a$$, the likelihood is maximized when $$b-a$$ is minimized. This means we want $$a$$ as big as possible, and $$b$$ as small as possible. But for one of the $$x\_i$$ to be sampled, $$a$$ must be smaller than that value, (and $$b$$ must be larger), so the maximum likelihood estimation is $$a=\min(x\_1, x\_2, \ldots , x\_n)$$ and $$b=\max(x\_1, x\_2, \ldots , x\_n)$$

</details>

<details>

<summary>[MCKINSEY] Flipping 576 Times</summary>

You flip a fair coin 576 times. Without using a calculator, calculate the probability of flipping at least 312 heads.

**Answer**

Fair coin, $$p(H)=0.5$$ Since this experiment has only $$2$$ outcomes hence we can use a binomial distribution,

mean = $$np$$ = $$576*0.5 = 288$$, var= $$np(1-p)= 576*0.5\*0.5 = 144$$, stddev = sqrt(var) = $$12$$.

For normal distribution, *68% of the data falls within one standard deviation, 95% percent within two standard deviations, and 99.7% within three standard deviations from the mean.*

$$312= 288$$(mean)$$+2\*12$$(stddev), which means the probability of flipping at least $$312$$ heads or tails is $$5%$$. Since we are only looking at the probability of at least $$321$$ heads, it is the right tail area of the distribution, which is $$5%/2= 2.5%$$. So the probability of flipping at least $$312$$ heads is $$2.5%$$.

</details>

<details>

<summary>[GOOGLE] Non-normal Probability Distribution</summary>

Explain how a probability distribution could be not normal and give an example scenario.

**Answer**

[Source](https://www.interviewquery.com/questions/non-normal-probability-distribution?ref=question_email)

Normal probability distributions are characterized by their famous bell shaped probability density function. The observations are centered around the mean and are equally spread around as per the standard deviation of the distribution, in case the probability distribution is a standard normal. They occur frequently in the nature, for e.g. distribution of heights

There are other types of distributions which are not normal; since normal distributions are for continuous random variable, all discrete random variables do not follow normal distributions.

There can be many examples of Non-Normal distribution:

* Flip a coin ten times and count the number of heads you get. That follows a binomial distribution
* Flip a coin until you get five heads and count the number of flips. That follows a negative binomial distribution
* Take a well-shuffled deck of cards and count how many red cards there are in the first ten. That follows a hypergeometric distribution

</details>


# Central Limit Theorem

The theorem gives us the ability to quantify the likelihood that our sample will deviate from the population without having to take any new sample to compare it with.

## Sampling distribution

Let's start with an example, suppose from the SAT math scores

* You take a sample of $$10$$ random students from a population of $$100$$. You might get a mean of $$502$$ for that sample.
* Then, you do it again with a new sample of $$10$$ students. You might get a mean of $$480$$ this time.
* Then, you do it again. And again. And again...... and get the following means for each of those three new samples of $$10$$ people: $$550$$, $$517$$, $$472$$

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2Fgit-blob-dde60767fa96a3e74cc7cb560e155e0780d28ab0%2Fimage8.png?alt=media" alt=""><figcaption></figcaption></figure>

The sampling distribution, which is basically the distribution of sample means of a population, has some interesting properties which are collectively called the **central limit theorem**, *which states that no matter how the original population is distributed, the sampling distribution will follow these three properties* –

* Sampling Distribution’s Mean $$(\mu\_{\bar{x}}) =$$ Population Mean $$(\mu)$$
* Sampling Distribution’s Standard Deviation (Standard Error) $$= \sigma\sqrt{n}$$, where $$\sigma$$ is the population’s standard deviation and $$n$$ is the sample size
* For $$n > 30$$, the sampling distribution becomes a normal distribution. Strongly skewed distributions can require larger sample sizes.

To prove the thrid point let's take a uniform distribution, in the image below $$500,000$$ times for each sample size $$(5, 20, 40)$$ have been drawn and their mean plotted. We’d expect the average to be $$(1 + 2 + 3 + 4 + 5 + 6 / 6 = 3.5)$$. The sampling distributions of the means center on this value. Just as the central limit theorem predicts, as we increase the sample size, the sampling distributions more closely approximate a normal distribution and have a tighter spread of values.

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2Fgit-blob-4c8148f948cf48e236ed4786948bd2560ac56a10%2Fimage9.png?alt=media" alt=""><figcaption></figcaption></figure>

### Confidence Interval <a href="#what_is_confidence_interval" id="what_is_confidence_interval"></a>

([Source](https://www.simplilearn.com/tutorials/data-analytics-tutorial/confidence-intervals-in-statistics))

A confidence interval shows the probability that a parameter will fall between a pair of values around the mean. Confidence intervals show the degree of uncertainty or certainty in a sampling method. They are constructed using confidence levels of 95% or 99%.

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2F6vu1mipix0gtRk6s9CWV%2Fimage.png?alt=media&amp;token=461a1bc9-961e-4d46-876a-a2cb6f1cd188" alt=""><figcaption></figcaption></figure>

The formula for CI is given by: (*mean – (z\* (std\_dev/sqrt(n))*

#### Calculating A Confidence Interval <a href="#calculating_a_confidence_interval" id="calculating_a_confidence_interval"></a>

Imagine a group of researchers who are trying to decide whether or not the oranges produced on a certain farm are large enough to be sold to a potential grocery chain. This will serve as an example of how to compute a confidence interval.

#### Step 1: Determine the sample size (n).

46 oranges are chosen at random by the researchers from farm trees. Consequently, n is 46.

#### Step 2: Determine the samples' means (x).

The researchers next determine the sample's mean weight, which comes out to be 86 grammes.

X = 86.

#### Step 3: Determine the standard deviation (s).

Although utilising the population-wide standard deviation is ideal, this data is frequently unavailable to researchers. If this is the case, the researchers should apply the sample's determined standard deviation.

Let's assume, for our example, that the researchers have chosen to compute the standard deviation from their sample. They get a 6.2-gramme standard deviation.

S = 6.2.

#### Step 4: Determine the confidence interval

In ordinary market research studies, 95% and 99% are the most popular selection for confidence intervals. For this example, let's assume that the researchers employ a 95% confidence interval.

#### Step 5: Find the Z value for the chosen confidence interval

The researchers would subsequently use the following table to establish their Z value:

Confidence Interval Z

80% 1.282

85% 1.440

90% 1.645

95% 1.960

99% 2.576

99.5% 2.807

99.9% 3.291

#### Step 6: Calculate the following formula

The next step would be for the researchers to enter their known values into the formula. Following our example, this formula would look like this:

86 ± 1.960 (6.2/6.782). This calculation yields a value of 86±1.79, which the researchers use as their confidence interval.

#### Step 7: Come to a decision.

According to the study's findings, the real mean of the larger population of oranges is probably (with a 95% confidence level) between 84.21 grammes and 87.79 grammes.


# Bayesian vs Frequentist Reasoning

Bayesian statistics concerns itself with trying to represent your beliefs in a DeFinetti consistent manner, simply put - you want a logically rigorous way of describing your initial (prior) belief and the update in your belief as you observe new data (thus creating posterior belief).

Frequentist statistics concerns itself with methods that have **long run** guarantees.

E.g., *If a person shoots bullseye of the target 9 times, the frequentist approach predicts that 10th shot will also be bullseye. Whereas, as a human we are prone to making some error. The Bayesian approach will take this prior behavior into account. It will not predict 1 as the probability of hitting bullseye in 10th shot. It will take into account prior shots in similar situations and predict accordingly.*

*In case of Web data, if a person has liked something 9 times, frequentist will predict that the person will like it 10th time as well. Whereas, Bayesian approach knows that humans can get bored. Liking something 9 times does not guarantee that it will be liked 10th time as well.*


# Hypothesis Testing

Hypothesis testing is the process used to evaluate the strength of evidence from the sample and provides a framework for making determinations related to the population

## Inferential Statistics

Sometimes, you may require a very large amount of data for your analysis which may need too much time and resources to acquire. In such situations, you are forced to work with a smaller sample of the data, instead of having the entire data to work with.

Situations like these arise all the time at big companies like Amazon. For example, say the Amazon QC department wants to know what proportion of the products in its warehouses are defective. Instead of going through all of its products (which would be a lot!), the Amazon QC team can just check a small sample of 1,000 products and then find, for this sample, the defect rate (i.e. the proportion of defective products). Then, based on this sample's defect rate, the team can "infer" what the defect rate is for all the products in the warehouses. **This process of “inferring” insights from sample data is called “Inferential Statistics”.**

## Hypothesis Testing

Hypothesis testing is used to confirm your conclusions about the population parameter. Through this we can conclude if there is enough evidence to confirm the hypothesis about the population.

* $$H\_0 =$$ null hypothesis, what is already present; always has the following signs: = OR ≤ OR ≥
* $$H\_1 =$$ alternate hypothesis, a challenge to the null hypothesis; always has the following signs: ≠ OR > OR <

### Steps

[📖Source](https://en.wikipedia.org/wiki/Statistical_hypothesis_testing)

* There is an initial research hypothesis of which the truth is unknown.
* The first step is to state the relevant null and alternative hypotheses. This is important, as mis-stating the hypotheses will muddy the rest of the process.
* The second step is to consider the statistical assumptions being made about the sample in doing the test; for example, assumptions about the statistical independence or about the form of the distributions of the observations. This is equally important as invalid assumptions will mean that the results of the test are invalid.
* Decide which test is appropriate, and state the relevant test statistic $$T$$.
* Derive the distribution of the test statistic under the null hypothesis from the assumptions. In standard cases this will be a well-known result. For example, the test statistic might follow a Student's t distribution with known degrees of freedom, or a normal distribution with known mean and variance. If the distribution of the test statistic is completely fixed by the null hypothesis, we call the hypothesis simple, otherwise it is called composite.
* Select a significance level ($$\alpha$$), a probability threshold below which the null hypothesis will be rejected. Common values are $$5%$$ and $$1%$$.
* The distribution of the test statistic under the null hypothesis partitions the possible values of T into those for which the null hypothesis is rejected—the so-called critical region—and those for which it is not. The probability of the critical region is $$\alpha$$. In the case of a composite null hypothesis, the maximal probability of the critical region is $$\alpha$$.
* Compute from the observations the observed value $$t\_{obs}$$ of the test statistic $$T$$.
* Decide to either reject the null hypothesis in favor of the alternative or not reject it. The decision rule is to reject the null hypothesis $$H\_0$$ if the observed value $$t\_{obs}$$ is in the critical region, and not to reject the null hypothesis otherwise.

A common alternative formulation of this process goes as follows:

* Compute from the observations the observed value $$t\_{obs}$$of the test statistic $$T$$.
* Calculate the $$p$$-value. This is the probability, under the null hypothesis, of sampling a test statistic at least as extreme as that which was observed (the maximal probability of that event, if the hypothesis is composite).
* Reject the null hypothesis, in favor of the alternative hypothesis, if and only if the $$p$$-value is less than (or equal to) the significance level (the selected probability) threshold ($$\alpha$$), for example $$0.05$$ or $$0.01$$.

The former process was advantageous in the past when only tables of test statistics at common probability thresholds were available. It allowed a decision to be made without the calculation of a probability. It was adequate for classwork and for operational use, but it was deficient for reporting results. The latter process relied on extensive tables or on computational support not always available. The explicit calculation of a probability is useful for reporting. The calculations are now trivially performed with appropriate software.

The difference in the two processes applied to the Radioactive suitcase example (below):

* "The Geiger-counter reading is $$10$$. The limit is $$9$$. Check the suitcase."
* "The Geiger-counter reading is high; $$97%$$ of safe suitcases have lower readings. The limit is $$95%$$. Check the suitcase."

The former report is adequate, the latter gives a more detailed explanation of the data and the reason why the suitcase is being checked.

### Example

A manufacturer claims that the average life of its products is $$36$$ months. An auditor selects a sample of $$49$$ units of the product and calculates the average life to be $$34.5$$ months. The population standard deviation is $$4$$ months. Test the manufacturer's claim at $$3%$$ significance level.

So as per the above problem:

* $$H\_0 : \mu = 36$$ months
* $$H\_1 : \mu ≠ 36$$ months

$$\alpha = 3%$$ or $$0.03$$ $$\sigma = 4$$ $$N = 49$$

First, we will take a look into the Critical-value method:

In this case we will have critical region at both sides with total area of $$0.03$$.

So, the area to the right = $$0.015$$, which means that area till UCV (cumulative probability till that point) $$= 1-0.015=0.985$$

$$z$$-score of cumulative probability of UCV ($$Z\_c$$ in this case) $$= z$$-score of $$0.985 = 2.17$$

Calculating the critical values UCV/LCV $$= \mu \pm Z\_c \* \frac{\sigma}{\sqrt{N}}$$

UCV/LCV $$= 36 \pm 2.17 \* \frac{4}{\sqrt{49}} = 37.24 \text{ and } 34.76$$

Now as the sample mean $$34.5$$ is not between UCV and LCV hence we reject the null hypothesis.

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2Fgit-blob-3c0a5ab2cdb83f711eed8c68e88e247385ded12c%2Fimage10.png?alt=media" alt=""><figcaption></figcaption></figure>

Now let's solve it using the $$p$$-value method:

Calculate the $$z$$-score of $$34.5 = \frac{\bar{x} -\mu}{\frac{\sigma}{\sqrt{N}}} = -2.62$$

Calculate the p-value from [table](https://goodcalculators.com/p-value-calculator/) $$= 0.0044$$, which is cumulative area till sample point

As this it is in the left-hand side hence there is no need to subtract from $$1$$.

Now this will be a 2-tailed test as we are checking for inequality, we need to multiply by $$2 = 0.0088$$, *remember 1-tailed test provides more power to detect an effect because the entire weight is allocated to one direction only.*

As $$p$$-value is $$< \alpha$$ so, we reject null-hypothesis.

### Errors in Hypothesis Testing

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2Fgit-blob-03498ec2f92426e39b1b7d4425253931e263f18b%2Fimage11.png?alt=media" alt=""><figcaption><p>Types of Error in Hypothesis Testing (<a href="https://www.sixsigmadaily.com/type-i-and-type-ii-errors-in-hypothesis-testing/">SOURCE</a>)</p></figcaption></figure>

### Types of Test

1. **Z-Test**:
   * **Purpose**: Used to test hypotheses about a population mean when the population standard deviation is known.
   * **Example**: Suppose you want to test if the average weight of a sample of 50 apples is significantly different from 150 grams (population mean). You know the population standard deviation is 10 grams.
2. **T-Test**:
   * **Purpose**: Used to test hypotheses about a population mean when the population standard deviation is unknown or when dealing with small sample sizes.
   * **Example**: You want to test if a new drug has a statistically significant effect on blood pressure. You collect data from 30 patients before and after treatment and perform a t-test to compare the means.
3. **Chi-Square Test**:
   * **Purpose**: Used to test the independence of categorical variables or goodness-of-fit of observed data to an expected distribution.
   * **Example**: You want to determine if there is an association between smoking habits (smoker, non-smoker) and the incidence of lung cancer (yes, no) in a population. You create a contingency table and perform a chi-square test for independence.
4. **ANOVA (Analysis of Variance)**:
   * **Purpose**: Used to compare means of more than two groups to determine if there are statistically significant differences among them.
   * **Example**: You have data on test scores from three different teaching methods (A, B, C). You want to determine if there is a significant difference in mean test scores between the methods.
5. **Paired T-Test**:
   * **Purpose**: Used to compare the means of two related groups (e.g., before and after treatment) to determine if there is a significant difference.
   * **Example**: You measure the blood pressure of the same group of patients before and after a 6-week exercise program to see if there is a significant change.
6. **Wilcoxon Rank-Sum Test (Mann-Whitney U Test)**:
   * **Purpose**: Used to compare two independent groups when the data is not normally distributed or when ordinal data is involved.
   * **Example**: You want to determine if there is a significant difference in test scores between students who received tutoring and those who did not. The data is not normally distributed.
7. **Fisher's Exact Test**:
   * **Purpose**: Used to test the independence of two categorical variables in a 2x2 contingency table, especially when sample sizes are small.
   * **Example**: You want to determine if there's an association between gender (male, female) and the success of a medical treatment (success, failure) in a small sample of patients.
8. **K-Sample Anderson-Darling Test**:
   * **Purpose**: Used to compare the distribution of multiple independent samples to determine if they come from the same population.
   * **Example**: You have three different groups of people, and you want to test if their ages are drawn from the same population distribution.

These are just a few examples of hypothesis tests, and there are many more specific tests designed for different types of data and research questions. The choice of which test to use depends on the nature of your data, the research question, and the assumptions of the test.

## Questions

<details>

<summary>Limitations of p-value</summary>

Can you tell some limitations of p-value?

**Answer**

Some limitations of p-value are as follows:

* p-value does not give the probability of how true the null hypothesis was. It just gives a binary decision on if it can be rejected or not
* p-value does not consider how precise the effect is (as it assumes we know the sample size, does not tell much about sample size)

</details>

<details>

<summary>Considerations for t-test</summary>

You are testing hundreds of hypotheses with a t-test, what considerations should be made?

**Answer**

Type 1 error will scale the more the number of t-tests are run. If $$\alpha = 0.05$$ then there is $$5%$$ chance of Type 1 error on a single test, then across many tests $$p(Type I)$$ will increase. For example, with 2 tests:

$$P(\text{type I error}) = p(\text{type I error on A OR type I error on B})$$ $$= 2p(\text{type I error on single test}) - p(\text{type I error on A AND type I error on B})$$ $$= 2\*.05 - .05^2 (\text{assuming independence of tests}) = 0.5 - .025 = .075$$

If you want your $$p(type I error)$$ across n-tests to remain at $$5%$$, you will need to decrease the $$\alpha$$ in each individual test. Bonferroni correction can be applied. Basically, alpha will be reduced to alpha/n. n is number of experiments you are running.

Otherwise, you can try and run an F-test to start in order to identify if a least $$1$$ test sees some significant effect. Then run a t-test on the specific experiment with the highest effect size. Granted, the p-value of the test will also depend on the variance of the sample in the given test, if we assume constant variance across tests, then the test with the highest effect size is in expectation the best performing test. Only running a single t-test will keep your p(type I error) low.

</details>

<details>

<summary>Explain selection bias</summary>

📖[Source](https://towardsdatascience.com/40-statistics-interview-problems-and-answers-for-data-scientists-6971a02b7eee)

Explain selection bias (with regard to a dataset, not variable selection). Why is it important? How can data management procedures such as missing data handling make it worse?

**Answer**

Selection bias is the phenomenon of selecting individuals, groups or data for analysis in such a way that proper randomization is not achieved, ultimately resulting in a sample that is not representative of the population. Understanding and identifying selection bias is important because it can significantly skew results and provide false insights about a particular population group. Types of selection bias include:

* sampling bias: a biased sample caused by non-random sampling
* time interval: selecting a specific time frame that supports the desired conclusion. e.g., conducting a sales analysis near Christmas.
* exposure: includes clinical susceptibility bias, protopathic bias, indication bias. Read more here.
* data: includes cherry-picking, suppressing evidence, and the fallacy of incomplete evidence.
* attrition: attrition bias is similar to survivorship bias, where only those that ‘survived’ a long process are included in an analysis, or failure bias, where those that ‘failed’ are only included
* observer selection: related to the Anthropic principle, which is a philosophical consideration that any data we collect about the universe is filtered by the fact that, in order for it to be observable, it must be compatible with the conscious and sapient life that observes it.

Handling missing data can make selection bias worse because different methods impact the data in different ways. For example, if you replace null values with the mean of the data, you are adding bias in the sense that you’re assuming that the data is not as spread out as it might actually be.

</details>


# A/B test

{% @mermaid/diagram content="flowchart TD
A\["`Business Problem`"] --> B\["`Business Sponsor`"]
A --> C\["`Engineering`"]
A --> D\["`Data Science`"]
D --> E\["`Understand the Business Problem`"]
E --> F\["`**North Star Metric (NSM)/Overall Evaluation Criterion (OEC)**
This is the primary KPI that measures the performance of the company towards its goal. Every function of the company - marketing, engineering, etc. works on improving this KPI.
*ex. NSM of Meta would be to increase the total amount of time spent by users in their ecosystem.*`"]
E --> G\["`**Driver Metric** 
NSM on a whole can be difficult to track in the shorter timeframe of the experiment we use driver metric instead. This is a product or feature level measure which is used as a proxy for the NSM in the AB tests.`"]
E --> H\["`**Guardrails**
These are a set of metrics which measures the trade-offs in a business and ensures that the results are not skewed due to any bias`"]
E --> I\["`**Secondary Metrics**
It is a set of metrics which measures the micro-interactions of a feature.`"]

subgraph State-the-Experiment
direction TB
K\["Null (H0) and Alternate (H1) Hypothesis"]
K --> L\["Significance of the test"]
L --> M\["Statistical Power of the test"]
M --> N\["MDE of the test"]
end

subgraph Design-the-Experiment
P\["Sample Size Calculation"]
P --> Q\["Experiment Duration"]
end
I --> State-the-Experiment
G --> State-the-Experiment
H --> State-the-Experiment
F --> State-the-Experiment
State-the-Experiment --> Design-the-Experiment
Design-the-Experiment --> R\["`**Run the Experiment**`"]

subgraph Access-Validity-Threats
Sa\["Stable Unit Treatment
Value Assumption"]
Sb\["Survivorship
Bias"]
Sc\["Sample Ratio
Mismatch"]

Sd\["Primacy Effect
& Novelty Effect"]
Se\["Holiday
Effect"]
Sf\["Other
Effects"]
Sg\["AA
test"]
end
R --> Access-Validity-Threats
Access-Validity-Threats --> T\["`**Statistical Inference**`"]
T --> U\["`**Launch Decision**`"]

" %}

### Metrics

We should also understand a few key terms in A/B testing before diving in:

1. **North Star Metric (NSM):** It is also known as Overall Evaluation Criterion (OEC). This is the primary KPI that measures the performance of the company towards its goal. For example, the NSM of Meta would be to increase the total amount of time spent by users in their ecosystem. Every function of the company — marketing, engineering, etc. works on improving this KPI.
2. **Driver Metric:** Since NSM can be difficult to track in the shorter timeframe of the experiment we use driver metric instead. This is a product or feature level measure which is used as a proxy for the NSM in the AB tests.
3. **Guardrails:** These are a set of metrics which measures the trade-offs in a business and ensures that the results are not skewed due to any bias
4. **Secondary Metrics:** It is a set of metrics which measures the micro-interactions of a feature.

Let’s look at a few examples of the definitions that we just discussed.

<figure><img src="https://cdn-images-1.medium.com/max/1200/1*zqEgwqJVfYbo05LJub-Scw.png" alt=""><figcaption></figcaption></figure>

### **Effects**

**1. Network Effect:**

**Effect Explanation:** Network effect, also known as social influence or social network effect, occurs when the behavior or decisions of one user in a group influence the behavior of other users. In the context of A/B testing, it can affect user interactions and outcomes in ways that are not solely dependent on the changes being tested.

**Example:** Suppose you are testing a new recommendation algorithm on an e-commerce website. Users who see the new recommendations may be influenced by what other users are buying or viewing, impacting their behavior. This can create a network effect that affects the test results.

**Mitigation Strategy:** To mitigate the network effect in A/B testing, you can use the following strategies:

* **Use Random Assignment:** Ensure that users are randomly assigned to control and treatment groups to minimize the influence of social networks on group composition.
* **Sequential Testing:** Consider sequential testing methods like Bayesian Bandit algorithms that continuously adapt based on the ongoing results. These methods can adapt to changes in user behavior influenced by the network effect.

**2. Weekend Effect:**

**Effect Explanation:** The weekend effect refers to the phenomenon where user behavior or website performance significantly differs on weekends compared to weekdays. This effect can skew A/B test results if not properly accounted for.

**Example:** If you are testing changes to a financial news website, you might observe that user engagement is significantly higher during weekdays when the stock market is open compared to weekends. This can impact the test results, making it seem like the changes had a more significant effect than they actually did.

**Mitigation Strategy:** To mitigate the weekend effect in A/B testing, consider the following strategies:

* **Stratified Sampling:** Stratify your sample to ensure that both control and treatment groups have a similar distribution of weekends and weekdays.
* **Time-of-Day Segmentation:** Analyze the data separately for weekends and weekdays or during different time intervals to understand how user behavior varies.
* **Extend the Test Duration:** If possible, run the A/B test for a longer duration to capture both weekday and weekend patterns.

**3. Novelty Effect:**

**Effect Explanation:** The novelty effect occurs when users initially respond positively to a change simply because it is new, but their behavior may revert to the baseline over time. In A/B testing, this effect can lead to inaccurate conclusions about the long-term impact of changes.

**Example:** Imagine you redesign the user interface of a mobile app, and users in the treatment group initially engage more because they find the new design exciting. However, this enthusiasm may wear off after some time.

**Mitigation Strategy:** To mitigate the novelty effect in A/B testing, you can employ these strategies:

* **Monitor Over Time:** Analyze user behavior over an extended period to determine if the effect is sustained or diminishes over time.
* **Segmentation:** Segment users based on their interaction history to identify whether the effect is more pronounced in certain user groups.
* **Retest Over Time:** Consider conducting follow-up A/B tests to validate the long-term impact of changes.

**4. Seasonality Effect:**

**Effect Explanation:** Seasonality effect occurs when user behavior or performance metrics vary predictably due to external factors like holidays, weather, or cultural events. Failing to account for seasonality can lead to misleading A/B test results.

**Example:** If you run an A/B test for a travel booking website during the holiday season, user behavior may be significantly different compared to non-holiday periods. This can impact the interpretation of test results.

**Mitigation Strategy:** To mitigate the seasonality effect in A/B testing, consider these strategies:

* **Use Historical Data:** Analyze historical data to identify and account for seasonal patterns.
* **Control for Seasonal Factors:** Use statistical methods like time series analysis to control for seasonal effects in the data.
* **Extend Test Duration:** Run the A/B test over a longer period that covers multiple seasons to balance out seasonal variations.

In A/B testing, it's crucial to be aware of these effects and implement appropriate mitigation strategies to ensure accurate and actionable results. Additionally, robust statistical analysis and a clear understanding of user behavior are essential for drawing meaningful conclusions from A/B tests.

### Case Study

([Source](https://towardsdatascience.com/cracking-a-b-testing-data-science-interviews-bc66e399b109))

#### Question <a href="#id-3a35" id="id-3a35"></a>

<mark style="color:red;">INTERVIEWER —</mark> Doordash is expanding into other categories such as convenience store delivery. Their notifications have had good success in the past and they are considering sending an in app notification to promote this newly launched category.

How would you design and analyze an experiment to decide if they should roll out the notification?

#### Solution <a href="#c7ff" id="c7ff"></a>

**Part 1 — Ask clarifying questions to understand business goals and product feature details well**

*What Interviewer is looking for -*

* *Did you begin by stating the product/business goal before diving into the experiment details? Talking about the experiment without knowing the product goal is a red flag.*

<mark style="color:green;">INTERVIEWEE —</mark> Before we begin with the experiment details, I would like to make sure my understanding of the background is clear. There could be multiple goals with a feature like this one — <mark style="background-color:yellow;">such as increasing new user acquisition, increase conversion for this category, increasing # of orders in the category or increasing total order value. Can you help me understand what the</mark> <mark style="background-color:yellow;"></mark><mark style="background-color:yellow;">**goal**</mark> <mark style="background-color:yellow;"></mark><mark style="background-color:yellow;">is here?</mark>

<mark style="color:red;">INTERVIEWER —</mark> That’s a fair question. With the in-app notification, we are primarily trying to increase the conversion rate for the new category — i.e. % of users that place an order in the new category out of all users that login to the app.

<mark style="color:green;">INTERVIEWEE —</mark> Ok, that’s helpful. Now I would like to also understand more about the notification — what is the messaging and who is the intended audience?

<mark style="color:red;">INTERVIEWER —</mark> We are not offering any discount at this point. The messaging is simply going to be to let them know we have a new category that they can start ordering from. If the experiment is successful, we intend to roll out the notification to all users.

<mark style="color:green;">INTERVIEWEE —</mark> Ok. Thanks for that background. I am now ready to dive into the experiment details.

**Part 2 — State Business Hypothesis, Null Hypothesis & define metrics to be evaluated**

*What Interviewer is looking for -*

* *That you think through secondary metrics and guardrail metrics in addition to the primary metrics*

<mark style="color:green;">INTERVIEWEE —</mark> So to state the **business hypothesis** -we expect that if we send in-app notification, then the daily number of orders in the new category will increase. That means our <mark style="background-color:yellow;">**Null Hypothesis (Ho)**</mark> <mark style="background-color:yellow;"></mark><mark style="background-color:yellow;">is that there is no change in the conversion rate due to the notification.</mark>

Now let me state the different metrics that we will want to include for the experiment — Since the goal of the notifications is to increase the conversion rate in the new category. That will be our **primary metric**. In terms of **Secondary metrics,** we should also watch the average order value to see what the impact is. It is possible that the conversion rate increases but the average order value decreases such that the resulting impact is lower overall revenue. That is something we may want to watch out for.

We should also consider **guardrail metrics** — these are metrics that are critical to the business that we do not want to impact through the experiment such as time spent on app or app uninstalls for example. Are there any such metrics that we should include in this case?

<mark style="color:red;">INTERVIEWER —</mark> I agree with your choice of primary metric but you can ignore the secondary metrics for this exercise. And you are spot on in terms of guardrail metrics — Doordash wants to be judicious about any features or releases when it comes to their app because we know that the LTV of a customer who has installed the app is much higher. We want to be careful so as not to drive users to uninstall the app.

<mark style="color:green;">INTERVIEWEE —</mark> Ok — that’s good to know. So we will include % of uninstalls as our guardrail metric.

**Part 3 — Choose significance level, power, MDE and calculate the required sample size and duration for the test**

*What Interviewer is looking for -*

* *Your knowledge of the statistical concepts and the calculation for sample size and duration*
* *Whether you consider factors such as network effect (common in two sided marketplaces such as Doordash, Uber, Lyft, Airbnb or social networks such as FB, LinkedIn), day of week effect, seasonality or novelty effect that may affect the validity of the test and need to be considered while arriving at the experiment design*

<mark style="color:green;">INTERVIEWEE —</mark> Now I would like to get into the design of the experiment.

Let’s first see if we need to consider **network effects** — these occur when the behavior of the control is influenced by the treatment given to the test group. Since Doordash is a double sided marketplace, it is more prone to seeing network effects. In this specific case, it is possible that if the treatment given to the test increases the demand from the test group, that may result in a deficit of supply (i.e. dashers) which could in-turn affect the performance of the control group.

To account for network effects, we will need to choose the **randomization unit** differently than we would normally do. There are many ways to do this — we could do **geo-based randomization or time-based randomization or network-cluster randomization or network ego-centric randomization.** Would you like me to go into the details for these?

<mark style="color:red;">INTERVIEWER —</mark> I am glad you brought up network effects as it is in fact something we carefully look for for our experiments in Doordash. In the interest of time, let’s assume there are no network effects in play here and move on.

<mark style="color:green;">INTERVIEWEE —</mark> So if we are assuming there are no network effects to be accounted for, the **randomization unit** for the experiment is simply the user — i.e. we will randomly select users and assign them to treatment and control. Treatment will receive notifications while control will not receive any notifications. Next, I would like to calculate the sample size and duration. For this I need a few inputs.

* **Baseline conversion** — which is the existing conversion of the control before changes are made
* **Minimum detectable difference or MDE** — which is the smallest change in conversion rate we are interested in detecting. A smaller change than this will not be practically significant to the business — it is typically chosen such that the improvement in the desired outcome will justify the cost of implementing and maintaining the feature
* **Statistical Power** — the statistical power can be thought of as the probability of accepting an alternative hypothesis, when the alternative hypothesis is true.
* **Significance Level** — which is the probability of rejecting a null hypothesis when it is true

A 5% significance level and power of 80% are usually chosen and I will assume these unless you say otherwise. Also I will assume a 50–50 split between the control and treatment. Once I have these inputs finalized, I will use power analysis to calculate the sample size.

<mark style="color:red;">INTERVIEWER —</mark> Yes, let’s say based on the analysis we get a sample size of 10,000 users per variation needed. How will you calculate the duration for the test?

<mark style="color:green;">INTERVIEWEE —</mark> Sure, for this we will need the daily number of users that login to the app.

<mark style="color:red;">INTERVIEWER —</mark> Assume we have 10,000 users that login to the app daily.

<mark style="color:green;">INTERVIEWEE —</mark> Ok, in that case, we would need at a minimum 2 days to run the experiment — I arrived at this by taking the total sample size of Control & Treatment and dividing by daily user count. However, there are other factors we should consider when finalizing the duration -

* **Day of week effect —** You may have a different population of users on weekend than weekdays — hence it is important to run long enough to capture weekly cycles.
* **Seasonality —** There can be times when users behave differently that are important to consider, such as holidays.
* **Novelty effect** — When you introduce a new feature, especially one that’s easily noticed, initially it attracts users to try it. A test group may appear to perform well at first, but the effect will quickly decline over time.
* **External effects —** for e.g. let’s say the market is doing really well and more people are likely to ignore the notification with the expectation of making high returns. This will lead us to draw spurious conclusions from the experiment

Due to the above, I would recommend running the experiment for at least one week.

<mark style="color:red;">INTERVIEWER —</mark> Ok, that’s fair. How would you analyze the test results?

**Part 4 — Analyze the results and draw valid conclusions**

*What interviewer is looking for -*

* *Your knowledge of the appropriate statistical tests to be used in different scenarios (for e.g. t-test for sample mean and z-test for sample proportions)*
* *You check for randomization — this will get you some brownie points*
* *You provide a final recommendation (or a framework to get there)*

<mark style="color:green;">INTERVIEWEE —</mark> Sure. There are two key parts to the analysis -

* **Check for randomization** — As best practice, we should check that the randomization was done correctly when assigning test and control. For this, we can look at some baseline metrics that we do not expect to be influenced by the test and compare them for the two groups. We can do this comparison by comparing the histograms or density curves for these metrics between the two groups. If there is no difference, we can conclude that randomization was done correctly.
* **Significance test for all metrics** (including primary and guardrail metrics) — Both our primary metric (conversion rate) and guardrail metric (uninstall rate) are proportions. We can use the z-test to test for statistical significance. We can do this using a programming language such as R or Python.

If there is a statistically significant increase in conversion rate, and uninstall rate is not impacted negatively, I would recommend implementing the test.

If there is a statistically significant increase in conversion rate, and uninstall rate is impacted negatively, I would recommend not implementing the test.

And lastly, if there is no statistically significant increase in conversion rate — I would recommend not implementing the test.

<mark style="color:red;">INTERVIEWER —</mark> That all sounds good. Thanks for your response.


# Overview

This page broadly summarizes the steps needed to go from data gathering to model building

1. Gather the data
2. Import and understand the data: do things like the following to understand more about the data at hand
   1. check the shape of data
   2. check the number of unique values in each column, drop the ones which have the same value
   3. if it is a classification problem check for class imbalance
   4. check for column datatypes and fix if necessary
3. Check and fix for missing values: [Read more about this here](/model-building/data/missing-value)
4. Perform feature engineering to build new columns from the existing ones, it can be YoY growth, etc.
5. Run univariate and multivariate analysis to understand more about the features
6. Detect and treat outliers: [Read more about this here](/model-building/data/outlier)
7. Encode categorical variables: [Read more about this here](/model-building/data/categorical-variable)
8. Standardize the data
9. If needed run different sampling techniques to reduce imbalance or run PCA etc.
10. Split the data to train, test and validation: Training data are collections of examples or samples that are used to 'teach' or 'train the machine learning model. In contrast, validation datasets contain different samples to evaluate trained ML models. It is still possible to tune and control the model at this stage. Working on validation data is used to assess the model performance and fine-tune the parameters of the model. This becomes an iterative process wherein the model learns from the training data and is then validated and fine-tuned on the validation set. Finally, a test data set is a separate sample, an unseen data set, to provide an unbiased final evaluation of a model fit.&#x20;

    **Cross-validation** involves one or more splits of the training data set and validation data set. In particular, K-fold cross-validation aims to maximize accuracy in testing by dividing the source data into several bins or groups. All except one of these are for training and validation purposes. The last is for testing.
11. Determine which metric you want to track
12. Run some basic algorithms to understand which models might give the best results, PYCARET is a good option to quickly prototype
13. Select the promising ones and deep dive on those
14. Check if there are any assumptions of the model --> Perform checks on Collinearity etc.
15. [Tune the model ](/model-building/hyperparameter-optimization)--> Use regularization, Cross validation etc. to reduce overfitting
16. Check for feature importance using built in functions or SHAP


# Data

{% hint style="info" %}
This page is a Work In Progress
{% endhint %}


# Scaling

**Feature scaling is a method used to standardize the range of independent variables or features of data. In data processing, it is also known as data normalization and is generally performed during the data preprocessing step.**

Since the range of values of raw data varies widely, in some machine learning algorithms, objective functions will not work properly without normalization. For example, the majority of classifiers calculate the distance between two points by the Euclidean distance. If one of the features has a broad range of values, the distance will be governed by this particular feature. Therefore, the range of all features should be normalized so that each feature contributes approximately proportionately to the final distance. Another reason why feature scaling is applied is that gradient descent converges much faster with feature scaling than without it.

## Rescaling (min-max normalization)

Also known as min-max scaling or min-max normalization, is the simplest method and consists in rescaling the range of features to scale the range in $$\[0,1]$$ or $$\[-1,1]$$. Selecting the target range depends on the nature of the data. The general formula is given as:

$$
x^{\prime} = \frac{x - min(x)}{max(x) - min(x)}
$$

## Mean normalization

This is similar to above with a slight change: $$x^{\prime} = \frac{x - average(x)}{max(x) - min(x)}$$

## Standardization

In machine learning, we can handle various types of data, e.g., audio signals and pixel values for image data, and this data can include multiple dimensions. Feature standardization makes the values of each feature in the data have zero-mean (when subtracting the mean in the numerator) and unit-variance. This method is widely used for normalization in many machine learning algorithms (e.g., support vector machines, logistic regression, and artificial neural networks).

$$
x^{\prime} = \frac{x - \bar x}{\sigma}
$$

## Scaling to Unit length

Another option that is widely used in machine-learning is to scale the components of a feature vector such that the complete vector has length one. This usually means dividing each component by the Euclidean length of the vector:

$$
x^{\prime} = \frac{x}{\Vert x \Vert}
$$

In some applications (e.g. Histogram features) it can be more practical to use the L1 norm (i.e., Manhattan Distance, City-Block Length or Taxicab Geometry) of the feature vector. This is especially important if in the following learning steps the Scalar Metric is used as a distance measure.


# Missing Value

[📖Source](https://www.analyticsvidhya.com/blog/2016/01/guide-data-exploration/)

Data can have missing values and it is very common. The reason of such missing data can be any of the following:

* **Data Extraction:** It is possible that there are problems with extraction process. In such cases, we should double-check for correct data with data guardians.
* **Data collection:** These errors occur at time of data collection and are harder to correct. They can be categorized in four types:
  * **Missing completely at random:** This is a case when the probability of missing variable is same for all observations
  * **Missing at random:** This is a case when variable is missing at random and missing ratio varies for different values / level of other input variables. For example: We are collecting data for age and female has higher missing value compare to male.
  * **Missing that depends on unobserved predictors:** This is a case when the missing values are not random and are related to the unobserved input variable. For example: In a medical study, if a particular diagnostic causes discomfort, then there is higher chance of drop out from the study. This missing value is not at random unless we have included “discomfort” as an input variable for all patients.
  * **Missing that depends on the missing value itself:** This is a case when the probability of missing value is directly correlated with missing value itself. For example: People with higher or lower income are likely to provide non-response to their earning.

There are many standard practices of dealing with missing data-

* **Deletion:** It is of two types: *List Wise Deletion* and *Pair Wise Deletion*. In list wise deletion, we delete observations where any of the variable is missing. Simplicity is one of the major advantage of this method, but this method reduces the power of model because it reduces the sample size. In pair wise deletion, we perform analysis with all cases in which the variables of interest are present. Advantage of this method is, it keeps as many cases available for analysis. One of the disadvantage of this method, it uses different sample size for different variables.

**Deletion methods are used when the nature of missing data is “Missing completely at random” else non random missing values can bias the model output.**

* **Mean/ Mode/ Median Imputation:** Imputation is a method to fill in the missing values with estimated ones. The objective is to employ known relationships that can be identified in the valid values of the data set to assist in estimating the missing values. Mean/Mode/Median imputation is one of the most frequently used methods. It consists of replacing the missing data for a given attribute by the mean or median (quantitative attribute) or mode (qualitative attribute) of all known values of that variable. It can be of two types:-
  * **Generalized Imputation:** In this case, we calculate the mean or median for all non missing values of that variable then replace missing value with mean or median.
  * **Similar case Imputation:** In this case, for example, we calculate average for gender “Male” (29.75) and “Female” (25) individually of non missing values then replace the missing value based on gender. For “Male“, we will replace missing values of manpower with 29.75 and for “Female” with 25.
* **Prediction Model:** Prediction model is one of the sophisticated method for handling missing data. Here, we create a predictive model to estimate values that will substitute the missing data. In this case, we divide our data set into two sets: One set with no missing values for the variable and another one with missing values. First data set become training data set of the model while second data set with missing values is test data set and variable with missing values is treated as target variable. Next, we create a model to predict target variable based on other attributes of the training data set and populate missing values of test data set.We can use regression, ANOVA, Logistic regression and various modeling technique to perform this. There are 2 drawbacks for this approach:
  * The model estimated values are usually more well-behaved than the true values
  * If there are no relationships with attributes in the data set and the attribute with missing values, then the model will not be precise for estimating missing values
* **KNN Imputation:** In this method of imputation, the missing values of an attribute are imputed using the given number of attributes that are most similar to the attribute whose values are missing. The similarity of two attributes is determined using a distance function.

  **Advantages:**

  * k-nearest neighbour can predict both qualitative & quantitative attributes
  * Creation of predictive model for each attribute with missing data is not required
  * Attributes with multiple missing values can be easily treated
  * Correlation structure of the data is taken into consideration

  **Disadvantage:**

  * KNN algorithm is very time-consuming in analyzing large database. It searches through all the dataset looking for the most similar instances.
  * Choice of k-value is very critical. Higher value of k would include attributes which are significantly different from what we need whereas lower value of k implies missing out of significant attributes.


# Outlier

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FsycEYPHfMRmr3ilII2G3%2Fimage1.png?alt=media&amp;token=16a4f955-037c-4a1c-b8ae-9e5db3f225e7" alt=""><figcaption></figcaption></figure>

Outliers are the extreme values that exhibit significant deviation from the other observations in our data set. By looking at the outlier, it initially seems that this data probably does not belong with the rest of the data set as they look different from the rest.

An outlier may occur due to the variability in the data, or due to experimental error/human error. They may indicate an experimental error or heavy skewness in the data (heavy-tailed distribution).

In the cases when you have a small sample size, outliers can significantly mess up all your results. For statistical analysis of data, outliers can impact the normality test results of our data, invalidate the basic assumptions like constant variances for regression testing etc.

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FcwNQSXHqgPUwmvSi7FgN%2Fimage2.png?alt=media&amp;token=e3aa1980-9875-415d-b82c-621baca2b71a" alt=""><figcaption><p>Outliers tend to affect mean more than median or mode</p></figcaption></figure>

## Detecting Outliers

When starting an outlier detection quest you need to answer 2 important questions about your dataset:

* *Which and how many features am I taking into account to detect outliers? (univariate / multivariate)*
* *Can I assume a distribution(s) of values for my selected features? (parametric / non-parametric)*

Here are some of the techniques for detecting outliers:

* **IQR and Boxplots:** Usually data points which lie 1.5 times of IQR above Q3 and below Q1 are considered outliers

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FQLwmlSposfPqq3lTwC8q%2Fimage3.png?alt=media&amp;token=9e31e0d6-8e2e-45fa-a15e-ae3c6a36c96b" alt=""><figcaption></figcaption></figure>

* **Z-Score or Extreme Value Analysis (parametric):** The z-score or standard score of an observation is a metric that indicates how many standard deviations a data point is from the sample’s mean, assuming a gaussian distribution. This makes z-score a parametric method. Very frequently data points are not to described by a gaussian distribution, this problem can be solved by applying transformations to data ie: scaling it. It is a very effective method if you can describe the values in the feature space with a gaussian distribution. However, it is only convenient to use in a low dimensional feature space, in a small to medium sized dataset.

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2Fyyo7FRBPXNd43i3FM5iC%2Fimage4.png?alt=media&amp;token=7c2bd8bb-39c1-4794-a4fc-3fbbe4fc5a99" alt=""><figcaption><p>Any data point whose Z-score falls out 3rd standard deviation is usually considered an outlier</p></figcaption></figure>

* **Dbscan:** Dbscan is a density based clustering algorithm, it is focused on finding neighbors by density (MinPts) on an ‘n-dimensional sphere’ with radius $\epsilon$. A cluster can be defined as the maximal set of ‘density connected points’ in the feature space.Dbscan then defines different classes of points:
  * **Core point:** A is a core point if its neighborhood (defined by $\epsilon$) contains at least the same number or more points than the parameter MinPts.
  * **Border point:** C is a border point that lies in a cluster and its neighborhood does not contain more points than MinPts, but it is still ‘density reachable’ by other points in the cluster.
  * **Outlier:** N is an outlier point that lies in no cluster and it is not ‘density reachable’ nor ‘density connected’ to any other point. Thus this point will have “his own cluster”.

It is an unsupervised model and needs to be re-calibrated each time a new batch of data is analyzed.

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2F26UkyUyZXOe8HHdnbs1d%2Fimage5.png?alt=media&amp;token=41237aad-0083-4fdd-8324-560af34870af" alt=""><figcaption></figcaption></figure>

* **Isolation Forests:** Isolation forests are an effective method for detecting outliers or novelties in data. The basic principle is that outliers are few and far from the rest of the observations. To build a tree (training), the algorithm randomly picks a feature from the feature space and a random split value ranging between the maximums and minimums. This is made for all the observations in the training set. To build the forest a tree ensemble is made averaging all the trees in the forest.

Then for prediction, it compares an observation against that splitting value in a “node”, that node will have two node children on which another random comparisons will be made. The number of “splittings” made by the algorithm for an instance is named: “path length”. As expected, outliers will have shorter path lengths than the rest of the observations.

If not correctly optimized, training time can be very long and computationally expensive.

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FyToAdpJlmC8an022jYGp%2Fimage6.png?alt=media&amp;token=13b4db2b-7016-4e7c-b14c-1f366c26e659" alt=""><figcaption></figcaption></figure>

## Dealing with Outliers

Below are a few common practices to deal with Outliers:

* Drop the outlier records
* Cap your outliers data or even you can try binning them
* Assign a new value: If an outlier seems to be due to a mistake in the data, you try imputing a value. Common imputation methods include using the mean of a variable or utilizing a regression model to predict the missing value
* Try a transformation: A different approach to true outliers could be to try creating a transformation of the data rather than using the data itself. For example, try creating a percentile version of your original field and working with that new field instead.


# Sampling

{% hint style="warning" %}
This page is a Work In Progress
{% endhint %}

[📚 Source](https://www.kaggle.com/code/residentmario/undersampling-and-oversampling-imbalanced-data/notebook) [📚 Source](https://medium.com/analytics-vidhya/undersampling-and-oversampling-an-old-and-a-new-approach-4f984a0e8392) [📚 Source](https://www.analyticsvidhya.com/blog/2020/07/10-techniques-to-deal-with-class-imbalance-in-machine-learning/)

Oftentimes in practical machine learning problems there will be significant differences in the rarity of different classes of data being predicted. For example, when detecting cancer, we can expect to have datasets with large numbers of false outcomes, and a relatively smaller number of true outcomes. Same with fraud analysis, the number of actual fraud cases will be much lower.

The overall performance of any model trained on such data will be constrained by its ability to predict rare points. In problems where these rare points are only equally important or perhaps less important than non-rare points, this constraint may only become significant in the later "tuning" stages of building the model. But in problems where the rare points are important, or even the point of the classifier (as in a cancer example), dealing with their scarcity is a first-order concern for the model builder.

Tangentially, note that the relative importance of performance on rare observations should inform your choice of error metric for the problem you are working on; the more important they are, the more your metric should penalize underperformance on them.

Several different techniques exist in the practice for dealing with imbalanced dataset. The most naive class of techniques is sampling: *changing the data presented to the model by undersampling common classes, oversampling (duplicating) rare classes, or both.*

## Undersampling vs Oversampling

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FC0Ktl35oE9hrtgC75NLo%2Fimage7.png?alt=media&amp;token=e05840bb-5dad-4da4-a20f-562bc754e5af" alt=""><figcaption><p><a href="https://medium.com/analytics-vidhya/undersampling-and-oversampling-an-old-and-a-new-approach-4f984a0e8392">Source</a></p></figcaption></figure>

**Undersampling** means to get all of the classes to the same amount as the minority class or the one with the least amount of rows. To put this in an example: We have a dataset of 100 rows with three independent columns and one dependent feature, otherwise known as the class column. The class column has three labels: 1, 2, and 3. Label 1 has 39 instances, label 2 has 32 instances and label 3 has 29 instances. In order to apply undersampling to the aforementioned dataset, we would have to reduce label 1 and label 2 to the same amount of instances as label 3. Thus, each label would have, in this particular case, 29 instances each.

In **Oversampling** we try to duplicate the other classes' rows to be equal to that of the majority class.

However, there is a caveat, over-sampling can cause overfitting and in case of under-sampling it can result in loss of information.


# Categorical Variable

In the realm of data analysis, categorical variables play a vital role in representing non-numeric data. To utilize these variables effectively it is essential to convert them into numerical form.

In the realm of data analysis and machine learning, categorical variables play a vital role in representing non-numeric data, such as gender, color, or country of origin. To utilize these variables effectively in statistical models and algorithms, it is essential to convert them into numerical form through a process called encoding.

### Common types of Encoding

1. **One-Hot Encoding:** One-hot encoding is one of the most popular methods for handling categorical variables. In this technique, each category is transformed into a binary vector, with a "1" representing the presence of a category and "0" indicating absence. One-hot encoding is simple, intuitive, and particularly useful for nominal categorical variables (categories without an inherent order).

Pros:

* Preserves distinct categories effectively.
* Suitable for nominal variables.
* Intuitive representation.

Cons:

* Increases dimensionality for high-cardinality variables.
* Requires additional memory and computational resources.

2. **Ordinal Encoding:** Ordinal encoding is used for ordinal categorical variables where the categories have a specific order or rank. In this method, the categories are assigned integer values based on their order. While this approach preserves the ordinal relationship between categories, it assumes a linear relationship, which might not always be appropriate for all datasets.

Pros:

* Preserves ordinal relationship between categories.
* Simple to implement.

Cons:

* Assumes a linear relationship that might not hold in all cases.

3. **Label Encoding:** Label encoding is a straightforward technique where each category is assigned a unique integer label. It is commonly used for binary categorical variables or when working with ordinal variables where a linear relationship can be assumed. However, this method may introduce unintended ordinal relationships when used inappropriately.

Pros:

* Simple and quick to apply.
* Suitable for binary and ordinal variables.

Cons:

* Introduces artificial ordinal relationships.
* Not suitable for nominal variables.

4. **Binary Encoding:** Binary encoding is a compromise between one-hot encoding and label encoding. In this method, each category is represented by a binary code, which reduces dimensionality compared to one-hot encoding while still capturing distinct categories.

Pros:

* Reduces dimensionality compared to one-hot encoding.
* Preserves distinct categories effectively.

Cons:

* May introduce artificial relationships if not used carefully.

5\. **Hash Encoder:** Just like one-hot encoding, the Hash encoder represents categorical features using the new dimensions. Here, the user can fix the number of dimensions after transformation using ***n\_component*** argument. Here is what I mean – A feature with 5 categories can be represented using N new features similarly, a feature with 100 categories can also be transformed using N new features. Doesn’t this sound amazing?

By default, the Hashing encoder uses **the md5** hashing algorithm but a user can pass any algorithm of his choice. Since Hashing transforms the data in lesser dimensions, it may lead to loss of information. Another issue faced by hashing encoder is the **collision.** Since here, a large number of features are depicted into lesser dimensions, hence multiple values can be represented by the same hash value, this is known as a collision.

Moreover, hashing encoders have been very successful in some Kaggle competitions. It is great to try if the dataset has high cardinality features.

***

Encoding categorical variables is a crucial step in data preprocessing for effective data analysis and machine learning. One-hot encoding, ordinal encoding, label encoding, and binary encoding each have their pros and cons, and the choice of method depends on the nature of the categorical data and the specific requirements of the analysis or model. Data scientists and machine learning practitioners must carefully consider these factors when selecting the appropriate encoding method to ensure accurate and reliable results in their applications.


# Hyperparameter Optimization

While building a machine learning model, many design choices need to be made as in how to define the model architecture. For example what should be the depth of my Decision Tree, how many trees should be present in my Random Forest, how many layers should be there in my Meural network, so on and so forth.

In most cases the optimal model architecture is not obvious or readily available to us. **This exploration to select the optimal model architecture automatically from a defined range of options is referred to as hyperparameter optimization or tuning.**

## What's a Hyperparameter?

[📚 Source](https://cloud.google.com/ai-platform/training/docs/hyperparameter-tuning-overview)

Hyperparameters contain the data that govern the training process itself.

Your training application handles **three categories of data** as it trains your model:

* Your **input data (also called training data)**, however, the values in your input data never directly become part of your model.
* Your **model's parameters** are the variables that your chosen machine learning technique uses to adjust to your data. For example, a deep neural network (DNN) is composed of processing nodes (neurons), each with an operation performed on data as it travels through the network. When your DNN is trained, each node has a weight value that tells your model how much impact it has on the final prediction. Those weights are an example of your model's parameters. In many ways, your model's parameters are the model—they are what distinguishes your particular model from other models of the same type working on similar data.
* Your **hyperparameters** are the variables that govern the training process itself. For example, part of setting up a deep neural network is deciding how many hidden layers of nodes to use between the input layer and the output layer, and how many nodes each layer should use. These variables are not directly related to the training data. They are configuration variables. Note that parameters change during a training job, while hyperparameters are usually constant during a job.

## Tuning

The process of Hyperparameter Tuning usually involves the following steps:

* Decide on the hyperparameters that is applicable for the model
* Provide a range or set of values for all the hyperparameters
* Run the training process on the set of parameter combinations to find the best hyperparameter which optimizes your model performance

Now this last step can be an extremely time and resource intensive process depending upon the model and range of the hyperparameters provided. There are many ways to perfrom the last step. Let's start with an example, suppose you are trying to optimize a XGBoost model and this is the **parameter space**:

* 'max\_depth': \[3,6,10],
* 'learning\_rate': \[0.01, 0.05, 0.1],
* 'n\_estimators': \[100, 500, 1000]

### Grid Search

Grid search will search all 27 combinations and find out the combination which gives the best possible value for the parameter of your choice. But the caveat being it takes huge amount of time and resources to perform grid search.

### Randomized Search

Same as Grid search, just that instead of searching the entire parameter space you randomly sample a subset of it, say you sample 10 of the 27 possible combinations. Random search might not give the best possible combination but in most cases it gives reasonably good result.

### Halving Grid Search

[📚 Source](https://scikit-learn.org/stable/modules/grid_search.html#successive-halving-user-guide)

Successive halving (SH) is like a tournament among candidate parameter combinations. SH is an iterative selection process where all candidates (the parameter combinations) are evaluated with a small amount of resources at the first iteration. Only some of these candidates are selected for the next iteration, which will be allocated more resources. For parameter tuning, the resource is typically the number of training samples, but it can also be an arbitrary numeric parameter such as `n_estimators` in a random forest.

### Bayesian Optimization

[📚 Additional Material](http://neupy.com/2016/12/17/hyperparameter_optimization_for_neural_networks.html#bayesian-optimization)

Bayesian optimization (BO) is a global optimization method for noisy black-box functions. Applied to hyperparameter optimization, BO builds a probabilistic model of the function mapping from hyperparameter values to the objective evaluated on a validation set. By iteratively evaluating a promising hyperparameter configuration based on the current model, and then updating it, BO aims to gather observations revealing as much information as possible about this function and, in particular, the location of the optimum. It tries to balance exploration (hyperparameters for which the outcome is most uncertain) and exploitation (hyperparameters expected close to the optimum). In practice, BO has been shown to obtain better results in fewer evaluations compared to grid search and random search, due to the ability to reason about the quality of experiments before they are run.

In plain English, BO evaluates hyperparameters that appear more promising from past results, and finds better settings, rather than using random search with fewer iterations. The performance of the past hyperparameter affects future decisions.

### Evolutionary Optimization

Evolutionary hyperparameter optimization follows a process inspired by the biological concept of evolution:

* Create an initial population of random solutions (i.e., randomly generate tuples of hyperparameters, typically 100+)
* Evaluate the hyperparameters tuples and acquire their fitness function (e.g., 10-fold cross-validation accuracy of the machine learning algorithm with those hyperparameters)
* Rank the hyperparameter tuples by their relative fitness
* Replace the worst-performing hyperparameter tuples with new hyperparameter tuples generated through crossover and mutation
* Repeat steps 2-4 until satisfactory algorithm performance is reached or algorithm performance is no longer improving

### Population-based training

Population Based Training (PBT) learns both hyperparameter values and network weights. Multiple learning processes operate independently, using different hyperparameters. As with evolutionary methods, poorly performing models are iteratively replaced with models that adopt modified hyperparameter values and weights based on the better performers. This replacement model warm starting is the primary differentiator between PBT and other evolutionary methods. PBT thus allows the hyperparameters to evolve and eliminates the need for manual hypertuning. The process makes no assumptions regarding model architecture, loss functions or training procedures.

PBT and its variants are adaptive methods: they update hyperparameters during the training of the models. On the contrary, non-adaptive methods have the sub-optimal strategy to assign a constant set of hyperparameters for the whole training.

## Open Source Tools

[📚 Source](hhttps://neptune.ai/blog/best-tools-for-model-tuning-and-hyperparameter-optimization)

### Ray Tune

Ray provides a simple, universal API for building distributed applications. Tune is a Python library for experiment execution and hyperparameter tuning at any scale. Tune is one of the many packages of Ray. Ray Tune is a Python library that speeds up hyperparameter tuning by leveraging cutting-edge optimization algorithm such as Ax/Botorch, HyperOpt, and Bayesian Optimization at scale.

### Optuna

Optuna is designed specially for machine learning. It’s a black-box optimizer, so it needs an objective function. This objective function decides where to sample in upcoming trials, and returns numerical values (the performance of the hyperparameters). It uses different algorithms, such as GridSearch, Random Search, Bayesian and Evolutionary algorithms to find the optimal hyperparameter values.

### HyperOpt

Hyperopt is a Python library for serial and parallel optimization over awkward search spaces, which may include real-valued, discrete, and conditional dimensions. It uses Bayesian optimization algorithms for hyperparameter tuning, to choose the best parameters for a given model. It can optimize a large-scale model with hundreds of hyperparameters. HyperOpt requires 4 essential components for the optimization of hyperparameters: the search space, the loss function, the optimization algorithm, a database for storing the history (score, configuration).

### Scikit-Optimize

Scikit-Optimize is an open-source library for hyperparameter optimization in Python. It was developed by the team behind Scikit-learn. It’s relatively easy to use compared to other hyperparameter optimization libraries. It has sequential model-based optimization libraries known as Bayesian Hyperparameter Optimization (BHO). The advantage of BHO is that they find better model settings than random search in fewer iterations.


# Overview

This page discusses the building blocks of an algorithm.

As said in the very popular [The Hundred Page Machine Learning Book](https://themlbook.com/) each learning algorithm usually consists of three parts:

1. a loss function
2. an [optimization criterion](#optimization-criterion) based on the loss function (a cost function, for example)
3. an [optimization routine](#optimizers) that leverages training data to find a solution to the optimization criterion.&#x20;

These are the building blocks of any learning algorithm. Some algorithms were designed to explicitly optimize a specific criterion (both linear and logistic regressions, SVM). Some others, including decision tree learning and kNN, optimize the criterion implicitly. Decision tree learning and kNN are among the oldest machine learning algorithms and were invented experimentally based on intuition, without a specific global optimization criterion in mind, and (like it often happens in scientific history) the optimization criteria were developed later to explain why those algorithms work.

A common overview of different algorithms is given below:

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FmpAUMDyPqdBSKp4c1zLf%2Fimage26.png?alt=media&amp;token=d7c8c849-e5dd-4e6d-8de2-807c512ed043" alt=""><figcaption><p>Overview of Different Algorithms (<a href="https://learn.microsoft.com/en-us/azure/machine-learning/algorithm-cheat-sheet">Source</a>) (<a href="https://scikit-learn.org/stable/tutorial/machine_learning_map/">Source</a>)</p></figcaption></figure>

### Optimization Criterion

Cost functions, also known as loss functions or objective functions, are used in data science to quantify the error or discrepancy between the predicted values generated by a model and the actual observed values in the dataset. The choice of a cost function depends on the specific type of problem you are trying to solve. Here are some common examples of cost functions used in data science:

1. **Mean Squared Error (MSE)**:
   * **Use Case**: Regression problems.
   * **Formula**: MSE = $$(1/n) Σ(y\_i - ŷ\_i)^2$$
   * **Description**: Measures the average squared difference between the predicted values (ŷ\_i) and the actual values (y\_i). Penalizes larger errors more.
2. **Mean Absolute Error (MAE)**:
   * **Use Case**: Regression problems.
   * **Formula**: MAE = $$(1/n) Σ|y\_i - ŷ\_i|$$
   * **Description**: Measures the average absolute difference between predicted and actual values. Less sensitive to outliers compared to MSE.
3. **Binary Cross-Entropy (Log Loss)**:
   * **Use Case**: Binary classification problems.
   * **Formula**: BCE = $$-Σ(y\_i \* log(ŷ\_i) + (1 - y\_i) \* log(1 - ŷ\_i))$$
   * **Description**: Evaluates the difference between predicted probabilities ($$ŷ\_i$$) and actual binary labels ($$y\_i$$). Commonly used with logistic regression and neural networks.
4. **Categorical Cross-Entropy (Multiclass Log Loss)**:
   * **Use Case**: Multiclass classification problems.
   * **Formula**: Cross-Entropy = $$-Σ(y\_i \* log(ŷ\_i))$$
   * **Description**: Generalization of binary cross-entropy for multiple classes. Measures the dissimilarity between predicted class probabilities and actual class labels.
5. **Hinge Loss (SVM Loss)**:
   * **Use Case**: Support Vector Machines (SVM) for binary classification.
   * **Formula**: Hinge Loss = $$Σmax(0, 1 - y\_i \* ŷ\_i)$$
   * **Description**: Encourages the correct classification of data points while allowing a margin of error for some misclassified points.
6. **Huber Loss**:
   * **Use Case**: Regression problems, robust to outliers.
   * **Formula**: Huber Loss = $$Σ{0.5 \* (y\_i - ŷ\_i)^2 for |y\_i - ŷ\_i| ≤ δ} + Σ{δ \* |y\_i - ŷ\_i| for |y\_i - ŷ\_i| > δ}$$
   * **Description**: Combines the properties of MSE and MAE by using a quadratic loss for small errors and linear loss for large errors. More robust to outliers.
7. **Kullback-Leibler Divergence (KL Divergence)**:
   * **Use Case**: Used in probabilistic models and when comparing probability distributions.
   * **Formula**: KL Divergence = $$Σ(p(x) \* log(p(x) / q(x)))$$
   * **Description**: Measures the difference between two probability distributions, p(x) and q(x). Used in tasks like variational autoencoders and generative adversarial networks (GANs).
8. **Custom Loss Functions**:
   * **Use Case**: Tailored to specific problems.
   * **Description**: In some cases, custom loss functions are designed to address the unique requirements of a particular problem. For example, in recommendation systems, a loss function might be created to optimize recommendations based on user behavior.

The choice of a cost function depends on the nature of your data, the type of problem you're trying to solve, and your modeling goals. Selecting an appropriate cost function is a crucial step in designing and training machine learning models.

In the context of optimization problems, such as those encountered in machine learning and mathematical optimization, the terms "convex" and "non-convex" refer to the shape of the cost or objective function. These terms describe how the function's curvature changes as you move through the parameter space. Let's break down what each of these terms means:

1. **Convex Cost Function**:
   * **Definition**: A cost function is convex if, when you draw a straight line between any two points on the function's graph, the line lies above the graph of the function for all points in between.
   * **Characteristics**:
     * Has a single global minimum (and no other local minima).
     * Gradient descent and similar optimization algorithms can efficiently find the global minimum.
     * Converges reliably to the optimal solution.
2. **Non-Convex Cost Function**:
   * **Definition**: A cost function is non-convex if the straight line connecting two points on the graph may lie below the function at some intermediate points.
   * **Characteristics**:
     * Can have multiple local minima, making it challenging to find the global minimum.
     * Gradient-based optimization methods may get stuck in local minima.
     * Requires careful initialization and possibly more sophisticated optimization techniques, such as random restarts or simulated annealing, to avoid suboptimal solutions.

In machine learning, when training models, you often encounter non-convex cost functions because the relationships between model parameters and the objective function can be complex. Deep learning, for example, involves optimizing highly non-convex functions due to the complexity of neural network architectures.

The convexity or non-convexity of the cost function has significant implications for optimization. Convex problems are generally easier to solve because they have a single global minimum, while non-convex problems can be more challenging and may require more sophisticated optimization techniques to find good solutions.

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2F0uSYSDSmeRl5tClhNV40%2Fimage.png?alt=media&amp;token=7d153015-6f96-44e8-9bc4-c20904c17285" alt=""><figcaption></figcaption></figure>

### Optimizers

Optimizers are algorithms used in data science and machine learning to adjust the parameters of a model during training to minimize the error or loss function. Each optimizer has its own way of updating these parameters, and they come with their own advantages and disadvantages.&#x20;

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2Fso64q7XVeSw4Z9QMmRsZ%2F56201contours_evaluation_optimizers.gif?alt=media&amp;token=1459cd7c-a30d-44e7-bdd9-ae048fc54567" alt=""><figcaption><p>(<a href="https://www.analyticsvidhya.com/blog/2021/10/a-comprehensive-guide-on-deep-learning-optimizers/">Source</a>)</p></figcaption></figure>

Here are some common optimizers explained in simple terms with their pros and cons:

1. **Gradient Descent**:
   * **How it works**: Gradient Descent computes the gradient (slope) of the loss function and takes small steps in the direction that reduces the loss.
   * **Pros**: Simple, widely used, and works well in many cases.
   * **Cons**: Can be slow to converge to the optimal solution, especially in high-dimensional spaces.
2. **Stochastic Gradient Descent (SGD)**:
   * **How it works**: Similar to Gradient Descent but updates parameters using a random subset (mini-batch) of the training data.
   * **Pros**: Faster convergence, works well with large datasets, and helps escape local minima.
   * **Cons**: Noisy updates can lead to oscillations, and it may require tuning of the learning rate.
3. **Mini-Batch Gradient Descent**:

   * **How it works**: A compromise between Gradient Descent and SGD, where updates are made using small batches of data.
   * **Pros**: Faster than Gradient Descent, less noisy than SGD, and suitable for a wide range of problems.
   * **Cons**: Still requires tuning of the learning rate, and convergence depends on the batch size.

   <mark style="color:blue;">In the case of Stochastic Gradient Descent, we update the parameters after every single observation and we know that every time the weights are updated it is known as an iteration. In the case of Mini-batch Gradient Descent, we take a subset of data and update the parameters based on every subset.</mark>
4. **Adam (Adaptive Moment Estimation)**:
   * **How it works**: Combines ideas from RMSprop and Momentum methods by adapting the learning rates for each parameter based on past gradients.
   * **Pros**: Fast convergence, works well with noisy data, and requires less hyperparameter tuning.
   * **Cons**: Might not perform as well on all problem types and could converge to suboptimal solutions in some cases.
5. **RMSprop (Root Mean Square Propagation)**:
   * **How it works**: Adjusts the learning rate for each parameter based on the magnitude of recent gradients.
   * **Pros**: Effective in handling non-stationary or noisy environments, requires less tuning compared to traditional SGD.
   * **Cons**: Not suitable for all problem types and may converge to local minima.
6. **Adagrad (Adaptive Gradient Algorithm)**:
   * **How it works**: Adapts the learning rate for each parameter based on the historical gradient information.
   * **Pros**: Automatically adapts learning rates, which can be beneficial for sparse data.
   * **Cons**: Learning rates can become too small over time, causing slow convergence and potentially overshooting the optimal solution.
7. **Nadam**:
   * **How it works**: A combination of Nesterov Accelerated Gradient (NAG) and Adam optimizers, which combines their advantages.
   * **Pros**: Fast convergence, good for complex models, and less sensitive to the choice of hyperparameters.
   * **Cons**: Computationally more expensive than some other optimizers.

The choice of optimizer depends on the specific problem you are solving, the dataset size, and the architecture of your neural network. It's common to experiment with different optimizers and hyperparameters to find the one that works best for your particular task.

### **Parametric and non-parametric models**

Parametric and non-parametric models are two fundamental approaches to modeling in statistics and machine learning. They differ in how they represent the underlying data distribution and make assumptions about the functional form of that distribution.

**Parametric Models**:

1. **Definition**: Parametric models make specific assumptions about the functional form or shape of the data distribution. These assumptions are typically described by a fixed number of parameters.
2. **Characteristics**:
   * The number of parameters in a parametric model is fixed, regardless of the amount of data available.
   * Examples of parametric models include linear regression, logistic regression, Gaussian Naive Bayes, and parametric probability distributions like the normal distribution.
   * Parametric models are computationally efficient and require relatively little data to estimate their parameters.
   * They are often interpretable because the model structure is well-defined.
   * However, they may not perform well when the assumptions about the data distribution are incorrect.

**Non-Parametric Models**:

1. **Definition**: Non-parametric models make fewer assumptions about the functional form of the data distribution. Instead, they aim to let the data determine the model's complexity.
2. **Characteristics**:
   * The number of parameters in a non-parametric model can grow with the amount of data. As more data becomes available, non-parametric models can become more complex and flexible.
   * Examples of non-parametric models include k-nearest neighbors (KNN), decision trees, random forests, support vector machines (SVMs), and kernel density estimation.
   * Non-parametric models are more flexible and can capture complex relationships in the data without strong assumptions about distribution shape.
   * They may require larger datasets to perform well, and their computational complexity can be higher.
   * Non-parametric models are often less interpretable because their structure is more data-driven.

**Key Differences**:

1. **Assumptions**: Parametric models make specific assumptions about the data distribution, while non-parametric models make fewer assumptions and let the data dictate the model's complexity.
2. **Complexity**: Parametric models have fixed complexity, while non-parametric models can adapt their complexity to the data.
3. **Interpretability**: Parametric models are often more interpretable due to their well-defined structure, whereas non-parametric models can be more challenging to interpret because they are data-driven.
4. **Data Requirements**: Parametric models can work well with relatively small datasets, while non-parametric models often require larger datasets to capture complex patterns.


# Bias/Variance Tradeoff

## Bias

* Error between average model prediction and ground truth
* The bias of the estimated function tells us the capacity of the underlying model to predict the values

$$bias = \mathbb{E}\[f'(x)] - f(x)$$

## Variance

* Average variability in the model prediction for the given dataset
* The variance of the estimated function tells you how much the function can adjust to the change in the dataset

$$variance = \mathbb{E}\[(f'(x) - \mathbb{E}\[f'(x)])^2]$$

**High Bias:**

* Overly simplified Model
* Under-fitting
* High error on both test and train data

**High Variance:**

* Overly complex Model
* Over-fitting
* Low error on train data and high on test
* Starts modelling the noise in the input

***

The bias-variance tradeoff is a fundamental concept in machine learning that helps us understand the tradeoff between two types of errors that a model can make: bias and variance. It's crucial to strike a balance between these two types of errors to create a model that generalizes well to new, unseen data.

**Bias** refers to the error due to overly simplistic assumptions in the learning algorithm. A high bias model might underfit the data, meaning it fails to capture the underlying patterns and relationships in the data, leading to poor performance on both the training and test sets.

**Variance** refers to the error due to the model's sensitivity to small fluctuations in the training data. A high variance model might overfit the data, meaning it fits the training data very well but fails to generalize to new data points, resulting in poor performance on the test set.

**Tradeoff Explanation with Example:**

Let's consider a simple example of fitting a polynomial to a set of data points. Imagine you have a set of data points that form a curve on a 2D plane. Your goal is to find a polynomial equation that fits these data points.

1. **High Bias, Low Variance:** Suppose you decide to fit a linear equation (a straight line) to the data points. This is a high bias model because it makes a simplistic assumption about the underlying relationship. As a result, the fitted line might not capture the curve's nuances, leading to a bias in predictions. However, this simple model is less sensitive to variations in the training data, so it might perform similarly on both the training and test sets.
2. **Low Bias, High Variance:** Now consider fitting a high-degree polynomial (e.g., a 10th-degree polynomial) to the data. This is a low bias model because it has the flexibility to closely follow the data points, even capturing their intricate details. However, this model is more sensitive to the noise and fluctuations in the training data, resulting in a high variance. It's likely to fit the training data extremely well but may not generalize to new data points, leading to poor performance on the test set.

Finding the right balance between bias and variance is crucial. You want a model that is complex enough to capture the underlying patterns but not overly complex to avoid fitting noise. This balance can be achieved by techniques like cross-validation, regularization, and choosing an appropriate model complexity based on the problem's nature and the available data.

In essence, the bias-variance tradeoff reminds us that while we aim to minimize both bias and variance, there's often a tradeoff between the two. The goal is to find the "sweet spot" where the model generalizes well to new data while still capturing the essential patterns in the training data.

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2F6gVALwWTI4Kq6Nfmzdap%2Fimage24.png?alt=media&amp;token=bf9c4e8c-2cab-4e5f-aa3f-1b642cff9e34" alt=""><figcaption><p><a href="https://stanford.edu/~shervine/teaching/cs-229/cheatsheet-machine-learning-tips-and-tricks">📖Source</a></p></figcaption></figure>

## Questions

<details>

<summary>High Bias</summary>

How can you identify a High Bias model? How can you fix it?

**Answer**

A **High Bias** model is due to a simple model and can be easily identified when you see:

1. High training error
2. Validation error or test error is the same as training error

To **fix** a High Bias model, you can:

1. Add more input features
2. Add more complexity by introducing polynomial features
3. Decrease the regularization term

</details>

<details>

<summary>Data Bias</summary>

Can you name a few types of data biases?

**Answer** [(Source)](https://www.telusinternational.com/articles/7-types-of-data-bias-in-machine-learning)

Though not exhaustive, this list contains common examples of data bias in the field, along with examples of where it occurs.

**Sample bias:** Sample bias occurs when a dataset does not reflect the realities of the environment in which a model will run. An example of this is certain facial recognition systems trained primarily on images of white men. These models have considerably lower levels of accuracy with women and people of different ethnicities. Another name for this bias is selection bias.

**Exclusion bias:** Exclusion bias is most common at the data preprocessing stage. Most often it's a case of deleting valuable data thought to be unimportant. However, it can also occur due to the systematic exclusion of certain information. For example, imagine you have a dataset of customer sales in America and Canada. 98% of the customers are from America, so you choose to delete the location data thinking it is irrelevant. However, this means you model will not pick up on the fact that your Canadian customers spend two times more.

**Measurement bias:** This type of bias occurs when the data collected for training differs from that collected in the real world, or when faulty measurements result in data distortion. A good example of this bias occurs in image recognition datasets, where the training data is collected with one type of camera, but the production data is collected with a different camera. Measurement bias can also occur due to inconsistent annotation during the data labeling stage of a project.

**Recall bias:** This is a kind of measurement bias, and is common at the data labeling stage of a project. Recall bias arises when you label similar types of data inconsistently. This results in lower accuracy. For example, let's say you have a team labeling images of phones as damaged, partially-damaged, or undamaged. If someone labels one image as damaged, but a similar image as partially damaged, your data will be inconsistent.

**Observer bias:** Also known as confirmation bias, observer bias is the effect of seeing what you expect to see or want to see in data. This can happen when researchers go into a project with subjective thoughts about their study, either conscious or unconscious. We can also see this when labelers let their subjective thoughts control their labeling habits, resulting in inaccurate data.

**Racial bias:** Though not data bias in the traditional sense, this still warrants mentioning due to its prevalence in AI technology of late. Racial bias occurs when data skews in favor of particular demographics. This can be seen in facial recognition and automatic speech recognition technology which fails to recognize people of color as accurately as it does caucasians.

**Association bias:** This bias occurs when the data for a machine learning model reinforces and/or multiplies a cultural bias. Your dataset may have a collection of jobs in which all men are doctors and all women are nurses. This does not mean that women cannot be doctors, and men cannot be nurses. However, as far as your machine learning model is concerned, female doctors and male nurses do not exist. Association bias is best known for creating gender bias.

</details>

<details>

<summary>KNN vs Bias/Variance</summary>

How can you relate the *KNN Algorithm* to the *Bias-Variance tradeoff*?

**Answer** ([Source](https://teazrq.github.io/stat542/rlab/knn.html))

$$K$$ nearest neighbor is a simple nonparametric method.

Suppose we collect a set of observations $$({x\_i, y\_i}*{i=1}^n)$$, the prediction at a new target point is $$\widehat y = \frac{1}{k} \sum*{x\_i \in N\_k(x\_0)} y\_i$$, where $$(N\_k(x\_0))$$ defines the $$(k)$$ samples from the training data that are closest to $$(x\_0)$$. As default, closeness is defined using distance measures, such as the Euclidean distance.

If we consider different values of $$k$$, we can observe the trade-off between bias and variance.

<img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FnYnCHvZv9YmxWCuzR24V%2FKNN%20vs%20bias%20var.png?alt=media&amp;token=584d79b5-0408-4730-abb6-d63c5ee2294e" alt="" data-size="original">

&#x20;As $$k$$ increases, we have a more stable model, i.e., smaller variance, however, the bias is also increased. As $$k$$ decreases, the bias also decreases, but the model is less stable. Formally, the prediction error (at a given target point $$x\_0$$) can be broken into three parts: the irreducible error, the bias squared, and variance.$$\[\begin{aligned} E\Big\[ \big( Y - \widehat f(x\_0) \big)^2 \Big] &= E \Big\[ \big( Y - f(x\_0) + f(x\_0) -  E\[\widehat f(x\_0)] + E\[\widehat f(x\_0)] - \widehat f(x\_0) \big)^2 \Big] \ &= E \Big\[ \big( Y - f(x\_0) \big)^2 \Big] + E \Big\[ \big(f(x\_0) -  E\[\widehat f(x\_0)] \big)^2 \Big] + E\Big\[ \big(E\[\widehat f(x\_0)] - \widehat f(x\_0) \big)^2 \Big] + \text{Cross Terms}\ &= \underbrace{E\Big\[ ( Y - f(x\_0))^2 \big]}*{\text{Irreducible Error}} + \underbrace{\Big(f(x\_0) - E\[\widehat f(x\_0)]\Big)^2}*{\text{Bias}^2} +  \underbrace{E\Big\[ \big(\widehat f(x\_0) - E\[\widehat f(x\_0)] \big)^2 \Big]}\_{\text{Variance}} \end{aligned}]$$

</details>


# Regression

## Linear Regression

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FnpPYCODPofIEqKuAqrrm%2Fimage8.png?alt=media&amp;token=58db9047-f562-4e9c-8157-860932e164f9" alt=""><figcaption><p><a href="https://xkcd.com/605/">Source</a></p></figcaption></figure>

$$Y = b\_0 + b\_1 \* x\_1 + b\_2 \* x\_2 + \epsilon$$

The idea is to find the line or plane which best fits the data. Collectively, $$b\_0, b\_1, b\_2$$ are called regression coefficients. $$\epsilon$$ is the error term, the part of $$Y$$ the regression model is unable to explain.

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FZpoBOBsysVwpj0oNf99t%2Fimage2.png?alt=media&amp;token=48e4f28b-66e2-4de7-8d6e-8fedfb051960" alt=""><figcaption><p><a href="https://www.shutterstock.com/image-illustration/annotated-diagram-explaining-components-graph-showing-1406041139">Source</a></p></figcaption></figure>

1. **Loss Function:** The loss function in linear regression quantifies how well the model's predictions match the actual target values. In linear regression, the most common loss function is the Mean Squared Error (MSE). The MSE calculates the average squared difference between the predicted values and the actual target values for all data points.
2. **Optimization Criterion:** The optimization criterion is the goal of finding the best-fitting line (or hyperplane in higher dimensions) that minimizes the chosen loss function. In the case of linear regression, the goal is to find the coefficients (slope and intercept) of the linear equation that minimize the MSE. This involves adjusting the coefficients to minimize the overall squared difference between the predicted and actual values.
3. **Optimization Routine:** To find the optimal coefficients that minimize the MSE, an optimization routine is used. Gradient Descent is a widely used optimization algorithm for linear regression.&#x20;

### Metrics

Now once you have the model fit next comes the metrics to measure how good the fit is, some of the common metrics are as follows:

* $$RSS$$ (Residual sum of squares) $$= (Y\_{actual} - Y\_{predicted})^2$$, it changes with scale change
* $$TSS$$ (Total sum of squares) $$= (Y\_{actual} - Y\_{avg})^2$$
* $$R^2$$ $$= 1-\frac{RSS}{TSS}$$, more the better, increases with more coefficients
* $$RSE$$ (Residual Standard Error) $$= \sqrt{\frac{RSS}{d.o.f}}$$, here $$d.o.f = n-2$$

### Feature selection

[📖Explanation](https://towardsdatascience.com/log-book-practical-guide-to-linear-polynomial-regression-in-r-e0ed2e7f8031)

* Hypothesis testing and using p-values to understand if the feature is important or not
* Using metrics like $$\text{Adjusted} R^2$$, $$AIC$$, $$BIC$$, etc. which takes into consideration the number of features used to build the model and penalizes accordingly
* How do we find the model that minimizes a metric like $$AIC$$? One approach is to search through all possible models, called all **subset regression**. This is computationally expensive and is not feasible for problems with large data and many variables. An attractive alternative is to use **stepwise regression** about which we learned above, this successively adds and drops predictors to find a model that lowers $$AIC$$. Simpler yet are **forward selection** and **backward selection**. In forward selection, you start with no predictors and add them one-by-one, at each step adding the predictor that has the largest contribution to , stopping when the contribution is no longer statistically significant. In backward selection, or backward elimination, you start with the full model and take away predictors that are not statistically significant until you are left with a model in which all predictors are statistically significant.
* **Penalized Regression** or **Regularization**:

  [📖Explanation](https://www.analyticsvidhya.com/blog/2016/01/ridge-lasso-regression-python-complete-tutorial/)

  Penalized regression is similar in spirit to AIC. Instead of explicitly searching through a discrete set of models, the model-fitting equation incorporates a constraint that penalizes the model for too many variables (parameters). Rather than eliminating predictor variables entirely — as with stepwise, forward, and backward selection — penalized regression applies the penalty by reducing coefficients, in some cases to near zero. Common penalized regression methods are ridge regression and lasso regression. Regularization is nothing but adding a penalty term to the objective function and control the model complexity using that penalty term. It can be used for many machine learning Algorithms. Both Ridge and Lasso regression uses $$L2$$ and $$L1$$ regularizations.

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FGJ71Wb2zXweKlLlEfk8B%2Fimage23.png?alt=media&amp;token=405c263c-d33b-438f-ace9-b726b7ec3968" alt=""><figcaption><p><a href="https://stanford.edu/~shervine/teaching/cs-229/cheatsheet-machine-learning-tips-and-tricks">Source</a></p></figcaption></figure>

{% hint style="info" %}
Ridge brings the coefficients close to $$0$$ but not exactly to $$0$$ which results in the model retaining all the features. Lasso on the other hand brings the coefficients to $$0$$ hence results in reduced features. Lasso shrinks the coefficients by same amount whereas Ridge shrinks them by same proportion.

Elastic Net is another useful technique which combines both L1 and L2 regularization.
{% endhint %}

### Assumptions

* The relationship between $$X$$ and $$Y$$ is **linear**. Because we are fitting a linear model, we assume that the relationship really is linear, and that the errors, or residuals, are simply random fluctuations around the true line.
* The error terms are **normally distributed**. This can be checked with a Q-Q plot
* Error terms are independent of each other. This can be checked with a ACF plot. This can be used while checking independence while using a time-series data
* Error terms are **homoscedastic**, i.e. they have constant variance. Residulas Vs Fitted graph should be flat. This means that the variability in the response is changing as the predicted value increases. This is a problem, in part, because the observations with larger errors will have more pull or influence on the fitted model.
* The independent variables are not multicollinear. **Multicollinearity** is when a variable can be explained as a combination of other variables. This can be checked by using **VIF(Variance inflation factor)** $$= \frac{1}{1-R\_i^2}$$.
  * A VIF score of $$>10$$ indicates there there is a problem
  * If a multicollinear variable is present the coefficients swing wildly thereby affecting the interpretability of the model. P-vales are not reliable. But it doesnot affect prediction or the goodness of fit statistics.
  * To deal with multicollinearity
    * drop variables
    * create new features from existing ones
    * PCA/PLS

{% hint style="info" %}
One very important point to remember is that Generalized Linear Regression is called so because $$Y$$ is linear w\.r.t its coefficients $$b\_0, b\_1, b\_2$$, etc. it is irrespective of whether the features $$x\_1, x\_2$$, etc. are linear or not. Meaning $$x\_1$$ can actually be $$x\_1^2$$ and it won't matter.
{% endhint %}

### OLS Stats Model (Ordinary Least Square)

OLS is a stats model, which will help us in identifying the more significant features that can has an influence on the output. OLS model in python is executed as: lm = smf.ols(formula = 'Sales \~ am+constant', data = data).fit() lm.conf\_int() lm.summary() And we get the output as below

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FF09UC9s1tw0Xc8Y4srQs%2Fimage.png?alt=media&amp;token=d7586a76-0cfe-4efc-bd5e-9e510c16a87e" alt=""><figcaption><p>The higher the t-value for the feature, the more significant the feature is to the output variable. And also, the p-value plays a rule in rejecting the Null hypothesis(Null hypothesis stating the features has zero significance on the target variable.). If the p-value is less than 0.05(95% confidence interval) for a feature, then we can consider the feature to be significant.</p></figcaption></figure>

## SVR (Support Vector Regression)

In simple linear regression, try to minimize the error rate. But in SVR, we try to fit the error within a certain threshold.

Our best fit line is the one where the hyperplane has the maximum number of points. We are trying to do here is trying to decide a decision boundary at ‘e’ distance from the original hyperplane such that data points closest to the hyperplane or the support vectors are within that boundary line.

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FRommcuSfn8y12T6oXZET%2Fimage.png?alt=media&amp;token=dffc9e54-ea91-420a-a620-046428f6dd9a" alt=""><figcaption></figcaption></figure>

## Non-Linear Regression

In some cases, the true relationship between the outcome and a predictor variable might not be linear. There are different solutions extending the linear regression model for capturing these nonlinear effects, some of these are covered below.

### Polynomial Regression

The equation of polynomial becomes something like this.

$$Y = b\_0 + b\_1 \* x\_1 + b\_2 \* x\_1^2 + b\_n \* x\_1^n$$and so on...

The degree of order which to use is a Hyperparameter, and we need to choose it wisely. But using a high degree of polynomial tries to overfit the data and for smaller values of degree, the model tries to underfit so we need to find the optimum value of a degree. **Polynomial Regression on datasets with high variability chances to result in over-fitting.**

### Regression Splines

[📖Explanation](https://www.analyticsvidhya.com/blog/2018/03/introduction-regression-splines-python-codes/)

In order to overcome the disadvantages of polynomial regression, we can use an improved regression technique which, instead of building one model for the entire dataset, divides the dataset into multiple bins and fits each bin with a separate model. Such a technique is known as Regression spline.

In polynomial regression, we generated new features by using various polynomial functions on the existing features which imposed a global structure on the dataset. To overcome this, we can divide the distribution of the data into separate portions and fit linear or low degree polynomial functions on each of these portions. The points where the division occurs are called **Knots**. Functions which we can use for modelling each piece/bin are known as Piecewise functions. There are various piecewise functions that we can use to fit these individual bins.

### Generalized additive models

It does the same thing as above but just removes the need to specifying the knots. It fits spline models with automated selection of knots.

## Questions

<details>

<summary>[UPSTART] Regression Coefficient</summary>

Suppose we have two variables, $$X$$ and $$Y$$, where $$Y = X +$$ some normal white noise.

1. What will our coefficient be of we run a regression of $$Y$$ on $$X$$?
2. What happens if we run a regression of $$X$$ on $$Y$$?

**Answer**

Let's start with $$Y = X$$, then the regression line is a perfect fit. The points of such a dataset is $$(1,1),(2,2),(3,3),(4,4),(5,5)$$

Adding some normal white noise to these points $$(1,1),(2,3),(3,5),(4,5),(5,5)$$. A regression line fit on these points will move up. Hence the coefficients of $$Y = mX+c$$, $$m$$ will increase, $$c$$ might still stay at $$0$$ or at max increase.

This movement will go in the negative direction if we predict $$X$$ based on $$Y$$

</details>

<details>

<summary>Linear Regression in Time Series</summary>

Do you think Linear Regression should be used in Time series analysis?

**Answer**

Linear Regression as per me can be used in Time Series but might not always give good results. Few reasons which come up are:

* Linear Regression is good for intrapolation but not for extrapolation so the results can vary wildly
* When Linear Regression is used but observations are correlated (as in time series data) you will have a biased estimate of the variance
* Moreover, time-series data have a pattern, such as during peak hours, festive seasons, etc., which would most likely be treated as outliers in the linear regression analysis

</details>

<details>

<summary>[AIRBNB] Booking Regression</summary>

Let's say we want to build a model to predict booking prices.

1. Explain the difference between a linear regression versus a random forest regression.
2. Which one would likely perform better?

**Answer**

Linear Regression is used to predict continuous outputs where there is a linear relationship between the features of the dataset and the output variable. It is used for regression problems where you are trying to predict something with infinite possible answers such as the price of a house.

In the case of regression, decision trees in random forest learn by splitting the training examples in a way such that the sum of squared residuals is minimized. To classify a new object based on attributes, each tree gives a classification and we say the tree “votes” for that class. The forest chooses the classification having the most votes (over all the trees in the forest) and in case of regression, it takes the average of outputs by different trees. It is useful when there are complex relationships between the features and the output variables. They also work well compared to other Algorithms when there are missing features, when there is a mix of categorical and numerical features and when there is a big difference in the scale of features.

It is difficult to tell which will perform better, it completely depends on the problem statement and the available data. Other than the points mentioned above some of the Key advantages of linear models over tree-based ones are:

* they can extrapolate (e.g., if labels are between 1-5 in train set, tree-based model will never predict 10, but linear will)
* could be used for anomaly detection because of extrapolation
* interpretability (yes, tree-based models have feature importance, but it's only a proxy, weights in linear model are better)
* need less data to get good results
* Random Forest is able to discover more complex relation at the cost of time

The first point becomes clearly important in this case as we would need booking price values which might not necessarily be in the training data range.

</details>

<details>

<summary>[GOOGLE] Adding Noise</summary>

Say we are running a probabilistic linear regression which does a good job modeling the underlying relationship between some $$y$$ and $$x$$. Now assume all inputs have some noise $$\epsilon$$ added, which is independent of the training data.

What is the new objective function? How do you compute it?&#x20;

**Answer** ([Source)](https://www.nicksingh.com/posts/30-machine-learning-interview-questions-ml-interview-study-guide)

The objective function for linear regression where $$x$$ is set of input vectors and $$w$$ are the weights: $$L(w) = E\[(w^Tx-y)^2]$$

Let's assume that the noise added is Gaussian as follows: $$\epsilon \sim N(0, \lambda I)$$, then the new objective function is given by:$$L(w) = E\[(w^T(x + \epsilon)-y)^2]$$.

To compute it, we simplify: $$L'(w) = E\[(w^T x -y + w^T\epsilon)^2]$$ $$L'(w) = E\[(w^T x - y)^2 + 2(w^Tx-y)w^T\epsilon +w^T\epsilon \epsilon^Tw]$$ $$L'(w) = E\[(w^T x - y)^2] + E\[2(w^Tx-y)w^T\epsilon] + E\[w^T\epsilon \epsilon^Tw]$$

We know that the expectation for $$\epsilon$$ is $$0$$ so, the middle term becomes $$0$$ and we are left with:$$L'(w) = L(w) + 0 + w^TE\[\epsilon \epsilon^T]w$$

The last term can be simplified as: $$L'(w) = L(w) + w^T\lambda Iw$$

And therefore, the objective function simplifies to that of L2-regularization: $$L'(w) = L(w) + \lambda||w||^2$$

</details>

<details>

<summary>[UBER] L1 vs L2</summary>

What is L1 and L2 regularization? What are the differences between the two?

**Answer** [(Source)](https://www.nicksingh.com/posts/30-machine-learning-interview-questions-ml-interview-study-guide)

$$L1$$ and $$L2$$ regularization are both methods of regularization that attempt to prevent overfitting in machine learning. For a regular regression model assume the loss function is given by $$L$$. $$L1$$ adds the absolute value of the coefficients as a penalty term, whereas $$L2$$ adds the squared magnitude of the coefficients as a penalty term.

The loss function for the two are:

* $$Loss(L\_1) = L + \lambda |w\_i|$$
* $$Loss(L\_2) = L + \lambda |w\_i^2|$$

Where the loss function $$L$$ is the sum of errors squared, given by the following, where $$f(x)$$ is the model of interest, for example, linear regression with $$p$$ predictors:

$$L = \sum\_{i=1}^{n} (y\_i - f(x\_i))^2 = \sum\_{i=1}^{n} (y\_i - \sum\_{j=1}^{p}(x\_{ij}w\_j) )^2 \space \text{for linear regression}$$

If we run gradient descent on the weights $$w$$, we find that $$L1$$ regularization will force any weight closer to $$0$$, irrespective of its magnitude, whereas, for the $$L2$$ regularization, the rate at which the weight goes towards $$0$$ becomes slower as the rate goes towards $$0$$. Because of this, $$L1$$ is more likely to “zero” out particular weights, and hence removing certain features from the model completely, leading to more sparse models.

</details>

<details>

<summary>[TESLA] Choice of Cost Function</summary>

You're working with several sensors that are designed to predict a particular energy consumption metric on a vehicle. Using the outputs of the sensors, you build a linear regression model to make the prediction. There are many sensors, and several of the sensors are prone to complete failure.

What are some cost functions you might consider, and which would you decide to minimize in this scenario?

**Answer** [(Source)](https://www.nicksingh.com/posts/30-machine-learning-interview-questions-ml-interview-study-guide)

There are two potential cost functions here, one using the $$L1$$ norm and the other using the $$L2$$ norm. Below are two basic cost functions using an L1 and L2 norm respectively:

* $$J(w) = ||Xw-y||$$
* $$J(w) = |Xw-y|^2$$

It would be more sensible to use an $$L1$$ norm in this case since the $$L1$$ norm penalizes the outliers harder and thus gives less weight to the complete failures than the $$L2$$ norm does.

Additionally, it would be prudent to involve a regularization term to account for noise. If we assume that the noise added to each sensor uniformly as follows: $$\epsilon \sim N(0, \lambda I)$$ then using traditional $$L2$$ regularization, we would have the cost function: $$J(w) = ||Xw-y|| + \lambda||w||^2$$

However, given the fact that there are many sensors (and a broad range of how useful they are), we could instead assume that noise is added by: $$\epsilon \sim N(0, \lambda D)$$ where each diagonal term in the matrix D represents the error term used for each sensor (and hence penalizing certain sensors more than others). Then our final cost function is given by: $$J(w) = ||Xw-y|| + \lambda w^TDw$$

</details>

<details>

<summary>[AIRBNB] Prove that maximizing the likelihood is equivalent to minimizing the sum of squared residuals</summary>

Suppose you are running a linear regression and model the error terms as being normally distributed. Show that in this setup, maximizing the likelihood of the data is equivalent to minimizing the sum of squared residuals.

**Answer** [(Source)](https://cppcodingzen.com/?p=1609)

A mathematical derivation like this requires us to:

* Define correct Mathematical symbols and their relationships through equations
* Recall and use the definitions of the terms like **likelihood** and **normally distributed**
* Perform Mathematical manipulation to derive the required result

**Problem Setup:**

A linear regression model proposes that the output $$y$$ is linearly dependent on the input vector $$X$$ by the relation,

$$y = W^T X + \beta$$

Where, $$X$$ is an $$m$$-dimensional vector, and $$(W; \beta) = {w\_1, w\_2, ..., w\_m; \beta}$$ are the parameters of the model.

Next, we are give a set of training data points, consisting of

* a set of input vectors, $$X = {X\_1, X\_2, ..., X\_n}$$. Note that every input $$X\_i$$ is a vector, $$X\_i = {X\_{i1}, X\_{i2}, ..., X\_{im}}$$
* and a set of outputs $$Y = {y\_1, y\_2, ..., y\_n}$$

Given the values of the parameters $$W$$ and $$\beta$$, the estimate $$\hat{y}$$ is given by

$$\hat{y}*i = \sum*{j=1}^m X\_{ij} \* w\_j + \beta = X\_{i1} \* w\_1 + X\_{i2} \* w\_2 + ... + X\_{im} \* w\_m + \beta$$.

Finally, the error term for $$i^{th}$$ input is simply the difference between the observed value $$y$$ and the estimate $$\hat{y}$$.

$$\epsilon\_i(W, \beta) = y\_i - \hat{y}*i = y\_i - \sum*{j=1}^m X\_{ij} \* w\_j - \beta$$

Note that the error term depends on the parameters of the model $$W$$ and $$\beta$$, and hence is denoted as $$\epsilon\_i(W, \beta)$$.

**Likelihood:**

Take a look at the problem statement again. We are assuming that the error terms are normally distributed. There is an implicit assumption that all the error terms are independent of each other. (Make sure you make this assumption explicit to your interviewer).

What does it mean for the error term to be normally distributed. It means that, by definition, the probability distribution function of the $$i^{th}$$ error term is given by

$$l(\epsilon\_i|W; \beta) = \frac{1}{\sqrt{2\pi}\sigma}exp{\frac{-\epsilon\_i^2}{2\sigma^2}}$$

This probability distribution function of the $$i^{th}$$ input is also called its likelihood function, and also depends on the parameters of the model, $$W$$ and $$\beta$$.

Since we are assuming that the error terms are also independent, their joint probability distribution, is given by the product of their likelihood.

$$l = \displaystyle \prod\_{i=1}^n l(\epsilon\_i|W; \beta)$$ $$= \displaystyle \prod\_{i=1}^n \frac{1}{\sqrt{2\pi}\sigma}exp{\frac{-\epsilon\_i^2}{2\sigma^2}}$$ $$= \displaystyle \frac{1}{(\sqrt{2\pi}\sigma)^n} \prod\_{i=1}^n exp{\frac{-\epsilon\_i^2}{2\sigma^2}}$$ $$= \displaystyle \frac{1}{(\sqrt{2\pi}\sigma)^n} \prod\_{i=1}^n exp{\frac{-(y\_i - \sum\_{j=1}^m X\_{ij} \* w\_j - \beta)^2}{2\sigma^2}}$$

**Maximum Likelihood Estimator:**

The maximum likelihood estimator seeks to maximize the likelihood function defined above. For the maximization,

* We can ignore the constant $$\frac{1}{(\sqrt{2\pi}\sigma)^n}$$
* We can also take the log of the likelihood function, converting the product into sum

The log likelihood function of the errors is given by

$$L = log(l)$$ $$= \displaystyle log(\prod\_{i=1}^n exp{\frac{-(y\_i - \sum\_{j=1}^m X\_{ij} \* w\_j - \beta)^2}{2\sigma^2}})$$ $$= \displaystyle \sum\_{i=1}^n {\frac{-(y\_i - \sum\_{j=1}^m X\_{ij} \* w\_j - \beta)^2}{2\sigma^2}}$$

As a final step, for the purpose of optimization, we can ignore the constant multiplier $$2\sigma^2$$ from the summation, giving us

$$L = \displaystyle \sum\_{i=1}^n -(y\_i - \sum\_{j=1}^m X\_{ij} \* w\_j - \beta)^2$$

But this is just the **negative of the sum of squared errors!**

Thus, if you want to maximize the likelihood (or log likelihood) of the errors, you better minimize the sum of squared errors of the estimates.

</details>


# Generative vs Discriminative Models

Machine learning models can be classified into two types of models – Discriminative and Generative models. In simple words, a discriminative model makes predictions on the unseen data based on conditional probability and can be used either for classification or regression problem statements. On the contrary, a generative model focuses on the distribution of a dataset to return a probability for a given example.

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2Fgit-blob-d9f0095fb4fe5cff9000b46853c700a9c35b8e9d%2Fimage7.png?alt=media" alt=""><figcaption><p>This essentially summarizes Generative vs Discriminative Models (<a href="https://dataisutopia.com/blog/discremenet-generative-models/">Source)</a></p></figcaption></figure>

## Discriminative Models

Discriminative models separate classes instead of modeling the conditional probability and don’t make any assumptions about the data points. But these models are not capable of generating new data points. Therefore, the ultimate objective of discriminative models is to separate one class from another.

In case of outliers present in the dataset, then discriminative models work better compared to generative models i.e, discriminative models are more robust to outliers. However, there is one major drawback of these models is the misclassification problem, i.e., wrongly classifying a data point.

Some examples are:

* Scalar Vector Machine(SVMs)
* Conditional Random Fields (CRFs)
* Decision Trees and Random Forest
* Logistic regression
* Traditional Neural Networks
* Nearest Neighbor

## Generative Models

Generative models focus on the distribution of individual classes in a dataset and the learning algorithms tend to model the underlying patterns or distribution of the data points. These models use the concept of joint probability and create the instances where a given feature ($$x$$) or input and the desired output or label ($$y$$) exist at the same time.

These models use probability estimates and likelihood to model data points and differentiate between different class labels present in a dataset. Unlike discriminative models, these models are also capable of generating new data points.

However, they also have a major drawback – If there is a presence of outliers in the dataset, then it affects these types of models to a significant extent.

Some examples are:

* Naïve Bayes
* Bayesian networks
* Markov random fields- Hidden Markov Models (HMMs)
* Latent Dirichlet Allocation (LDA)
* Generative Adversarial Networks (GANs)
* Autoregressive Model

## Questions

<details>

<summary>Generative vs Discriminative Model</summary>

Can you describe the distinction between Generative and Discriminative Models from the probability standpoint?

**Answer**

In mathematical terms, a discriminative machine learning trains a model which is done by learning parameters that maximize the conditional probability $$P(Y|X)$$, while on the other hand, a generative model learns parameters by maximizing the joint probability of $$P(X,Y)$$.

</details>

<details>

<summary>Generative and Discriminative Model Description</summary>

Describe both generative and discriminative models and give an example of each?

**Answer**

In simple words, a discriminative model makes predictions on the unseen data based on conditional probability and can be used either for classification or regression problem statements. On the contrary, a generative model focuses on the distribution of a dataset to return a probability for a given example.

Linear Discriminant Analysis (LDA) is a generative model, whereas Logistic Regression is a discriminative model.

</details>


# Classification

## Logistic Regression

It is easy to say that the linear regression predicts a “value” of the targeted variable through a linear combination of the given features, while on the other hand, a Logistic regression predicts “probability value” through a linear combination of the given features plugged inside a logistic function. Linear regression is unbounded, and this brings logistic regression into picture. Their value strictly ranges from 0 to 1.

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FUaaKrw9rF5WeqqqdWCNz%2Fimage3.png?alt=media&amp;token=b3e1cbc6-a7c8-4eda-af33-71ccb1b526be" alt=""><figcaption><p>Logistic regression uses Sigmoid function to transform linear regression into the logit function. Logit is nothing but log of Odds. Then using log of Odds it calculates the required probability. (<a href="https://www.vebuso.com/2020/02/linear-to-logistic-regression-explained-step-by-step/">Source</a>)</p></figcaption></figure>

### Cost Function

One more thing to note here is that logistic regression uses maximum likelihood estimation (MLE) instead of least squares method of minimizing the error which is used in linear models. In Linear regression we minimized SSE. In Logistic Regression we maximize log likelihood instead. Linear regression uses mean squared error as its cost function. If this is used for logistic regression, then it will be a non-convex function of parameters (theta). Gradient descent will converge into global minimum only if the function is convex.

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2F5HJih9QlopnEPfjSttki%2Fimage4.png?alt=media&amp;token=afa93874-06d1-4269-aff1-b0503fdacf23" alt=""><figcaption><p>Cost function (<a href="https://pvgisours.tistory.com/59">Source</a>)</p></figcaption></figure>

## Metrics

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FoOWXAeUjpOcs4KXz4UnX%2Fimage5.png?alt=media&amp;token=ea5b15f0-58ac-4349-a404-78d9b243c5d6" alt=""><figcaption><p>Confusion Matrix and key Metrics</p></figcaption></figure>

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2Fm2HHzVj54bPY1320upna%2Fimage6.png?alt=media&amp;token=061be3c2-49be-4b59-84d7-c09f0afeb3a3" alt=""><figcaption><p>(a) ROC curve (b) Precision-Recall curve. Both are a helpful diagnostic tool for evaluating a single classifier but challenging for comparing classifiers. Like ROC AUC, we can calculate the area under the curve as a score and use that score to compare classifiers. The focus on the minority class makes the Precision-Recall AUC more useful for imbalanced classification problems. (<a href="https://machinelearningmastery.com/tour-of-evaluation-metrics-for-imbalanced-classification/">Source</a>)</p></figcaption></figure>

In classification problems, various evaluation metrics are used to assess the performance of a machine learning model. The choice of metric depends on the specific characteristics of your problem and your priorities, such as the relative importance of false positives and false negatives. Here's an explanation of common classification metrics and when to use them, along with examples:

1. **Accuracy**:
   * **Use Case**: Suitable for balanced datasets where false positives and false negatives have similar consequences.
   * **Example**: In a spam email classifier, where both false positives (legitimate emails marked as spam) and false negatives (spam emails in the inbox) are undesirable.
2. **Precision**:
   * **Use Case**: When minimizing false positives is crucial, and you want to ensure that the positive predictions made by your model are highly accurate.
   * **Example**: Medical diagnoses like cancer detection, where false positives can lead to unnecessary treatments and stress.
3. **Recall (Sensitivity or True Positive Rate)**:
   * **Use Case**: When minimizing false negatives is critical, and you want to ensure that your model captures as many positive instances as possible.
   * **Example**: An airport security system for detecting prohibited items, where missing a threat (false negative) is far more serious than a false alarm.
4. **F1 Score**:
   * **Use Case**: Balances precision and recall, suitable when you want a single metric that considers both false positives and false negatives.
   * **Example**: Information retrieval systems, where you need to find relevant documents (recall) while minimizing the number of irrelevant ones (precision).
5. **Specificity (True Negative Rate)**:
   * **Use Case**: Relevant in scenarios where minimizing false positives is essential, like fraud detection.
   * **Example**: Credit card fraud detection, where it's important to correctly identify non-fraudulent transactions (true negatives) to prevent blocking legitimate transactions.
6. **ROC AUC (Receiver Operating Characteristic Area Under the Curve)**:
   * **Use Case**: Useful when comparing different models or assessing the overall performance of a classifier across different thresholds.
   * **Example**: Evaluating the performance of various machine learning algorithms in a credit scoring task.
7. **Matthews Correlation Coefficient (MCC)**:
   * **Use Case**: Appropriate for imbalanced datasets, where there is a significant difference in class frequencies.
   * **Example**: Anomaly detection in network security, where normal events far outnumber anomalous ones.
8. **F-beta Score**:
   * **Use Case**: Allows you to adjust the balance between precision and recall using the parameter beta.
   * **Example**: When you want to prioritize either precision (beta < 1) or recall (beta > 1) depending on the specific needs of your application.

Remember that the choice of metric should be based on the specific goals and trade-offs of your problem. It's often a good practice to consider multiple metrics, especially when the consequences of false positives and false negatives differ significantly in your application.

It seems like there might be a typo in your question. I assume you are referring to the **Naive Bayes algorithm**. Naive Bayes is a classification algorithm, not "naive bias." Let me provide a detailed explanation of the Naive Bayes algorithm.

### **Naive Bayes Algorithm**

Naive Bayes is a probabilistic machine learning algorithm used for classification tasks, such as spam email detection, sentiment analysis, and text categorization. It is based on Bayes' theorem, which calculates the probability of an event based on prior knowledge of conditions related to that event.

The Naive Bayes algorithm makes a simplifying assumption known as the "naive" assumption, which is that all features used in the classification are conditionally independent of each other given the class label. This means that the presence or absence of one feature does not affect the presence or absence of another feature.

**3. Model Representation:** In Naive Bayes, the goal is to calculate the probability of a particular class (C) given a set of features (X₁, X₂, ..., Xᵢ). This is represented as:

```
P(C | X₁, X₂, ..., Xᵢ) = P(C) * P(X₁ | C) * P(X₂ | C) * ... * P(Xᵢ | C)
```

Where:

* P(C | X₁, X₂, ..., Xᵢ) is the posterior probability of class C given the features X₁ through Xᵢ.
* P(C) is the prior probability of class C.
* P(Xᵢ | C) is the conditional probability of feature Xᵢ given class C.

&#x20;**Training:** To train a Naive Bayes classifier, you need labeled training data where you know both the features and the corresponding class labels. The training process involves:

a. Calculating Prior Probabilities (P(C)):

* Calculate the prior probability of each class, i.e., the probability that an example belongs to that class based on the training data.

b. Estimating Conditional Probabilities (P(Xᵢ | C)):

* For each feature Xᵢ and each class C, estimate the conditional probability that the feature Xᵢ occurs given the class C. This is typically done using techniques like Maximum Likelihood Estimation (MLE) or Laplace smoothing (to handle zero probabilities).

&#x20;**Types of Naive Bayes:** There are different variants of Naive Bayes classifiers, including:

* **Gaussian Naive Bayes**: Assumes that continuous features follow a Gaussian distribution.
* **Multinomial Naive Bayes**: Used for discrete data like text data, where features represent word counts or frequencies.
* **Bernoulli Naive Bayes**: Suitable for binary data, where features are binary variables.

**Advantages:**

* Naive Bayes is simple, computationally efficient, and scales well to high-dimensional data.
* It works well with small to moderate-sized datasets.
* It is particularly effective for text classification tasks like spam detection and sentiment analysis.

&#x20;**Limitations:**

* The "naive" assumption of feature independence may not hold in some real-world scenarios.
* It can perform poorly when features are highly correlated.
* Handling of continuous and numerical data may require additional preprocessing.

Despite its simplifying assumptions, Naive Bayes is a powerful and often surprisingly effective algorithm, especially for text classification tasks and situations where feature independence is a reasonable approximation.

## Questions

<details>

<summary>[ROBINHOOD] Interpret Coefficients</summary>

How would you interpret coefficients of logistic regression for categorical and boolean variables?

**Answer**

**Reference:** [Explanation](https://www.displayr.com/how-to-interpret-logistic-regression-coefficients/)

Let's explain this using an example. The table below shows the main outputs from the logistic regression. It is very obvious which are the categorial variables out here: ![](https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2F75rGQqO7eutsnERlIVMZ%2Fimage1.png?alt=media\&token=16186cb8-0a26-4a99-a604-548c7c697eb4)

The first category (usually not shown) has a coefficient of $$0$$. So, if we can say, for example, that:

* The effect of having a DSL service versus having no DSL service $$(0.92 - 0 = 0.92)$$ is a little more than twice as big in terms of leading to churn as is the effect of being a senior citizen $$(0.41)$$.
* The effect of having a Fiber optic service is approximately twice as big as having a DSL service.
* If somebody has a One-year contract and a DSL service, these two effects almost completely cancel each other out.

Consider the scenario of a senior citizen with a $$2$$ month tenure, with no internet service, a one-year contract and a monthly charge of $100. If we compute all the effects and add them up we have:

$$0.41$$ (Senior Citizen = Yes) $$- 0.06 (2\*-0.03$$; tenure) $$+ 0$$ (no internet service) $$- 0.88$$ (one year contract) $$+ 0 (100\*0$$; monthly charge) $$= -0.53$$.

We then need to add the (Intercept), also sometimes called the constant, which gives us $$-0.53- 1.41 = -1.94$$. To make the next bit a little more transparent, I am going to substitute $$-1.94$$ with $$x$$. The logistic transformation is:

Probability $$= \frac{1} {1 + \exp^{-x}} = \frac{1}{1 + \exp^{1.94}} = 0.13 = 13%$$.

Thus, the senior citizen with a $$2$$ month tenure, no internet service, a one-year contract, and a monthly charge of $$$100$$, is predicted as having a $$13%$$ chance of cancelling their subscription. By contrast if we redo this, just changing one thing, which is substituting the effect for no internet service $$(0)$$ with that for a fiber optic connection $$(1.86)$$, we compute that they have a $$48%$$ chance of cancelling.

</details>

<details>

<summary>Multinomial Logistic Regression</summary>

Can Logistic Regression be used for multi class classification?

**Answer**

Logistic regression, by default, is limited to two-class classification problems. Some extensions like one-vs-rest can allow logistic regression to be used for multi-class classification problems, although they require that the classification problem first be transformed into multiple binary classification problems.

Multinomial logistic regression algorithm is an extension to the logistic regression model that involves changing the loss function to cross-entropy loss and predict probability distribution to a multinomial probability distribution to natively support multi-class classification problems.

</details>

<details>

<summary>Choice of Cost Function</summary>

In what situations would you recommend using one metric over the another for classification models?

**Answer**

It all depends on the use case. For example, a diagnostic lab will be concerned with incorrect positive diagnosis. Hence, they will aim for a high specificity value. On the other hand, for a model predicting loan default rate the goal is to identify even a small chance of default, hence we need the model to maximize sensitivity.

</details>

<details>

<summary>[INTUIT] Logistic Regression vs Decision Tree</summary>

What's the difference between decision tree and logistic regression?

**Answer**

Decision trees and logistic regression are both machine learning algorithms used for classification tasks, but they have different approaches and characteristics. Here are the key differences between decision trees and logistic regression:

**1. Algorithm Type:**

* **Decision Tree**: Decision trees are non-linear models that use a tree-like structure to make decisions by recursively splitting the data into subsets based on the most informative features.
* **Logistic Regression**: Logistic regression is a linear model that estimates the probability of a binary outcome by fitting a linear equation to the input features.

**2. Model Complexity:**

* **Decision Tree**: Decision trees can capture complex relationships in the data and can fit highly non-linear decision boundaries.
* **Logistic Regression**: Logistic regression assumes a linear relationship between the input features and the log-odds of the output, making it less flexible for modeling complex, non-linear relationships.

**3. Interpretability:**

* **Decision Tree**: Decision trees are highly interpretable. You can easily visualize the tree structure and understand how decisions are made at each node.
* **Logistic Regression**: Logistic regression provides interpretable coefficients for each feature, indicating the direction and magnitude of their influence on the outcome.

**4. Handling of Numeric vs. Categorical Features:**

* **Decision Tree**: Decision trees can handle both numeric and categorical features without requiring one-hot encoding.
* **Logistic Regression**: Logistic regression typically requires one-hot encoding of categorical features to be included in the model.

**5. Overfitting:**

* **Decision Tree**: Decision trees are prone to overfitting, especially if they are deep and complex. Pruning or limiting the depth of the tree can help mitigate overfitting.
* **Logistic Regression**: Logistic regression is less prone to overfitting, especially when the number of features is limited relative to the number of training samples.

**6. Probability Output:**

* **Decision Tree**: Decision trees can provide class probabilities by counting the proportion of samples in each leaf node belonging to a particular class. However, this can lead to uneven class probability estimates.
* **Logistic Regression**: Logistic regression provides well-calibrated class probabilities, making it suitable for tasks where probability estimates are essential.

**7. Handling Imbalanced Data:**

* **Decision Tree**: Decision trees can struggle with imbalanced datasets, as they tend to favor the majority class in splits.
* **Logistic Regression**: Logistic regression can handle imbalanced datasets better by adjusting the decision threshold or using class weights.

**8. Performance on Linear Problems:**

* **Decision Tree**: Decision trees are not well-suited for linear problems where the decision boundary is best represented by a straight line.
* **Logistic Regression**: Logistic regression is appropriate for linear problems and can capture linear relationships effectively.

In practice, the choice between decision trees and logistic regression depends on the specific characteristics of your data and the problem you are trying to solve. Decision trees are more suitable for non-linear and interpretable problems, while logistic regression is a good choice for problems where linear relationships are predominant and well-calibrated probability estimates are required.

</details>


# Clustering

Clustering analysis is a method used in unsupervised machine learning to group similar data points together in a dataset. Its primary objective is to identify inherent patterns or structures within the data without any prior labels. Here's a summary of the different types of clustering along with their pros and cons:

1. **K-means Clustering**:
   * **Description**: K-means clustering partitions the data into a predetermined number of clusters, where each cluster is represented by its centroid. It iteratively assigns data points to the nearest centroid and updates the centroids until convergence.
   * **Pros**:
     * Simple and computationally efficient.
     * Scales well to large datasets.
   * **Cons**:
     * Sensitive to initial centroid selection.
     * Assumes spherical clusters and struggles with non-linear boundaries.
2. **Hierarchical Clustering**:
   * **Description**: Hierarchical clustering creates a tree-like hierarchy of clusters by recursively merging or splitting clusters based on their similarity.
   * **Pros**:
     * No need to specify the number of clusters beforehand.
     * Provides insights into the relationships between clusters through dendrograms.
   * **Cons**:
     * Computationally expensive for large datasets.
     * Less suitable for high-dimensional data.
3. **Density-based Clustering (DBSCAN)**:
   * **Description**: DBSCAN groups together densely packed data points as clusters based on a specified minimum number of points within a given distance threshold.
   * **Pros**:
     * Can discover clusters of arbitrary shapes.
     * Robust to outliers and noise.
   * **Cons**:
     * Sensitive to parameter settings.
     * May struggle with datasets of varying densities.
4. **Gaussian Mixture Models (GMM)**:
   * **Description**: GMM assumes that the data is generated from a mixture of several Gaussian distributions. It models each cluster as a Gaussian distribution and estimates the parameters (mean and covariance) using expectation-maximization (EM) algorithm.
   * **Pros**:
     * Flexible in capturing complex cluster shapes.
     * Provides probabilistic cluster assignments.
   * **Cons**:
     * Sensitive to initialization and local optima.
     * Requires a sufficient amount of data to estimate parameters accurately.

Each type of clustering algorithm has its own strengths and weaknesses, making them suitable for different types of datasets and applications. Choosing the appropriate clustering method depends on factors such as the nature of the data, the desired number of clusters, and computational resources available.

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FLQkot74TVSiM3gPPiSWu%2Fimage.png?alt=media&amp;token=2e4896a9-4ee4-4593-9268-f7c9279d7da3" alt=""><figcaption><p>(<a href="https://www.linkedin.com/newsletters/the-ai-vanguard-7043488558778626048/">Source</a>)</p></figcaption></figure>

***

## Questions

<details>

<summary>Significance Testing</summary>

Why do you need to perform significance testing for Clustering?

**Answer**

* **Significance testing** addresses an important aspect of cluster validation. Many cluster analysis methods will deliver clusterings even for homogeneous data. They assume implicitly that clustering has to be found, regardless of whether this is meaningful or not.

A critical and challenging question in cluster analysis is whether the identified clusters represent important underlying structure or are artifacts of natural sampling variation.

* **Significance testing** is performed to distinguish between a clustering that reflects meaningful *heterogeneity* in the data and an artificial clustering of *homogeneous* data.
* Significance testing is also used for more specific tasks in cluster analysis, such as; estimating the number of clusters, and for interpreting some or all of the individual clusters, to show the significance of the individual clusters.

</details>

<details>

<summary>Jaccard index</summary>

Can you explain Jaccard index?

**Answer**

Jaccard Index, also known as the Jaccard similarity coefficient, is a measure used to understand the similarity between two sets of data.

Imagine you have two baskets of fruits. One basket has apples, bananas, and cherries, while the other has bananas, cherries, and dates. The Jaccard Index helps us determine how similar these two baskets are.

Here’s how it works:

1. **Intersection**: First, we look at the common items in both baskets. In our case, it’s bananas and cherries.
2. **Union**: Next, we consider all unique items across both baskets. Here, it’s apples, bananas, cherries, and dates.
3. **Jaccard Index**: We then divide the number of common items (intersection) by the total number of unique items (union). So, our Jaccard Index would be 2 (bananas and cherries) divided by 4 (apples, bananas, cherries, dates), which equals 0.5.

So, the Jaccard Index for our two fruit baskets is 0.5, indicating they are 50% similar. If the baskets were identical, the Jaccard Index would be 1 (or 100%), and if they had no common items, the index would be 0.

This concept is widely used in data science and machine learning, especially in clustering and recommendation systems, to measure the similarity between different sets of data.

</details>

<details>

<summary>MeanShift vs <strong>Similarity-Based</strong> Clustering</summary>

Can you explain the difference between MeanShift  vs **Similarity-Based** clustering?

**Answer**

**MeanShift Clustering** and **Similarity-Based Clustering** are two different approaches to clustering data. Here’s a comparison of the two:

**MeanShift Clustering**:

* MeanShift is a non-parametric, density-based clustering algorithm.
* It does not require the number of clusters to be specified in advance.
* The algorithm works by finding regions of high density and iteratively shifting data points towards the highest density of points.
* MeanShift aims to discover “blobs” in a smooth density of samples. It is a centroid-based algorithm, which works by updating candidates for centroids to be the mean of the points within a given region.

**Similarity-Based Clustering**:

* Similarity-Based Clustering, also known as distance-based clustering, groups data points based on their similarity or distance from each other.
* The similarity or distance can be calculated using various metrics such as Euclidean distance, Manhattan distance, cosine similarity, etc.
* K-means is a popular example of similarity-based clustering where data points are grouped based on their distance from the centroid of the clusters.
* The number of clusters needs to be specified in advance in similarity-based clustering methods like K-means.

In summary, the choice between MeanShift and Similarity-Based Clustering depends on the specific requirements of your task, such as whether the number of clusters is known in advance and the nature of the data you are working with.

</details>

<details>

<summary>Gaussian Mixture Model (GMM)</summary>

Explain Gaussian Mixture Model (GMM) in easy to understand terms.

**Answer**

Sure, let’s think of a Gaussian Mixture Model (GMM) as a recipe for making a fruit smoothie.

Imagine you have different types of fruits like apples, bananas, and strawberries. Each type of fruit represents a Gaussian (or normal) distribution. The taste of each fruit is unique, just like each Gaussian distribution in a GMM has its own mean (average) and variance (spread).

Now, when you make a smoothie, you don’t use the same amount of each fruit. You might use more strawberries and fewer apples. This is similar to how a GMM assigns different weights to each Gaussian distribution, indicating their importance in the overall model.

When you blend all the fruits together, you get a smoothie that has a combined flavor of all the fruits. Similarly, a GMM is a mixture of multiple Gaussian distributions to form a more complex distribution.

One of the key features of GMMs is their ability to assign a “membership score” to each data point for each cluster (like asking how much a sip of the smoothie tastes like apple or strawberry). This is different from some other clustering algorithms that simply assign each data point to one cluster.

Remember, just like adjusting the amount of each fruit changes the taste of your smoothie, changing the parameters (mean, variance, and weight) of a GMM changes the shape of the resulting distribution.

</details>

<details>

<summary>Types of hierarchical clustering</summary>

Explain the types of hierarchical clustering and the difference between them?

**Answer**

Hierarchical clustering is a method of cluster analysis that seeks to build a hierarchy of clusters. There are two main types of hierarchical clustering: Agglomerative Clustering and Divisive Clustering.

**Agglomerative Clustering**:

* This is a “bottom-up” approach.
* Each observation starts in its own cluster, and pairs of clusters are merged as one moves up the hierarchy.
* The algorithm merges the closest (or most similar) clusters together, based on a specified measure of distance or similarity (like Euclidean distance).
* This process continues until all data points are merged into a single cluster.

**Divisive Clustering:**

* This is a “top-down” approach.
* All observations start in one cluster, and splits are performed recursively as one moves down the hierarchy.
* The algorithm starts with one large cluster and divides it into smaller ones based on the data’s structure.
* This process continues until each data point forms its own individual cluster.

The choice between agglomerative and divisive clustering depends on the specific requirements of your task, such as the size of your dataset and the nature of the data you are working with.

</details>

<details>

<summary>Latent Class Model</summary>

What is latent class model?

**Answer**

A Latent Class Model (LCM) is a statistical model used for clustering multivariate discrete data. Here’s an easy way to understand it:

Imagine you’re a teacher with a classroom full of students. You want to group these students based on their performance in different subjects like Math, English, and Science. However, you don’t know their exact capabilities in these subjects, which are “latent” or hidden.

In an LCM, each student represents a data point, and their capabilities in different subjects are the latent classes. The model assumes that the students’ grades arise from a mixture of these latent classes. For example, a student doing well in Math and Science but not in English might belong to a latent class of “STEM-oriented” students.

The LCM will try to detect these latent classes based on the observed data (the students’ grades). The goal is to find a model where, within each latent class, the observed variables (grades in this case) are statistically independent. This means that knowing a student’s grade in Math doesn’t give any information about their grade in English, given that we know the latent class to which they belong.

This model is used in various fields like market research, healthcare, and social sciences to uncover hidden groupings in data.

</details>

<details>

<summary>Outlier Analysis</summary>

Can you use clustering for outlier analysis?

**Answer**

Clustering can be a powerful tool for outlier analysis. The idea is that normal data objects will belong to clusters, while outliers will either belong to small or sparse clusters, or not belong to any clusters at all.

Here are a few ways clustering is used for outlier detection:

1. **DBSCAN Clustering**: DBSCAN (Density-Based Spatial Clustering of Applications with Noise) is a density-based clustering algorithm that can handle outliers. It works by identifying groups such that members of each group are densely packed together and then identifies outliers as data points that fall outside of any densely packed cluster.
2. **Cluster and Outlier Analysis (Anselin Local Moran’s I)**: This method identifies statistically significant hot spots, cold spots, and spatial outliers using the Anselin Local Moran’s I statistic. A high positive z-score for a feature indicates that the surrounding features have similar values (either high values or low values). A low negative z-score (for example, less than -3.96) for a feature indicates a statistically significant spatial data outlier.

Remember, the choice of clustering algorithm and the parameters used can greatly affect the results of outlier detection. It’s always a good idea to understand the nature of your data and the assumptions of the algorithm before applying it for outlier analysis.

</details>


# Tree based approaches

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FeofxJ4nkT2Db37qygz7f%2Fimage13.png?alt=media&amp;token=12560759-1715-4ab1-8753-bdaa6c5f4da3" alt=""><figcaption></figcaption></figure>

## Decision Tree

A decision tree is a flowchart-like structure in which each internal node represents a "test" on an attribute (e.g., whether a coin flip comes up heads or tails), each branch represents the outcome of the test, and each leaf node represents a class label (decision taken after computing all attributes). It is immensly popular primarily due to its ease of explanation which is often a critical requirement in business.

Decision Trees follow Sum of Product (SOP) representation. The Sum of product (SOP) is also known as Disjunctive Normal Form. For a class, every branch from the root of the tree to a leaf node having the same class is conjunction (product) of values, different branches ending in that class form a disjunction (sum).

The primary challenge in the decision tree implementation is to identify which attributes do we need to consider as the root node and each level. Handling this is to know as the attribute selection. We have different attributes selection measures to identify the attribute which can be considered as the root note at each level.

The decision of making strategic splits heavily affects a tree’s accuracy. The decision criteria are different for classification and regression trees.

Decision trees use multiple algorithms to decide to split a node into two or more sub-nodes. The creation of sub-nodes increases the homogeneity of resultant sub-nodes. In other words, we can say that the purity of the node increases with respect to the target variable. The decision tree splits the nodes on all available variables and then selects the split which results in most homogeneous sub-nodes.

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2F6r3bAAp0aUbMBT70jBgM%2Fimage11.png?alt=media&amp;token=e737d957-d155-48f9-8fe5-320961a5d517" alt=""><figcaption><p>Decision Tree</p></figcaption></figure>

The algorithm selection is also based on the type of target variables. Let us look at some algorithms used in Decision Trees:

* ID3 → (extension of D3)
* C4.5 → (successor of ID3)
* CART → (Classification And Regression Tree)
* CHAID → (Chi-square automatic interaction detection Performs multi-level splits when computing classification trees)
* MARS → (multivariate adaptive regression splines)

The ID3 algorithm builds decision trees using a top-down greedy search approach through the space of possible branches with no backtracking. A greedy algorithm, as the name suggests, always makes the choice that seems to be the best at that moment.

### Steps in the ID3 algorithm:

* It begins with the original set as the root node.
* On each iteration of the algorithm, it iterates through the unused attributes of the set and calculates a measure of Homogeneity:
  * **Gini Index:** Gini Index uses the probability of finding a data point with one label as an indicator for homogeneity — if the dataset is completely homogeneous, then the probability of finding a datapoint with one of the labels is 1 and the probability of finding a data point with the other label is zero
  * **Information Gain / Entropy-based:** The idea is to use the notion of entropy which is a central concept in information theory. Entropy quantifies the degree of disorder in the data. Entropy is always a positive number between zero and 1. Another interpretation of entropy is in terms of information content. A completely homogeneous dataset has no information content in it (there is nothing non-trivial to be learnt from the dataset) whereas a dataset with a lot of disorder has a lot of latent information waiting to be learnt.
  * For Regression models the split can happen by checking metrics like $$R^2$$
* It then selects the attribute which has the smallest Entropy or Largest Information gain
* The set is then split by the selected attribute to produce a subset of the data
* The algorithm continues to recur on each subset, considering only attributes never selected before

Decision trees have a strong tendency to overfit the data. So practical uses of the decision tree must necessarily incorporate some ’regularization’ measures to ensure the decision tree built does not become more complex than is necessary and starts to overfit. There are broadly two ways of regularization on decision trees:

* **Truncation:** Truncate the decision tree during the training (growing) process preventing it from degenerating into one with one leaf for every data point in the training dataset. Below criterion are used to decide if the decision tree needs to be grown further:
  * Minimum Size of the Partition for a Split: Stop partitioning further when the current partition is small enough.
  * Minimum Change in Homogeneity Measure: Do not partition further when even the best split causes an insignificant change in the purity measure (difference between the current purity and the purity of the partitions created by the split).
  * Limit on Tree Depth: If the current node is farther away from the root than a threshold, then stop partitioning further.
  * Minimum Size of the Partition at a Leaf: If any of partitions from a split has fewer than this threshold minimum, then do not consider the split. Notice the subtle difference between this condition and the minimum size required for a split.
  * Maxmimum number of leaves in the Tree: If the current number of the bottom-most nodes in the tree exceeds this limit then stop partitioning.
* **Pruning:** Let the tree grow to any complexity. However add a post-processing step in which we prune the tree in a bottom-up fashion starting from the leaves. It is more common to use pruning strategies to avoid overfitting in practical implementations. One popular approach to pruning is to use a validation set. This method called reduced-error pruning, considers every one of the test (non-leaf ) nodes for pruning. Pruning a node means removing the entire subtree below the node, making it a leaf, and assigning the majority class (or the average of the values in case it is regression) among the training data points that pass through that node. A node in the tree is pruned only if the decision tree obtained after the pruning has an accuracy that is no worse on the validation dataset than the tree prior to pruning. This ensures that parts of the tree that were added due to accidental irregularities in the data are removed, as these irregularities are not likely to repeat.

### Hyperparameters

Though there are various ways to truncate or prune trees, the `DecisionTreeClassifier` function in sklearn provides the following hyperparameters which you can control:

* `criterion (Gini/IG or entropy)`: It defines the function to measure the quality of a split. Sklearn supports “gini” criteria for Gini Index & “entropy” for Information Gain. By default, it takes the value “gini”.
* `max_features`: It defines the no. of features to consider when looking for the best split
* `max_depth`: denotes maximum depth of the tree. It can take any integer value or None. If None, then nodes are expanded until all leaves are pure or until all leaves contain less than min\_samples\_split samples. By default, it takes “None” value.
* `min_samples_split`: This tells above the minimum no. of samples reqd. to split an internal node
* `min_samples_leaf`: The minimum number of samples required to be at a leaf node

## Random Forest

*Ensemble means a group of things viewed as a whole rather than individually. In ensembles, a collection of models is used to make predictions, rather than individual models. Arguably, the most popular in the family of ensemble models is the random forest: an ensemble made by the combination of a large number of decision trees.*

For an ensemble to work, each model of the ensemble should comply with the following conditions:

* Each model should be diverse. Diversity ensures that the models serve complementary purposes, which means that the individual models make predictions independent of each other.
* Each model should be acceptable. Acceptability implies that each model is at least better than a random model.

Random forests are created using a special ensemble method called **bagging (Bootstrap Aggregation)**. Bootstrapping means creating bootstrap samples from a given data set. A bootstrap sample is created by sampling the given data set uniformly and with replacement. A bootstrap sample typically contains about $$30$$-$$70$$% data from the data set. Aggregation implies combining the results of different models present in the ensemble.

### Steps

* Create a bootstrap sample from the training set
* Now construct a decision tree using the bootstrap sample. While splitting a node of the tree, only consider a random subset of features. Every time a node has to split, a different random subset of features will be considered.
* Repeat the steps 1 and 2 for $$n$$ times, to construct $$n$$ trees in the forest. Remember each tree is constructed independently, so it is possible to construct each tree in parallel.
* While predicting a test case, each tree predicts individually, and the final prediction is given by the majority vote of all the trees

### OOB (Out-of-Bag) Error

The OOB error is calculated by using each observation of the training set as a test observation. Since each tree is built on a bootstrap sample, each observation can used as a test observation by those trees which did not have it in their bootstrap sample. All these trees predict on this observation and you get an error for a single observation. The final OOB error is calculated by calculating the error on each observation and aggregating it.

It turns out that the OOB error is as good as cross validation error.

### Advantages

* A random forest is more stable than any single decision tree because the results get averaged out; it is not affected by the instability and bias of an individual tree.
* A random forest is immune to the curse of dimensionality since only a subset of features is used to split a node.
* You can parallelize the training of a forest since each tree is constructed independently.
* You can calculate the OOB (Out-of-Bag) error using the training set which gives a really good estimate of the performance of the forest on unseen data. Hence there is no need to split the data into training and validation; you can use all the data to train the forest.

### Hyperparameters

* `n_estimators`: The number of trees in the forest.
* `criterion`: The function to measure the quality of a split. Supported criteria are “gini” for the Gini impurity and “entropy” for the information gain. This parameter is tree-specific.
* `max_features`: The number of features to consider when looking for the best split
* `max_depth`: The maximum depth of the tree
* `min_samples_split`: The minimum number of samples required to split an internal node
* `min_samples_leaf`: The minimum number of samples required to be at a leaf node
* `min_weight_fraction_leaf`: The minimum weighted fraction of the sum total of weights (of all the input samples) required to be at a leaf node. Samples have equal weight when sample\_weight is not provided
* `max_leaf_nodes`: Grow trees with max\_leaf\_nodes in best-first fashion. Best nodes are defined as relative reduction in impurity
* `min_impurity_split`: Threshold for early stopping in tree growth. A node will split if its impurity is above the threshold, otherwise it is a leaf

## Ensemble Learning

Ensemble Learning is called the Wisdom of the crowd. Combine multiple weak models/learners into one predictive model to reduce bias, variance and/or improve accuracy.

* **Bagging:** Trains N different weak models (usually of same types – homogenous) with N non-overlapping subset of the input dataset in parallel. In the test phase, each model is evaluated. The label with the greatest number of predictions is selected as the prediction. Bagging methods reduces variance of the prediction
* **Boosting:** Trains N different weak models (usually of same types – homogenous) with the complete dataset in a sequential order. The datapoints wrongly classified with previous weak model is provided more weights to that they can be classified by the next weak leaner properly. In the test phase, each model is evaluated and based on the test error of each weak model, the prediction is weighted for voting. Boosting methods decreases the bias of the prediction.
* **Stacking:** Trains N different weak models (usually of different types – heterogenous) with one of the two subsets of the dataset in parallel. Once the weak learners are trained, they are used to trained a meta learner to combine their predictions and carry out final prediction using the other subset. In test phase, each model predicts its label, these set of labels are fed to the meta learner which generates the final prediction.

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2F71HYuVItf521RNQaD4LS%2Fimage25.png?alt=media&amp;token=0068fa00-07c8-482a-a498-e79ad07697e6" alt=""><figcaption><p>Comparison for the above 3 methods <a href="https://www.cheatsheets.aqeel-anwar.com">(Source)</a></p></figcaption></figure>

## Boosting

Boosting was first introduced in 1997 by Freund and Schapire in the popular algorithm, AdaBoost. It was originally designed for classification problems. Since its inception, many new boosting algorithms have been developed those tackle regression problems also and have become famous as they are used in the top solutions of many Kaggle competitions.

An ensemble is a collection of models which ideally should predict better than individual models. The key idea of boosting is to create an ensemble which makes high errors only on the less frequent data points. Boosting leverages the fact that we can build a series of models specifically targeted at the data points which have been incorrectly predicted by the other models in the ensemble. If a series of models keep reducing the average error, we will have an ensemble having extremely high accuracy. Boosting is a way of generating a strong model from a weak learning algorithm.

## AdaBoost

[📖Read](https://towardsdatascience.com/log-book-adaboost-the-math-behind-the-algorithm-a014c8afbbcc)

* AdaBoost starts by assigning equal weight to each datapoint, the idea is to adjust the weights of each observation after every iteration such that the the algorithm is forced to take a harder look at these difficult to classify observations
* Post each iteration we will have a weak learner using which we will calculate 2 things:
  * the updated weights of each $$N$$ observation for the next iteration
  * the weight that the weak learner itself will have on the final output, in each of the $$t$$ iterations we will have a learner $$h\_1$$, $$h\_2$$, $$h\_3$$ .. $$h\_t$$ each of which will be combined to make the final model, the weight of each of these individual learners in the final output is given by $$\alpha\_t$$. The models with low error rate will have higher values of $$\alpha\_t$$ and hence higher weight in the final output.
* Before you apply the AdaBoost algorithm, you should specifically remove the Outliers. Since AdaBoost tends to boost up the probabilities of misclassified points and there is a high chance that outliers will be misclassified, it will keep increasing the probability associated with the outliers and make the progress difficult.

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2F1midF0wXivou9knT36SZ%2Fimage10.png?alt=media&amp;token=ba6140e2-e77b-46a2-93b7-6191700ca573" alt=""><figcaption><p>Adaboost Pseudocode</p></figcaption></figure>

***

## Gradient Boosting

* As always let's start with a crude initial function F₀, something like average of all values in case of regression. It will give us some output, however bad.
* Calculate the loss function
* Next, we should have fit a new model on the residuals given by the Loss function, but there is a subtle twist: we will instead fit on the negative gradient of the loss function (for mathematical proof check the [link](https://towardsdatascience.com/log-book-xgboost-the-math-behind-the-algorithm-54ddc5008850)
* This process of fitting the model iteratively on the -ve gradient will continue till we have reached the minima or the limit of the number of weak learners given by T, this is called the additive approach
* Recall that, in Adaboost,“shortcomings” are identified by high-weight data points. In Gradient Boosting, “shortcomings” are identified by gradients. This is in short of the intuition as to how Gradient Boosting works. In case of regression and classification the only thing that differs is the loss function that is used.

## XGBoost

[📖Read](https://towardsdatascience.com/log-book-xgboost-the-math-behind-the-algorithm-54ddc5008850)

XGBoost stands for "Extreme Gradient Boosting", where the term "Gradient Boosting" originates from the paper Greedy Function Approximation: A Gradient Boosting Machine, by Friedman. It initially started as a research project by Tianqi Chen as part of the Distributed (Deep) Machine Learning Community (DMLC) group. It became well known in the ML competition circles after its use in the winning solution of the Higgs Machine Learning Challenge.

XGBoost and GBM both follow the principle of gradient boosted trees, but XGBoost uses a more regularized (by taking the model complexity into account) model formulation to control over-fitting, which gives it better performance, which is why it’s also known as ‘regularized boosting’ technique. In Stochastic Gradient Descent, used by Gradient Boosting, we use less point to take less time to compute the direction we should go towards, in order to make more of them, in the hope we go there quicker. In Newton’s method, used by XGBoost, we take more time to compute the direction we want to go into, in the hope we have to take fewer steps in order to get there.

**Why is XGBoost so good?**

* *Parallel Computing:* when you run xgboost, by default, it would use all the cores of your laptop/machine enabling its capacity to do parallel computation
* *Regularization:* The biggest advantage of xgboost is that it uses regularization and controls the overfitting and simplicity of the model which gives it better performance.
* *Enabled Cross Validation:* XGBoost is enabled with internal Cross Validation function
* *Missing Values:* XGBoost is designed to handle missing values internally. The missing values are treated in such a manner that if there exists any trend in missing values, it is captured by the model
* *Flexibility:* XGBoost is not just limited to regression, classification, and ranking problems, it supports user-defined objective functions as well. Furthermore, it supports user-defined evaluation metrics as well.

## Questions

<details>

<summary>[AMAZON] Random Forest Explanation</summary>

How does random forest generate the forest? Additionally why would we use it over other algorithms such as logistic regression?

**Answer**

Random forest is a supervised learning algorithm. The "forest" it builds, is an ensemble of decision trees, usually trained with the “bagging” method. The general idea of the bagging method is that a combination of learning models increases the overall result. Put simply: random forest builds multiple decision trees and merges them together to get a more accurate and stable prediction.

Steps:

* Create a bootstrap sample from the training set
* Now construct a decision tree using the bootstrap sample. While splitting a node of the tree, only consider a random subset of features. Every time a node has to split, a different random subset of features will be considered.
* Repeat the steps 1 and 2 for $$n$$ times, to construct $$n$$ trees in the forest. Remember each tree is constructed independently, so it is possible to construct each tree in parallel.
* While predicting a test case, each tree predicts individually, and the final prediction is given by the majority vote of all the trees

As for why or when we use it over logistic regression, the answer is it depends:

* If your problem/data is linearly separable, then first try logistic regression. If you don’t know, then still start with logistic regression because that will be your baseline, followed by non-linear classifier such as random forest. Do not forget to tune the parameters of logistic regression / random forest for maximizing their performance on your data.
* If your data is categorical, then random forest should be your first choice; however, logistic regression can be dealt with categorical data
* If you want easy to understand results, logistic regression is a better choice because it leads to simple interpretation of the explanatory variables.
* If speed is your criteria, then logistic regression should be your choice
* If your data is unbalanced, then random forest may be a better choice
* If number of data objects are less than the number of features, logistic regression should not be used
* Lastly for either of the random forest or logistic regression “models appear to perform similarly across the datasets with performance more influenced by choice of dataset rather than model selection”

</details>

<details>

<summary>Explain CART</summary>

[📖Source](https://www.analyticsvidhya.com/blog/2021/05/25-questions-to-test-your-skills-on-decision-trees/)

Explain the CART Algorithm for Decision Trees.

**Answer**

The CART stands for Classification and Regression Trees is a greedy algorithm that greedily searches for an optimum split at the top level, then repeats the same process at each of the subsequent levels.

Moreover, it does verify whether the split will lead to the lowest impurity or not as well as the solution provided by the greedy algorithm is not guaranteed to be optimal, it often produces a solution that’s reasonably good since finding the optimal Tree is an NP-Complete problem that requires exponential time complexity.

As a result, it makes the problem intractable even for small training sets. This is why we must go for a “reasonably good” solution instead of an optimal solution.

</details>

<details>

<summary>Properties of Gini Impurity</summary>

[📖Source](https://www.analyticsvidhya.com/blog/2021/05/25-questions-to-test-your-skills-on-decision-trees/)

Briefly explain the properties of Gini Impurity.

**Answer**

Gini Impurity is the probability of incorrectly classifying a randomly chosen element in the dataset if it were randomly labeled according to the class distribution in the dataset. It’s calculated as

$$\text{Gini Impurity} = 1 - \text{Gini Index}$$

So there can be 2 cases:

* When all the data points belong to a single class: $$G = 1 - (1^2 + 0^2) = 0$$
* When $$50%$$ of the data points belong to a class: $$G = 1 - (0.5^2 + 0.5^2) = 0.5$$

<img src="https://github.com/dipranjan/dsinterviewqns/blob/gitbook/Algorithms/images/image12.PNG" alt="image 12" data-size="original">

Gini impurity tends to isolate the most frequent class in its own branch of the Tree, while entropy tends to produce slightly more balanced Trees.

</details>

<details>

<summary>CART vs ID3</summary>

[📖Source](https://www.analyticsvidhya.com/blog/2021/05/25-questions-to-test-your-skills-on-decision-trees/)

Explain the difference between the CART and ID3 Algorithms.

**Answer**

The CART algorithm produces only binary Trees: non-leaf nodes always have two children (i.e., questions only have yes/no answers).

On the contrary, other Tree algorithms such as ID3 can produce Decision Trees with nodes having more than two children.

</details>

<details>

<summary>Types of nodes</summary>

[📖Source](https://www.analyticsvidhya.com/blog/2021/05/25-questions-to-test-your-skills-on-decision-trees/)

List down the different types of nodes in Decision Trees.

**Answer**

The Decision Tree consists of the following different types of nodes:

* Root node: It is the top-most node of the Tree from where the Tree starts
* Decision nodes: One or more Decision nodes that result in the splitting of data into multiple data segments and our main goal is to have the children nodes with maximum homogeneity or purity
* Leaf nodes: These nodes represent the data section having the highest homogeneity

</details>

<details>

<summary>Information Gain</summary>

[📖Source](https://www.analyticsvidhya.com/blog/2021/05/25-questions-to-test-your-skills-on-decision-trees/)

What do you understand of Information gain? Any disadvantages that you can think of?

**Answer**

Information gain is the difference between the entropy of a data segment before the split and after the split i.e, reduction in impurity due to the selection of an attribute.

Some points keep in mind about information gain:

* The high difference represents high information gain.
* Higher the difference implies the lower entropy of all data segments resulting from the split.
* Thus, the higher the difference, the higher the information gain, and the better the feature used for the split. Mathematically, the information gain can be computed by the equation as follows:

Information Gain = $$E(S1) – E(S2)$$, $$E(S1)$$ denotes the entropy of data belonging to the node before the split and $$E(S2)$$ denotes the weighted summation of the entropy of children nodes by considering the weights as the proportion of data instances falling in specific children nodes.

As for the disadvantages, Information gain biases the Decision Tree against considering attributes with a large number of distinct values which might lead to overfitting. In order to solve this problem, the Information Gain Ratio is used.

</details>

<details>

<summary>Space Time complexity of a Decision tree</summary>

[📖Source](https://www.analyticsvidhya.com/blog/2021/05/25-questions-to-test-your-skills-on-decision-trees/)

Explain the time and space complexity of training and testing in the case of a Decision Tree.

**Answer**

What happens in the training stage is that for each of the features (dimensions) in the dataset we’ll sort the data which takes $$O(n log n)$$ time following which we traverse the data points to find the right threshold which takes $$O(n)$$ time. For $$d$$ dimensions, total time complexity would be:

$$O(n \* log n \* d) + O(n\*d) \text{which asymptotically is } O(n \* log n \* d)$$

Train space complexity: The things we need while training a decision tree are the nodes which are typically stored as if-else conditions. Hence, the train space complexity would be: $$O(nodes)$$

Test time complexity would be $$O(depth)$$ since we have to move from root to a leaf node of the decision tree. Test space complexity would be $$O(nodes)$$

For Random forest the same would be:

Training Time Complexity = $$O(n*log(n)*d*k)$$, $$k$$=number of Decision Trees Notes: When we have a large number of data with reasonable features. Then we can use multi-core to parallelize our model to train different Decision Trees. Run-time Complexity= $$O(depth of tree* k)$$ Space Complexity= $$O(depth of tree \*k)$$

Note: Random Forest is comparatively faster than other algorithms.

</details>

<details>

<summary>Training time</summary>

[📖Source](https://www.analyticsvidhya.com/blog/2021/05/25-questions-to-test-your-skills-on-decision-trees/)

If it takes one hour to train a Decision Tree on a training set containing 1 million instances, roughly how much time will it take to train another Decision Tree on a training set containing 10 million instances?

**Answer**

As we know that the computational complexity of training a Decision Tree is given by O(n × m log(m)). So, when we multiplied the size of the training set by 10, then the training time will be multiplied by some factor, say K.

Now, we have to determine the value of K. To finds K, divide the complexity of both:

$$K = (n × 10m × log(10m)) / (n × m × log(m)) = 10 × log(10m) / log(m)$$

For $$10$$ million instances i.e., $$m = 106$$, then we get the value of $$K ≈ 11.7$$.

Therefore, we can expect the training time to be roughly $$11.7$$ hours.

</details>

<details>

<summary>Missing data and neumerical values</summary>

[📖Source](https://www.analyticsvidhya.com/blog/2021/05/25-questions-to-test-your-skills-on-decision-trees/)

How does a Decision Tree handle missing attribute values? How does it deal with continuous(numerical) features?

**Answer**

Decision Trees handle missing values in the following ways:

* Fill the missing attribute value by the most common value of that attribute
* Fill the missing value by assigning a probability to each of the possible values of the attribute based on other samples

Decision Trees handle continuous features by converting these continuous features to a threshold-based boolean feature. To decide The threshold value, we use the concept of Information Gain, choosing that threshold that maximizes the information gain.

</details>

<details>

<summary>Inductive Bias</summary>

[📖Source](https://www.analyticsvidhya.com/blog/2021/05/25-questions-to-test-your-skills-on-decision-trees/)

What is the Inductive Bias of Decision Trees?

**Answer**

The ID3 algorithm preferred Shorter Trees over longer Trees. In Decision Trees, attributes having high information gain are placed close to the root are preferred over those that do not.

</details>

<details>

<summary>Compare different selection measures</summary>

[📖Source](https://www.analyticsvidhya.com/blog/2021/05/25-questions-to-test-your-skills-on-decision-trees/)

Compare the different attribute selection measures.

**Answer**

The three measures, in general, returns good results, but:

* Information Gain: It is biased towards multivalued attributes
* Gain ratio: It prefers unbalanced splits in which one data segment is much smaller than the other segment
* Gini Index: It is biased to multivalued attributes, has difficulty when the number of classes is large, tends to favor tests that result in equal-sized partitions and purity in both partitions

</details>

<details>

<summary>Effect of Outliers</summary>

[📖Source](https://www.analyticsvidhya.com/blog/2021/05/25-questions-to-test-your-skills-on-decision-trees/)

Are Decision Trees affected by the outliers? Explain

**Answer**

Decision Trees are not sensitive to noisy data or outliers since, extreme values or outliers, never cause much reduction in Residual Sum of Squares(RSS), because they are never involved in the split.

</details>

<details>

<summary>Advantages and Disadvantages of Decision Tree</summary>

[📖Source](https://www.analyticsvidhya.com/blog/2021/05/25-questions-to-test-your-skills-on-decision-trees/)

Discuss the advantages and disadvantages of Decision tree

**Answer**

**Advantages:**

* **Clear Visualization:** This algorithm is simple to understand, interpret and visualize as the idea is mostly used in our daily lives. The output of a Decision Tree can be easily interpreted by humans.
* **Simple and easy to understand:** Decision Tree works in the same manner as simple if-else statements which are very easy to understand.
* This can be used for both classification and regression problems.
* Decision Trees can handle both continuous and categorical variables.
* **No feature scaling required:** There is no requirement of feature scaling techniques such as standardization and normalization in the case of Decision Tree as it uses a rule-based approach instead of calculation of distances.
* **Handles nonlinear parameters efficiently:** Unlike curve-based algorithms, the performance of decision trees can’t be affected by the Non-linear parameters. So, if there is high non-linearity present between the independent variables, Decision Trees may outperform as compared to other curve-based algorithms.
* Decision Tree can automatically handle missing values.
* Decision Tree handles the outliers automatically, hence they are usually robust to outliers.
* **Less Training Period:** The training period of decision trees is less as compared to ensemble techniques like Random Forest because it generates only one Tree unlike the forest of trees in the Random Forest.

**Disadvantages:**

* **Overfitting:** This is the major problem associated with the Decision Trees. It generally leads to overfitting of the data which ultimately leads to wrong predictions for testing data points. it keeps generating new nodes in order to fit the data including even noisy data and ultimately the Tree becomes too complex to interpret. In this way, it loses its generalization capabilities. Therefore, it performs well on the training dataset but starts making a lot of mistakes on the test dataset.
* **High variance:** As mentioned, a Decision Tree generally leads to the overfitting of data. Due to the overfitting, there is more likely a chance of high variance in the output which leads to many errors in the final predictions and shows high inaccuracy in the results. So, in order to achieve zero bias (overfitting), it leads to high variance due to the bias-variance tradeoff.
* **Unstable:** When we add new data points it can lead to regeneration of the overall Tree. Therefore, all nodes need to be recalculated and reconstructed.
* **Not suitable for large datasets:** If the data size is large, then one single Tree may grow complex and lead to overfitting. So in this case, we should use Random Forest instead, an ensemble technique of a single Decision Tree.

</details>


# Time Series Analysis

**Reference:** [📖Explanation](https://nbviewer.jupyter.org/github/Yorko/mlcourse_open/blob/master/jupyter_english/topic09_time_series/topic9_part1_time_series_python.ipynb)

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2Fgit-blob-c051fca109239dab5b87b63256d44903003a50f1%2Fimage14.png?alt=media" alt=""><figcaption></figcaption></figure>

A time series is simply a series of data points ordered in time. In a time series, time is often the independent variable and the goal is usually to make a forecast for the future.

* **Moving Average:** Here the assumption is that future value of our variable depends on the average of its $$k$$ previous values. Moving average has another use case - smoothing the original time series to identify trends. The wider the window, the smoother the trend. In the case of very noisy data, which is often encountered in finance, this procedure can help detect common patterns. This can also be used to determine anamolies based on the confidence level.
* **Weighted average:** It is a simple modification to the moving average. The weights sum up to `1` with larger weights assigned to more recent observations. $$\hat{y}*{t} = \displaystyle\sum^{k}*{n=1} \omega\_n y\_{t+1-n}$$
* **Exponential smoothing:** Here instead of weighting the last $$k$$ values of the time series, we start weighting all available observations while exponentially decreasing the weights as we move further back in time. $$\hat{y}*{t} = \alpha \cdot y\_t + (1-\alpha) \cdot \hat y*{t-1}$$

  The $$\alpha$$ weight is called a smoothing factor. It defines how quickly we will "forget" the last available true observation. The smaller $$\alpha$$ is, the more influence the previous observations have and the smoother the series is.
* **Double Exponential smoothing:** Up to now, the methods that we've discussed have been for a single future point prediction (with some nice smoothing). That is cool, but it is also not enough. Let's extend exponential smoothing so that we can predict two future points (of course, we will also include more smoothing).

  Series decomposition will help us -- we obtain two components: intercept (i.e. level) $$\ell$$ and slope (i.e. trend) $$b$$. We have learnt to predict intercept (or expected series value) with our previous methods; now, we will apply the same exponential smoothing to the trend by assuming that the future direction of the time series changes depends on the previous weighted changes. As a result, we get the following set of functions:

  $$\ell\_x = \alpha y\_x + (1-\alpha)(\ell\_{x-1} + b\_{x-1})$$

  $$b\_x = \beta(\ell\_x - \ell\_{x-1}) + (1-\beta)b\_{x-1}$$

  $$\hat{y}\_{x+1} = \ell\_x + b\_x$$

  The first one describes the intercept, which, as before, depends on the current value of the series. The second term is now split into previous values of the level and of the trend. The second function describes the trend, which depends on the level changes at the current step and on the previous value of the trend. In this case, the $$\beta$$ coefficient is a weight for exponential smoothing. The final prediction is the sum of the model values of the intercept and trend.
* **Triple exponential smoothing a.k.a. Holt-Winters:**

  The idea is to add a third component - seasonality. This means that we should not use this method if our time series is not expected to have seasonality. Seasonal components in the model will explain repeated variations around intercept and trend, and it will be specified by the length of the season, in other words by the period after which the variations repeat. For each observation in the season, there is a separate component; for example, if the length of the season is 7 days (a weekly seasonality), we will have 7 seasonal components, one for each day of the week.

  The new system of equations:

  $$\ell\_x = \alpha(y\_x - s\_{x-L}) + (1-\alpha)(\ell\_{x-1} + b\_{x-1})$$

  $$b\_x = \beta(\ell\_x - \ell\_{x-1}) + (1-\beta)b\_{x-1}$$

  $$s\_x = \gamma(y\_x - \ell\_x) + (1-\gamma)s\_{x-L}$$

  $$\hat{y}*{x+m} = \ell\_x + mb\_x + s*{x-L+1+(m-1)modL}$$

  The intercept now depends on the current value of the series minus any corresponding seasonal component. Trend remains unchanged, and the seasonal component depends on the current value of the series minus the intercept and on the previous value of the component. Take into account that the component is smoothed through all the available seasons; for example, if we have a Monday component, then it will only be averaged with other Mondays. You can read more on how averaging works and how the initial approximation of the trend and seasonal components is done [here](http://www.itl.nist.gov/div898/handbook/pmc/section4/pmc435.htm). Now that we have the seasonal component, we can predict not just one or two steps ahead but an arbitrary $$m$$ future steps ahead, which is very encouraging.

  Below is the code for a triple exponential smoothing model, which is also known by the last names of its creators, Charles Holt and his student Peter Winters. Additionally, the Brutlag method was included in the model to produce confidence intervals:

  $$\hat y\_{max\_x}=\ell\_{x−1}+b\_{x−1}+s\_{x−T}+m⋅d\_{t−T}$$

  $$\hat y\_{min\_x}=\ell\_{x−1}+b\_{x−1}+s\_{x−T}-m⋅d\_{t−T}$$

  $$d\_t=\gamma∣y\_t−\hat y\_t∣+(1−\gamma)d\_{t−T},$$

  where $$T$$ is the length of the season, $$d$$ is the predicted deviation. Other parameters were taken from triple exponential smoothing. You can read more about the method and its applicability to anomaly detection in time series [here](http://fedcsis.org/proceedings/2012/pliks/118.pdf).

  Exponentiality is hidden in the recursiveness of the function – we multiply by $$(1-\alpha)$$ each time, which already contains a multiplication by $$(1-\alpha)$$ of previous model values.

### Stationarity

Before we start modeling, we should mention such an important property of time series, **stationarity**.

So why is stationarity so important? Because it is easy to make predictions on a stationary series since we can assume that the future statistical properties will not be different from those currently observed. Most of the time-series models, in one way or the other, try to predict those properties (mean or variance, for example). Furture predictions would be wrong if the original series were not stationary.

When running a linear regression the assumption is that all of the observations are all independent of each other. In a time series, however, we know that observations are time dependent. It turns out that a lot of nice results that hold for independent random variables (law of large numbers and central limit theorem to name a couple) hold for stationary random variables. So by making the data stationary, we can actually apply regression techniques to this time dependent variable.

**Dickey-Fuller test** can be used as a check for stationarity. If ‘Test Statistic’ is greater than the ‘Critical Value’ then the time series is stationary.

There are a few ways to deal with non-stationarity:

* Deflation by CPI
* Logarithmic
* First Difference
* Seasonal Difference
* Seasonal Adjustment

Plot the ACF and PACF charts and find the optimal parameters.

### ARIMA family

We will explain this model by building up letter by letter. $$SARIMA(p, d, q)(P, D, Q, s)$$, Seasonal Autoregression Moving Average model:

* $$AR(p)$$ - autoregression model i.e. regression of the time series onto itself. The basic assumption is that the current series values depend on its previous values with some lag (or several lags). The maximum lag in the model is referred to as $$p$$. To determine the initial $$p$$, you need to look at the PACF plot and find the biggest significant lag after which **most** other lags become insignificant.
* $$MA(q)$$ - moving average model. Without going into too much detail, this models the error of the time series, again with the assumption that the current error depends on the previous with some lag, which is referred to as $$q$$. The initial value can be found on the ACF plot with the same logic as before.

Let's combine our first 4 letters:

$$AR(p) + MA(q) = ARMA(p, q)$$

What we have here is the Autoregressive–moving-average model! If the series is stationary, it can be approximated with these 4 letters. Let's continue.

* $$I(d)$$ - order of integration. This is simply the number of nonseasonal differences needed to make the series stationary.

Adding this letter to the four gives us the $$ARIMA$$ model which can handle non-stationary data with the help of nonseasonal differences. Great, one more letter to go!

* $$S(s)$$ - this is responsible for seasonality and equals the season period length of the series

With this, we have three parameters: $$(P, D, Q)$$

* $$P$$ - order of autoregression for the seasonal component of the model, which can be derived from PACF. But you need to look at the number of significant lags, which are the multiples of the season period length. For example, if the period equals 24 and we see the 24-th and 48-th lags are significant in the PACF, that means the initial $$P$$ should be 2.
* $$Q$$ - similar logic using the ACF plot instead.
* $$D$$ - order of seasonal integration. This can be equal to 1 or 0, depending on whether seasonal differeces were applied or not.

### Prophet

[📖 Check out this discussion](https://stats.stackexchange.com/questions/472266/inference-in-time-series-prophet-vs-arima)

ARIMA and similar models assume some sort of causal relationship between past values and past errors and future values of the time series. Facebook Prophet doesn't look for any such causal relationships between past and future. Instead, it simply tries to find the best curve to fit to the data, using a linear or logistic curve, and Fourier coefficients for the seasonal components. There is also a regression component, but that is for external regressors, not for the time series itself (The Prophet model is a special case of GAM - Generalized Additive Model).

### Metrics

* **R squared:** coefficient of determination (in econometrics, this can be interpreted as the percentage of variance explained by the model), $$(-\infty, 1]$$

$$R^2 = 1 - \frac{SS\_{res}}{SS\_{tot}}$$

* **Mean Absolute Error:** this is an interpretable metric because it has the same unit of measurment as the initial series, $$\[0, +\infty)$$

$$MAE = \frac{\sum\limits\_{i=1}^{n} |y\_i - \hat{y}\_i|}{n}$$

* **Median Absolute Error:** again, an interpretable metric that is particularly interesting because it is robust to outliers, $$\[0, +\infty)$$

$$MedAE = median(|y\_1 - \hat{y}\_1|, ... , |y\_n - \hat{y}\_n|)$$

* **Mean Squared Error:** the most commonly used metric that gives a higher penalty to large errors and vice versa, $$\[0, +\infty)$$

$$MSE = \frac{1}{n}\sum\limits\_{i=1}^{n} (y\_i - \hat{y}\_i)^2$$

* **Mean Squared Logarithmic Error:** practically, this is the same as MSE, but we take the logarithm of the series. As a result, we give more weight to small mistakes as well. This is usually used when the data has exponential trends, $$\[0, +\infty)$$

$$MSLE = \frac{1}{n}\sum\limits\_{i=1}^{n} (log(1+y\_i) - log(1+\hat{y}\_i))^2$$

* **Mean Absolute Percentage Error:** this is the same as MAE but is computed as a percentage, which is very convenient when you want to explain the quality of the model to management, $$\[0, +\infty)$$

$$MAPE = \frac{100}{n}\sum\limits\_{i=1}^{n} \frac{|y\_i - \hat{y}\_i|}{y\_i}$$

***

### Questions

<details>

<summary>Cross Validation with Time Series</summary>

Can cross validation be used with Time Series to estimate model parameters automatically?

**Answer**

Normal cross-validation cannot be used for time series because one cannot randomly mix values in a fold while preserving this structure. With randomization, all time dependencies between observations will be lost. But something like "cross-validation on a rolling basis" can be used.

The idea is rather simple -- we train our model on a small segment of the time series from the beginning until some $$t$$, make predictions for the next $$t+n$$ steps, and calculate an error. Then, we expand our training sample to $$t+n$$ value, make predictions from $$t+n$$ until $$t+2\*n$$, and continue moving our test segment of the time series until we hit the last available observation. As a result, we have as many folds as $$n$$ will fit between the initial training sample and the last observation.

</details>

<details>

<summary>CNNs in Time Series</summary>

How are CNNs used for Time Series Prediction?

**Answer**

* The ability of CNNs to learn and automatically extract features from raw input data can be applied to time series forecasting problems. A sequence of observations can be treated like a one-dimensional image that a CNN model can read and distill into the most salient elements.
* The capability of CNNs has been demonstrated to great effect on time series classification tasks such as automatically detecting human activities based on raw accelerator sensor data from fitness devices and smartphones.
* CNNs have the support for multivariate input, multivariate output, it can learn arbitrary but complex functional relationships, but does not require that the model learn directly from lag observations. Instead, the model can learn a representation from a large input sequence that is most relevant for the prediction problem.

</details>

<details>

<summary>Data Prep for Time Series</summary>

What are some of Data Preprocessing Operations you would use for Time Series Data?

**Answer**

It depends on the problem, but some common ones are:

* Parsing time series information from various sources and formats.
* Generating sequences of fixed-frequency dates and time spans.
* Manipulating and converting date times with time zone information.
* Resampling or converting a time series to a particular frequency.
* Performing date and time arithmetic with absolute or relative time increments.

</details>

<details>

<summary>Missing Value in Time Series</summary>

What are some of best ways to handle missing values in Time Series Data?

**Answer**

The most common methodology used for handling missing, unequally spaced, or unsynchronized values is *linear interpolation*.

The idea is to create estimated values at the desired time stamps. These can be used to generate multivariate time series that are synchronized, equally spaced, and have no missing values. Consider the scenario where $$y\_i$$ and $$y\_j$$ are values for the time series at times $$t\_i$$ and $$t\_j$$, repectively, where $$i < j$$. Let $$t$$ be a time drawn from the interval $$(t\_i, t\_j)$$. Then, the interpolated value of the series is given by: $$y = y\_i + \frac{t-t\_i}{t\_j-t\_i}\*(y\_j-y\_i)$$

</details>

<details>

<summary>Stationarity</summary>

Can you explain why time series has to be stationary?

**Answer**

Stationarity is important because, in its absence, a model describing the data will vary in accuracy at different time points. As such, stationarity is required for sample statistics such as means, variances, and correlations to accurately describe the data at all time points of interest.

Looking at the time series plots below, you can notice how the mean and variance of any given segment of time would do a good job representing the whole stationary time series but a relatively poor job representing the whole non-stationary time series. For instance, the mean of the non-stationary time series is much lower from $$600\<t<800$$ and its variance is much higher in this range than in the range from $$200\<t<400$$.

<img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2Fgit-blob-d7f17d9e61f99bcbbb0b7e2e39c1405e769f9e78%2Fimage20.png?alt=media" alt="" data-size="original">

What quantities are we typically interested in when we perform statistical analysis on a time series? We want to know

* Its expected value,
* Its variance, and
* The correlation between values $$s$$ periods apart for a set of values.

To calculate these things we use a mean across many time periods. The mean across many time periods is only informative if the expected value is the same across those time periods. If these population parameters can vary, what are we really estimating by taking an average across time?

(Weak) stationarity requires that these population quantities must be the same across time, making the sample average a reasonable way to estimate them.

</details>

<details>

<summary>IQR in Time Series</summary>

How is Interquartile range used in Time series?

**Answer**

It is mostly used to detect outliers in Time Series data.

</details>

<details>

<summary>Irregular Data</summary>

What does irregularly-spaced spatial data mean in Time series?

**Answer**

* A lot of techniques assume that data is sampled at regularly-spaced intervals of time. This interval between adjacent samples is called the *sampling period*.
* A lot of data is not or cannot be sampled with a fixed sampling period. For example, if we measure the atmosphere using sensors, the terrain may not allow us to place weather stations exactly 50 miles apart.

There are many different ways to deal with this kind of data which does not have a fixed sampling period. One approach is to interpolate the data onto a grid and then use a technique intended for gridded data.

</details>

<details>

<summary>Sliding Window</summary>

Explain the Sliding Window method in Time series?

**Answer**

* Time series can be phrased as supervised learning. Given a sequence of numbers for a time series dataset, we can restructure the data to look like a supervised learning problem.
* In the sliding window method, the previous time steps can be used as input variables, and the next time steps can be used as the output variable.

In statistics and time series analysis, this is called a lag or lag method. The number of previous time steps is called the window width or size of the lag. This sliding window is the basis for how we can turn any time series dataset into a supervised learning problem.

</details>

<details>

<summary>LSTM vs MLP</summary>

Can you discuss on the usage of LSTM vs MLP in Time Series?

**Answer**

Multilayer Perceptrons, or MLPs for short, can be applied to time series forecasting. A challenge with using MLPs for time series forecasting is in the preparation of the data. Specifically, lag observations must be flattened into feature vectors. My understanding is that LSTMs captures the relations between time steps, whereas simple MLPs treat each time step as a separated feature (doesn't take succession into consideration).

RNNs are known to be superior to MLP in case of sequential data. But complex models like LSTM and GRU require a lot of data to achieve their potential.

</details>

<details>

<summary>Sequential vs Non-Sequential</summary>

Can a CNN (or other non-sequential deep learning models) outperform LSTM (or other sequential models) in time series data

**Answer**

[*Source*](https://ai.stackexchange.com/questions/16818/can-non-sequential-deep-learning-models-outperform-sequential-models-in-time-ser)

You are right CNN based models can outperform RNN. You can take a look at this [paper](https://arxiv.org/pdf/1803.01271.pdf) where they compared different RNN models with TCN (temporal convolutional networks) on different sequence modeling tasks. Even though there are no big differences in terms of results there are some nice properties that CNN based models offers such as: parallelism, stable gradients and low training memory footprint. In addition to CNN based models there are also attention based models (you might want to take a look at the [transformer](https://papers.nips.cc/paper/7181-attention-is-all-you-need.pdf))

</details>

<details>

<summary>Long Time Series</summary>

What's the best architecture for time series prediction with a long dataset?

**Answer**

[*Source*](https://ai.stackexchange.com/questions/17802/whats-the-best-architecture-for-time-series-prediction-with-a-long-dataset)

LSTM is ideal for this. For even stronger representational capacity, make your LSTM's multi-layered. Using 1-dimensional convolutions in a CNN is a common way to exctract information from time series too, so there's no harm in trying. Typically, you'll test many models out and take the one that has best validation performance.

</details>

<details>

<summary>Correlation</summary>

How to use Correlation in Time series data?

**Answer**

[*Source*](https://stats.stackexchange.com/questions/133155/how-to-use-pearson-correlation-correctly-with-time-series)

Pearson correlation is used to look at correlation between series ... but being time series, the correlation is looked at across different lags -- the cross-correlation function. The cross-correlation is impacted by dependence within-series, so in many cases the within-series dependence should be removed first. So, to use this correlation, rather than smoothing the series, it's actually more common (because it's meaningful) to look at dependence between residuals - the rough part that's left over after a suitable model is found for the variables.

You probably want to begin with some basic resources on time series models before delving into trying to figure out whether a Pearson correlation across (presumably) nonstationary, smoothed series is interpretable.

In particular, you'll probably want to look into spurious correlation. The point about spurious correlation is that series can appear correlated, but the correlation itself is not meaningful. Consider two people tossing two distinct coins counting number of heads so far minus number of tails so far as the value of their series.

(So, if person-1 tosses $$HTHH$$... they have $$3-1 = 2$$ for the value at the $$4^{th}$$ time step, and their series goes $$1,0,1,2,....$$)

Obviously, there's no connection whatever between the two series. Clearly neither can tell you the first thing about the other!

But look at the sort of correlations you get between pairs of coins:

<img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2Fgit-blob-c9ffabf2812cca0162912300fbf724d1c8e31098%2Fimage21.png?alt=media" alt="" data-size="original">

If I didn't tell you what those were, and you took any pair of those series by themselves, those would be impressive correlations would they not?

But they're all meaningless. Utterly spurious. None of the three pairs are really any more positively or negatively related to each other than any of the others -- it's just cumulated noise. The spuriousness isn't just about prediction, the whole notion of considering association between series without taking account of the within-series dependence is misplaced.

All you have here is within-series dependence. There's no actual cross-series relation whatever.

Once you deal properly with the issue that makes these series auto-dependent - they're all integrated (Bernoulli random walks), so you need to difference them - the "apparent" association disappears (the largest absolute cross-series correlation of the three is 0.048).

What that tells you is the truth -- the apparent association is a mere illusion caused by the dependence within-series.

If your question asked "*how to use Pearson correlation correctly with time series*" -- so please understand: if there's within-series dependence and you don't deal with it first, you won't be using it correctly.

Further, smoothing won't reduce the problem of serial dependence; quite the opposite -- it makes it even worse! Here are the correlations after smoothing (default loess smooth - of series vs index - performed in R):

```
|       | coin1      | coin2      |
|-------|------------|------------|
| coin2 | 0.9696378  |            |
| coin3 | -0.8829326 | -0.7733559 |
```

They all got further from 0. They're all still nothing but meaningless noise, though now it's smoothed, cumulated noise. (By smoothing, we reduce the variability in the series we put into the correlation calculation, so that may be why the correlation goes up.)

*Check the link given above for a detailed discussion*

</details>

<details>

<summary>State Space Model and Kalman Filtering</summary>

Describe in details how State Space Model and Kalman Filtering are used in Time Series forecasting?

**Answer**

[*Resource*](https://www.google.com/url?sa=t\&rct=j\&q=\&esrc=s\&source=web\&cd=\&cad=rja\&uact=8\&ved=2ahUKEwjDh_W1_dD6AhWU-DgGHY5OAlMQFnoECBMQAQ\&url=https%3A%2F%2Ftowardsdatascience.com%2Fstate-space-model-and-kalman-filter-for-time-series-prediction-basic-structural-dynamic-linear-2421d7b49fa6\&usg=AOvVaw0u-sCm-kITU5I-Ptdc9K8s) \_\_ [*Source*](https://mfe.baruch.cuny.edu/wp-content/uploads/2014/12/TS_Lecture5_2019.pdf)

A state space model (SSM) is a time series model in which the time series $$Y\_t$$ is interpreted as the result of a noisy observation of a stochastic process $$X\_t$$. The values of the variables $$X\_t$$ and $$Y\_t$$ can be continuous (scalar or vector) or discrete. Graphically, an SSM is represented as follows:

<img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2Fgit-blob-3f95844a753713d98a7e03f1b5eacf252b4b378b%2Fimage22.png?alt=media" alt="" data-size="original">

SSMs belong to the realm of Bayesian inference, and they have been successfully applied in many fields to solve a broad range of problems. It is usually assumed that the state process $$X\_t$$ is Markovian. The most well studied SSM is the Kalman filter, for which the processes above are linear and the sources of randomness are Gaussian.

Let $$T$$ denote the time horizon. Our broad goal is to make inference about the states Xt based on a set of observations $$Y\_1, . . . , Y\_t$$. Three questions are of particular interest:

* **Filtering:** $$t < T$$. What can we infer about the current state of the system based on all available observations?
* **Smoothing:** $$t = T$$. What can be inferred about the system based on the information contained in the entire data sample? In particular, how can we back fill missing observations?
* **Forecasting:** $$t > T$$. What is the optimal prediction of a future observation and/or a future state of the system?

In principle, any inference for this model can be done using the standard methods of multivariate statistics. However, these methods require storing large amounts of data and inverting $$tn × tn$$ matrices. Notice that, as new data arrive, the storage requirements and matrix dimensionality increase. This is frequently computationally intractable and impractical. Instead, the Kalman filter relies on a recursive approach which does not require significant storage resources and involves inverting $$n × n$$ matrices only.

</details>

<details>

<summary>State Space Model vs Conventional Methodologies</summary>

What are disadvantages of state-space models and Kalman Filter for time-series modelling over let's say conventional methodologies like ARIMA, VAR or ad-hoc/heuristic methods?

**Answer**

[*Source*](https://stats.stackexchange.com/questions/78287/what-are-disadvantages-of-state-space-models-and-kalman-filter-for-time-series-m)

* Overall - compared to ARIMA, state-space models allow you to model more complex processes, have interpretable structure and easily handle data irregularities; but for this you pay with increased complexity of a model, harder calibration, less community knowledge.
* ARIMA is a universal approximator - you don't care what is the true model behind your data and you use universal ARIMA diagnostic and fitting tools to approximate this model. It is like a polynomial curve fitting - you don't care what is the true function, you always can approximate it with a polynomial of some degree.
* State-space models naturally require you to write-down some reasonable model for your process (which is good - you use your prior knowledge of your process to improve estimates). Of course, if you don't have any idea of your process, you always can use some universal state-space model also - e.g. represent ARIMA in a state-space form. But then ARIMA in its original form has more parsimonious formulation - without introducing unnecessary hidden states.
* Because there is such a great variety of state-space models formulations (much richer than class of ARIMA models), behavior of all these potential models is not well studied and if the model you formulated is complicated - it's hard to say how it will behave under different circumstances. Of course, if your state-space model is simple or composed of interpretable components, there is no such problem. But ARIMA is always the same well studied ARIMA so it should be easier to anticipate its behavior even if you use it to approximate some complex process.
* Because state-space allows you directly and exactly model complex/nonlinear models, then for these complex/nonlinear models you may have problems with stability of filtering/prediction (EKF/UKF divergence, particle filter degradation). You may also have problems with calibrating complicated-model's parameters - it's a computationally-hard optimization problem. ARIMA is simple, has less parameters (1 noise source instead of 2 noise sources, no hidden variables) so its calibration is simpler.
* For state-space there is less community knowledge and software in statistical community than for ARIMA.

</details>

<details>

<summary>Discrete Wavelet Transform</summary>

Have you heard of Discrete Wavelet Transform in time series?

**Answer**

[*Source*](https://stackoverflow.com/questions/68303601/wavelet-for-time-series)

A Discrete Wavelet Transform (DWT) allows you to decompose your input data into a set of discrete levels, providing you with information about the frequency content of the signal i.e. determining whether the signal contains high frequency variations or low frequency trends. Think of it as applying several band-pass filters to your input data.

</details>


# Anomaly Detection

An anomaly is something that differs from a norm: a deviation, an exception. In software engineering, by anomaly we understand a rare occurrence or event that doesn’t fit into the pattern, and, therefore, seems suspicious.

## Types

* **Global outliers:** When a data point assumes a value that is far outside all the other data point value ranges in the dataset, it can be considered a global anomaly. In other words, it’s a rare event.
* **Contextual outliers:** When an outlier is called contextual it means that its value doesn’t correspond with what we expect to observe for a similar data point in the same context. Contexts are usually temporal, and the same situation observed at different times can be not an outlier.
* **Collective outliers:** Collective outliers are represented by a subset of data points that deviate from the normal behavior.

## Methods of detecting Anomalies

There are three categories of outlier detection, namely, supervised, semi-supervised, and unsupervised:

* **Supervised:** Requires fully labeled training and testing datasets. An ordinary classifier is trained first and applied afterward. Here, the quality of the training dataset is very important, and it is a lot of manual work involved since somebody needs to collect and label examples. Due to these mostly unsupervised methods are used in Anamoly detection.
* **Semi-supervised:** Uses training and test datasets, whereas training data only consists of normal data without any outliers. A model of the normal class is learned, and outliers can be detected afterward by deviating from that model.
* **Unsupervised:** Does not require any labels; there is no distinction between a training and a test dataset Data is scored solely based on intrinsic properties of the dataset.

And three fundamental approaches to detect anomalies are based on:

* **By Density:** Normal points occur in dense regions, while anomalies occur in sparse regions
* **By Distance:** Normal point is close to its neighbors and anomaly is far from its neighbors
* **By Isolation:** The term isolation means ‘separating an instance from the rest of the instances’. Since anomalies are ‘few and different’ and therefore they are more susceptible to isolation.

### Isolation Forest

Isolation Forest is an unsupervised anomaly detection algorithm that uses a random forest algorithm (decision trees) under the hood to detect outliers in the dataset. The algorithm tries to split or divide the data points such that each observation gets isolated from the others. Usually, the anomalies lie away from the cluster of data points, so it's easier to isolate the anomalies compare to the regular data points.

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FNGvpuR7IPEnzW61EDtlq%2Fimage18.png?alt=media&amp;token=52fd0ebd-7fcb-4e65-ab3e-a8e55e67bd5c" alt=""><figcaption><p>Partitioning of Anomaly and Regular data point (<a href="https://towardsdatascience.com/5-anomaly-detection-algorithms-every-data-scientist-should-know-b36c3605ea16">Source)</a></p></figcaption></figure>

### Local Outlier Factor

It takes the density of data points into consideration to decide whether a point is an anomaly or not. The local outlier factor computes an anomaly score called anomaly score that measures how isolated the point is with respect to the surrounding neighborhood. It takes into account the local as well as the global density to compute the anomaly score.

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FmQfAWk9wL4JeEJiKzPZQ%2Fimage19.png?alt=media&amp;token=d2a9bce0-788e-4507-ae6f-6d87c7727325" alt=""><figcaption><p>Local Outlier Factor Formulation (<a href="https://medium.com/mlpoint/local-outlier-factor-a-way-to-detect-outliers-dde335d77e1a">Source)</a></p></figcaption></figure>

### Robust Covariance

For gaussian independent features, simple statistical techniques can be employed to detect anomalies in the dataset. For a gaussian/normal distribution, the data points lying away from 3rd deviation can be considered as anomalies.

For a dataset having all the feature gaussian in nature, then the statistical approach can be generalized by defining an elliptical hypersphere that covers most of the regular data points, and the data points that lie away from the hypersphere can be considered as anomalies.

### One Class SVM

A regular SVM algorithm tries to find a hyperplane that best separates the two classes of data points. For one-class SVM where we have one class of data points, and the task is to predict a hypersphere that separates the cluster of data points from the anomalies.

### One Class SVM (SGD)

One-class SVM with SGD solves the linear One-Class SVM using Stochastic Gradient Descent. The implementation is meant to be used with a kernel approximation technique to obtain results similar to `sklearn.svm.OneClassSVM` which uses a Gaussian kernel by default.

### Cluster-based Local Outlier Factor (CBLOF)

The CBLOF calculates the outlier score based on cluster-based local outlier factor. An anomaly score is computed by the distance of each instance to its cluster center multiplied by the instances belonging to its cluster.

### Histogram-based Outlier Detection (HBOS)

HBOS assumes the feature independence and calculates the degree of anomalies by building histograms. In multivariate anomaly detection, a histogram for each single feature can be computed, scored individually and combined at the end.

### KNN

It is one of the simplest methods in anomaly detection. For a data point, its distance to its kth nearest neighbor could be viewed as the outlier score.

For Univariate Analysis Isolation Forest can be used to detect outliers that returns the anomaly score of each sample. Isolation Forest is a tree-based model. In these trees, partitions are created by first randomly selecting a feature and then selecting a random split value between the minimum and maximum value of the selected feature. In multivariate anomaly detection, outlier is a combined unusual score on at least two variables. For Multivariate Analysis multiple options are avaliable other than isolation forest:

## Questions

<details>

<summary>[AKAMAI] Anomaly in Univariate Dataset</summary>

If given a univariate dataset, how would you design a function to detect anomalies?

What if the data is bivariate?

**Answer**

**Reference:** [📖Explanation](https://towardsdatascience.com/anomaly-detection-for-dummies-15f148e559c1)

Anomaly detection is the process of identifying unexpected items or events in data sets, which differ from the norm. And anomaly detection is often applied on unlabeled data which is known as unsupervised anomaly detection. Anomaly detection has two basic assumptions:

* Anomalies only occur very rarely in the data.
* Their features differ from the normal instances significantly.

This is part is mentioned above too, for Univariate Analysis Isolation Forest can be used to detect outliers that returns the anomaly score of each sample. Isolation Forest is a tree-based model. In these trees, partitions are created by first randomly selecting a feature and then selecting a random split value between the minimum and maximum value of the selected feature.

In multivariate anomaly detection, outlier is a combined unusual score on at least two variables. For Multivariate Analysis multiple options are avaliable other than isolation forest:

* The Cluster-based Local Outlier Factor (CBLOF)
* Histogram-based Outlier Detection (HBOS)
* KNN

An ensemble of these methods can be used to finalize the anamolies. Always visually investigate some of the anomalies.

</details>

<details>

<summary>Swamping VS Masking</summary>

What are the Swamping and Masking problems in Anomaly Detection?

**Answer**

* Since anomalies are rare events, making it very difficult to label them with high accuracy, swamping is the phenomenon of labeling normal events as anomalies.
* When clustering algorithms are used, the data points belonging to different clusters get merged into one cluster, if the number of segments in the dataset is not known, this causes the outlier cluster to be merged to a cluster with normal data points. This causes the outliers to not be detected. This is defined as masking.

</details>

<details>

<summary>Uniform Distribution VS Normal Distribution</summary>

What are the differences in Anomalies for Uniform Distribution and Normal Distribution in One-Dimensional Data?

**Answer**

Keep in mind how the Uniform and Normal Distribution looks like.

**Uniform**

* When data is distributed uniformly over a finite range, the mean and standard deviation merely characterize the range of values.
* One possible indication of anomalous behavior could be that a small neighborhood contains substantially fewer or more data points than expected from a uniform distribution.

**Normal**

* A normal distribution follows the empirical rule, which states that 68%, 95%, and 99.7% of the values lie within one, two, and three standard deviations of the mean, respectively.
* About 0.1% of the points are more than $$3 \*\sigma$$ (three standard deviations) away from the mean, hence, it is taken as the threshold and points beyond that distance from the mean are declared to be anomalous.

</details>

<details>

<summary>SVM VS Logistic Regression</summary>

Compare SVM and Logistic Regression in handling outliers

**Answer**

* For Logistic Regression, outliers can have an unusually large effect on the estimate of logistic regression coefficients. It will find a linear boundary if it exists to accommodate the outliers. To solve the problem of outliers, sometimes a sigmoid function is used in logistic regression.
* For SVM, outliers can make the decision boundary deviate severely from the optimal hyperplane. One way for SVM to get around the problem is to intrduce slack variables. There is a penalty involved with using slack variables, and how SVM handles outliers depends on how this penalty is imposed.

</details>

<details>

<summary>Outlier VS Novelty</summary>

Explain the difference between Outlier Detection vs Novelty Detection

**Answer**

* The training data contains outliers which are defined as observations that are far from the others. Outlier detection estimators thus try to fit the regions where the training data is the most concentrated, ignoring the deviant observations.
* The training data is not polluted by outliers and we are interested in detecting whether a new observation is an outlier. In this context an outlier is also called a novelty.

Outlier detection and novelty detection are both used for anomaly detection, where one is interested in detecting abnormal or unusual observations. Outlier detection is then also known as unsupervised anomaly detection and novelty detection as semi-supervised anomaly detection.

In the context of outlier detection, the outliers/anomalies cannot form a dense cluster as available estimators assume that the outliers/anomalies are located in low density regions. On the contrary, in the context of novelty detection, novelties/anomalies can form a dense cluster as long as they are in a low density region of the training data, considered as normal in this context.

</details>

<details>

<summary>Out of Distribution VS Anomaly</summary>

What is the difference between Out of Distribution and Anomaly Detection?

**Answer**

* Out of distribution (OOD) data refers to data that was collected at a different time, and possibly under different conditions or in a different environment than the data collected to create the model. It can be said that the data is from a different distribution.
* After the out of distribution data is collected, the model can perform either Novelty detection or Anomaly detection.
* Novelty data is the data that is in-distribution. Novelty detection checks whether the new data is in-distribution or not.
* Anomaly detection is used to test the data to see if it is different than what the model was trained on.

</details>

<details>

<summary>Resolution-Based Outlier Detection</summary>

What is a Resolution-Based Outlier Detection?

**Answer**

* The resolution-based outlier detection is an approach to address the problem of parameter value determination by measuring the outlierness of an observation $$p ∈ D$$ at different resolutions, and aggregating the results.
* In this algorithm, at the highest resolution, all observations are isolated points and thus considered to be outliers whereas at the lowest resolution all observations belong to one cluster and none is considered to be an outlier.
* As the resolution decreases from its highest value to the lowest value some observations in D begin to form clusters leaving other observations out of the clusters, and this phenomenon is captured in the resolution based outlier detection approach.

</details>

<details>

<summary>Change Detection</summary>

What is the Change Detection problem in Anomaly Detection?

**Answer**

* Change detection or change point detection tries to identify times when the probability distribution of a time series changes.
* Change detection is generally used to detect anomalous behavior.

Change point detection is great for the following cases:

* Detecting anomalous sequences/states in a time series.
* Detecting the average velocity of unique states in a time series.
* Detecting a sudden change in a time series state in real-time.

</details>

<details>

<summary>Drawbacks of Density Based Method</summary>

Can you tell some shortcomings of density based anomaly detection methods?

**Answer**

Density-based outlier detection method investigates the density of an object and that of its neighbors. Here, an object is identified as an outlier if its density is relatively much lower than that of its neighbors. Many real-world data sets demonstrate a more complex structure, where objects may be considered outliers with respect to their local neighborhoods, rather than with respect to the global data distribution.

</details>

<details>

<summary>Distance based Outlier detection</summary>

Explain Distance based Outlier detection methods?

**Answer**

A distance-based outlier detection method consults the neighborhood of an object, which is defined by a given radius. An object is then considered an outlier if its neighborhood does not have enough other points. This is termed as Distance-Based Outlier Detection Methods.

* Distance-Based Methods usually depend on a Multi-dimensional Index, Which is used to retrieve the neighborhood of each object to see if it contains sufficient points. If there are insufficient points, then the object is termed an outlier.
* Distance-Based methods scale better to multi-dimensional space and can be computed more efficiently than the statistical-based method. Identifying Distance-based outliers is an important and useful data mining activity. The main disadvantage of distance-based methods is that distance-based outlier detection is based on a single value of a custom parameter. This can cause significant problems if the dataset contains both dense and sparse regions.

Outlier detection methods can be categorized according to whether the sample of data for analysis is given with expert-provided labels that can be used to build an outlier detection model. In this case, the detection methods are supervised, semi-supervised, or unsupervised. Alternatively, outlier detection methods may be organized according to their assumptions regarding normal objects versus outliers. This categorization includes statistical methods, proximity-based methods, and clustering-based methods.

Algorithms For Mining Distance-Based Outliers:

* **Index-based algorithm:** The index-based algorithm facilitates multidimensional indexing structures, including R-trees or $$k-d$$ trees, to search for neighbors of each object $$o$$ inside radius $$d$$ around that object. Once $$K (K = N(1-p))$$ neighbors of object $$o$$ are discovered, it is accessible that $$o$$ is not an outlier. This algorithm has the lowest case complexity of $$O (k \* n^2)$$, where $$k$$ is the dimensionality, and $$n$$ is the number of objects in the data set.
* **Nested-loop algorithm:** The nested loop algorithm has the same evaluation complexity as the index-based algorithm but avoids building index structures and minimizes the amount of I/O. It splits the memory buffer in half and puts the data into several logical blocks.
* **Cell-based algorithm:** It avoids the $$O(n^2)$$ computational complexity and develops a cell-based algorithm for memory-resident datasets. Its complexity is $$O(c\*k + n)$$, where $$c$$ is a constant based on the number of cells and $$k$$ is the dimension.

</details>

<details>

<summary>Dictionary learning in Anamoly detection</summary>

Can you explain Dictionary learning in Anamoly detection?

**Answer**

Dictionary learning generalizes the assumption that “typical” data points inhabit a low-dimensional subspace of the ambient space by proposing that points may all lie in the union of many very low-dimensional subspaces. Specifically, using sparse coding techniques, dictionary learning represents each data point as a linear combination of only a few basis elements, ie. dictionary atoms. Since each data point is closely related to its corresponding atoms, points using the same atoms are presumably semantically related, naturally grouping the data. This suggests that anomalous data will exhibit at least one of three properties:

* Anomalous data will not be well represented by a learned dictionary as long as the dictionary is constrained to a sufficiently small number of atoms. Therefore, anomalies can be identified as having large residuals.
* The learned dictionary will be more influenced by anomalies than by data points that follow the greater trend, since anomalies will be much farther from the union of the spans of small subsets of the dictionary than from the span of the entire dictionary. That is, anomalous data will have high leverage on the model relative to “typical” data.
* If the dictionary contains a sufficiently large number of atoms or the anomalies occur frequently, these data points will fit the model well but will use “rare” basis atoms. These atoms are typically not used by regular data and are included in the learned model primarily to “accommodate” the anomalous data. Alternatively, typical basis elements may be used in atypical combinations to represent anomalous points.

</details>

<details>

<summary>Autoencoder in Anamoly Detection</summary>

How are Autoencoders used in Anamoly Detection?

**Answer**

One of the predominant use cases of the Autoencoder is anomaly detection. Think about cases like IoT devices, sensors in CPU, and memory devices which work very nicely as per functions. Still, when we collect their fault data, we have majority positive classes and significantly less percentage of minority class data, also known as imbalance data. Sometimes it is tough to label the data or expensive labelling the data, so we know the expected behaviour of data.

We pass Autoencoder with majority classes(normal data). The training objective is to minimize the reconstruction error, and the training objective is to minimize this. as training progresses, the model weights for the encoder and decoder are updated. The encoder is a downsampler, and the decoder is an upsampler. Encoder and decoder can be ANN, CNN, or LSTM neural network.

What AutoEncoder does? It learns the reconstruction function that works with normal data, and we can use this Model for anomaly detection. We get low reconstruction error for normal data and high for abnormal data(minority class).

</details>


# Big O

Big O notation also called Landau's symbol, is a symbolism used in complexity theory, computer science, and mathematics to describe the asymptotic behavior of functions. Basically, it tells you how fast a function grows or declines.

### Why is Algorithm Analysis Important?

To understand why algorithm analysis is important, we will take help of a simple example.

Suppose a manager gives a task to two of his employees to design an algorithm in Python that calculates the factorial of a number entered by the user.

The manager has to decide which algorithm to use. To do so, he has to find the complexity of the algorithm. One way to do so is by finding the time required to execute the algorithms.

In the Jupyter notebook, you can use the `%timeit` literal followed by the function call to find the time taken by the function to execute. Look at the following script:

```python
'''Alogrithm by Employee 1'''
def fact(n):
    product = 1
    '''Uses for loop'''
    for i in range(n):
        product = product * (i+1)
    return product

%timeit fact(50)

################################

'''Alogrithm by Employee 2'''
def fact(n):
    if n == 0:
        return 1
    else:
        '''Uses recursive call'''
        return n * fact(n-1)

%timeit fact(50)
```

**Output -**

* 4.16 µs ± 15 ns per loop (mean ± std. dev. of 7 runs, 100000 loops each)
* 7.41 µs ± 142 ns per loop (mean ± std. dev. of 7 runs, 100000 loops each)

The execution time shows that the first algorithm is faster compared to the second algorithm involving recursion. This example shows the importance of algorithm analysis. In the case of large inputs, the performance difference can become more significant. However, **execution time is not a good metric to measure the complexity of an algorithm since it depends upon the hardware. A more objective complexity analysis metrics for the algorithms is needed.** This is where Big O notation comes to play.

### Algorithm Analysis with Big-O Notation

Big-O notation signifies the relationship between the input to the algorithm and the steps required to execute the algorithm. It is denoted by a big $$O$$ followed by opening and closing parenthesis. Inside the parenthesis, the relationship between the input and the steps taken by the algorithm is presented using $$n$$.

For instance, if there is a linear relationship between the input and the step taken by the algorithm to complete its execution, the Big-O notation used will be $$O(n)$$. Similarly, the Big-O notation for quadratic functions is $$O(n^2)$$

The following are some of the most common Big-O functions:

<table data-header-hidden><thead><tr><th width="208"></th><th></th></tr></thead><tbody><tr><td><strong>Name</strong></td><td><strong>Big O</strong></td></tr><tr><td>Constant</td><td><span class="math">O(c)</span></td></tr><tr><td>Linear</td><td><span class="math">O(n)</span></td></tr><tr><td>Quadratic</td><td><span class="math">O(n^2)</span></td></tr><tr><td>Cubic</td><td><span class="math">O(n^3)</span></td></tr><tr><td>Exponential</td><td><span class="math">O(2^n)</span></td></tr><tr><td>Logarithmic</td><td><span class="math">O(log(n))</span></td></tr><tr><td>Log Linear</td><td><span class="math">O(nlog(n))</span></td></tr></tbody></table>

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FIvMyGJTiA0cdtDretIHG%2Fimage16.png?alt=media&amp;token=cb99e65d-a29b-4bb8-b9c3-8fc9b90d532d" alt=""><figcaption><p>n is the input size and c is a positive constant</p></figcaption></figure>

### Analogy

Imagine the following scenario: *You've got a file on a hard drive, and you need to send it to your friend who lives across the country. You need to get the file to your friend as fast as possible. How should you send it?*

Most people's first thought would be email, FTP, or some other means of electronic transfer. That thought is reasonable, but only half correct. If it's a small file, you're certainly right. It would take 5 - 10 hours to get to an airport, hop on a flight, and then deliver it to your friend. But what if the file were really, really large? Is it possible that it's faster to physically deliver it via plane?

Yes, actually it is. A one-terabyte (1 TB) file could take more than a day to transfer electronically. It would be much faster to just fly it across the country. If your file is that urgent (and cost isn't an issue), you might just want to do that. What if there were no flights, and instead you had to drive across the country? Even then, for a really huge file, it would be faster to drive.

#### Time Complexity

This is what the concept of asymptotic runtime, or big $$O$$ time, means. We could describe the data transfer *algorithm* runtime as:

* Electronic Transfer: $$O(s)$$, where $$s$$ is the size of the file. This means that the time to transfer the file increases linearly with the size of the file. (Yes, this is a bit of a simplification, but that's okay for these purposes)
* Airplane Transfer: $$O(1)$$ with respect to the size of the file. As the size of the file increases, it won't take any longer to get the file to your friend. The time is constant.

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FFvt4t8DuZJYFHuqtPdMF%2Fimage17.PNG?alt=media&amp;token=824dd01d-5b91-4100-a97b-9708644aa96c" alt=""><figcaption><p>No matter how big the constant is and how slow the linear increase is, linear will at some point surpass the constant.</p></figcaption></figure>

There are many more runtimes than this. Some of the most common ones are $$O(log N),O(N log N), O(N), O(N^2), O(2^N)$$. There's no fixed list of possible runtimes, though. You can also have multiple variables in your runtime. For example, the time to paint a fence that's $$w$$ meters wide and $$h$$ meters high could be described as $$O(wh)$$. If you needed $$p$$ layers of paint, then you could say that the time is $$O(whp)$$.

#### Big O of DS Algorithms

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2Fh1c9uWeWwSixZ41glvx9%2Fimage15.png?alt=media&amp;token=d439fcd8-a54f-417c-9a07-65af94e238ed" alt=""><figcaption><p>Big O of some of the popular Machine Learning Algorithms</p></figcaption></figure>


# Neural Network

{% hint style="success" %}
[This is an excellent resource available for free on Neural Networks](http://neuralnetworksanddeeplearning.com/chap1.html)
{% endhint %}

Perceptrons were one of the earliest proposed models for learning simple classification tasks, which later became the fundamental building block of artificial neural networks.

## Artificial Neuron

An artificial neuron is very similar to a perceptron, except that the activation function is not a step function.

Like a perceptron, the input is the weighted sum of inputs. The output is the activation function applied on the input. The activation function could be any function, though it should have some important properties such as:

* Activation functions should be smooth i.e. they should have no abrupt changes when plotted
* They should also make the inputs and outputs non-linear with respect to each other to some extent. This is because non-linearity helps in making neural networks more compact

Some of the common activation functions are as follows:

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FbeHnvfkWEbK5ILFDGnbF%2Fimage1.png?alt=media&amp;token=af016620-ca76-4cfb-83cf-808f33248f01" alt=""><figcaption><p><a href="https://www.v7labs.com/blog/neural-networks-activation-functions">Source</a></p></figcaption></figure>

**An artificial neural network is a network of such neurons.** Neurons in a neural network are arranged in layers. The *first* and the *last* layer are called the *input* and *output* layers. Input layers have as many neurons as the number of attributes in the data set and the output layer has as many neurons as the number of classes of the target variable (for a classification problem). For a regression problem, the number of neurons in the output layer would be 1 (a numeric variable). There are six main things that need to be specified for specifying a neural network completely:

* Network Topology
* Input Layer
* Output Layer
* Weights
* Activation functions
* Biases

An important thing to note is that the inputs can only be numeric. For different types of input data, we use different ways to convert the inputs to a numeric form. In case of text data, we either use a one-hot vector or word embeddings corresponding to a certain word. Feeding images (or videos) is straightforward since images are naturally represented as arrays of numbers.

**The output layer used in case of multiclass classification problem is the softmax layer.** A softmax output is a multiclass logistic function commonly used to compute the 'probability' of an input belonging to one of the multiple classes.

Since large neural networks can potentially have extremely complex structures, certain assumptions are made to simplify the way information flows in them:

* Neurons are arranged in layers and the layers are arranged sequentially
* Neurons within the same layer do not interact with each other
* All the inputs enter the network through the input layer and all the outputs go out of the network through the output layer
* Neurons in consecutive layers are densely connected, i.e. all neurons in layer $$l$$ are connected to all neurons in layer $$l+1$$
* Every interconnection in the neural network has a weight associated with it, and every neuron has a bias associated with it
* All neurons in a particular layer use the same activation function

Neural networks require rigorous training. Recall that models such as linear regression, logistic regression, SVMs etc. are trained on their coefficients, i.e. the training task is to find the optimal values of the coefficients to minimize some cost function. Neural networks are no different - they are trained on weights and biases.

During training, the neural network learning algorithm fits various models to the training data and selects the best model for prediction. The learning algorithm is trained with a fixed set of hyperparameters - the network structure (number of layers, number of neurons in the input, hidden and output layers etc.). It is trained on the weights and the biases, which are the parameters of the network.

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FXrysYauGQGw3HkzdwcHh%2Fimage2.gif?alt=media&amp;token=5e159f4a-9d0b-40f3-8c06-642e5632ceac" alt=""><figcaption><p>Neural Networks can be used to <em>reasonably</em> approximate <em>most</em> functions</p></figcaption></figure>

## Feed forward

In artificial neural networks, the output from one layer is used as input to the next layer. Such networks are called feedforward neural networks. Feed forward neural network architecture consists of following main parts – Input Layer, Hidden Layer - If the number of hidden layers is one then it is known as a shallow neural network, else it is known as a deep neural network, Output Layer

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FhbLk5LDG8xJldVfSrIZp%2Fimage3.gif?alt=media&amp;token=73c18ff2-cbe8-417c-89ab-15e0d2eed8db" alt=""><figcaption><p><a href="https://machinelearningknowledge.ai/animated-explanation-of-feed-forward-neural-network-architecture/">Source</a></p></figcaption></figure>

## Backpropagation

In the neural network training task, the goal is to compute the optimal weights and biases by minimizing some cost function. The task of training neural networks is exactly the same as that of other ML models such as linear regression, SVMs etc. The desired output (output from the last layer) minus the actual output is the cost (or the loss), and we to tune the parameters $$w$$ and $$b$$ such that the total cost is minimized.

An important point to note is that if the data is large (which is often the case), loss calculation itself can get pretty messy. For example, if you have a million data points, they will be fed into the network (in batches), the output will be calculated using feedforward and the loss/cost $$L\_i$$ (for $$i^{th}$$ data point) will be calculated. The total loss is the sum of losses of all the individual data points. We minimize the average of the total loss and the not the total loss. Minimizing the average loss implies that the total loss is getting minimized.

For a large neural network, the number of weight elements and biases becomes so large and minimizing the loss with so many parameters is a difficult task. This complex task is achieved using gradient descent.

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2Fk39L0KFiaNjbPlS01uRE%2Fimage4.gif?alt=media&amp;token=cadef91c-76e2-4092-b379-9526b46740f8" alt=""><figcaption><p><a href="https://machinelearningknowledge.ai/animated-explanation-of-feed-forward-neural-network-architecture/">Source</a></p></figcaption></figure>

For updating weights and biases using plain backpropagation, you have to scan through the entire data set to make a single update to the weights. This is computationally very expensive for large datasets. Thus, you use multiple batches (or **mini-batches**) of data points, compute the average gradient for a batch, and update the weights based on that gradient.

But there is a danger in doing this - you are making weight updates based only on gradients computed for small batches, not the entire training set. Thus, you make multiple passes through the entire training set using epochs. An epoch is one pass through the entire training set, and you use multiple epochs (typically 10, 20, 50, 100 etc.) while training. In each epoch, you reshuffle all the data points, divide the reshuffled set into m batches, and update weights based on gradient of each batch.

### Stochastic Gradient Descent (SGD)

In most libraries such as TensorFlow, the SGD training procedure is as follows:

* You specify the number of epochs (typical values are 10, 20, 50, 100 etc.) - more epochs require more computational power
* You specify the number of batches $$m$$ (typical values are 32, 64, 128, etc.)
* At the start of each epoch, the data set is reshuffled and divided into $$m$$ batches
* The average gradient of each batch is then used to make a weight update
* The training is complete at the end of all the epochs

Apart from being computationally faster, the SGD training process has another big advantage - it actually helps you reach the global minima (instead of being stuck at a local minima). Also, to avoid the problem of getting stuck at a local optimum, you need to strike a balance between exploration and exploitation. Exploration means that you try to minimize the loss function with different starting points of $$W$$ and $$b$$, i.e., you initialize $$W$$ and $$b$$ with different values. On the other hand, exploitation means that you try to reach the global minima starting from a particular  $$W$$ and $$b$$ and do not explore the terrain at all. That might lead you to the lowest point locally, but not the necessarily the global minimum.

## Modifications in Neural Networks

Neural networks are usually large, complex models with tens of thousands of parameters, and thus tend to overfit the training data. As with many other ML models, regularization is a common technique used in neural networks to address this problem.

The *parameter norm* regularization is similar to that in linear regression in almost every aspect. As in lasso regression (L1 norm), we get a sparse weight matrix, which is not the case with the L2 norm. Despite this fact, the L2 norm is more common because the sum of the squares term is easily differentiable which comes in handy during backpropagation. Apart from using the parameter norm, there is another popular neural network regularization technique called **dropouts**.

### Dropouts

Some important points to note regarding dropouts are:

* Dropouts can be applied only to some layers of the network (in fact, that is a common practice - you choose some layer arbitrarily to apply dropouts to)
* The mask $$\alpha$$ is generated independently for each layer during feedforward, and the same mask is used in backpropagation.
* The mask changes with each minibatch/iteration, are randomly generated in each iteration (sampled from a Bernoulli with some $$p(1)=q$$)

Why the dropout strategy works well is explained through the notion of a manifold. Manifold captures the observation that in high dimensional spaces, the data points often actually lie in a lower-dimensional manifold. The dropout strategy uses this fact to find a lower-dimensional solution to the problem.

Dropouts help in symmetry breaking as well. There is every possibility of the creation of communities within neurons which restricts them from learning independently. Hence, by setting some random set of the weights to zero in every iteration, this community/symmetry is broken. Note that there, a different mini batch is processed in every iteration in an epoch, and dropouts are applied to each mini batch.

### Problems

<details>

<summary>Vanishing and Exploding Gradient</summary>

Explain the vanishing and exploding gradient problem in neural networks.

**Answer**

The vanishing and exploding gradient problems are issues that can occur during the training of deep neural networks, particularly in deep feedforward neural networks and recurrent neural networks (RNNs). These problems are related to the way gradients (derivatives of the loss function with respect to the model parameters) are propagated backward through the layers of a network during the training process using gradient descent or its variants.

1. **Vanishing Gradient Problem:**
   * The vanishing gradient problem occurs when the gradients of the loss function with respect to the model parameters become very small as they are propagated backward through the layers of a deep network.
   * This is problematic because small gradients lead to very slow weight updates during training, which can result in the network learning very slowly or not learning at all.
   * It typically occurs in networks with many layers, especially when activation functions like sigmoid or hyperbolic tangent (tanh) are used.
   * These activation functions squash their inputs into a small range, and the derivatives of these functions become very small for large or very small inputs. As gradients are backpropagated, these small gradients can compound, making earlier layers learn very slowly.
2. **Exploding Gradient Problem:**
   * The exploding gradient problem is the opposite of the vanishing gradient problem. It occurs when gradients become extremely large as they are propagated backward through the layers of a deep network.
   * When gradients are too large, they can lead to numerical instability during training, causing the model's weights to become NaN (Not-a-Number) or saturate, which can hinder convergence.
   * The exploding gradient problem can happen in networks with certain types of architectures, large learning rates, or poorly conditioned loss functions.

Solutions to these problems include:

**1. Weight Initialization:** Properly initializing the weights of neural networks can help alleviate the vanishing and exploding gradient problems. Techniques like He initialization or Xavier initialization set initial weights to values that help maintain gradient magnitudes.

**2. Activation Functions:** Using activation functions that have gradients that neither vanish nor explode across a wide range of inputs can help. Rectified Linear Units (ReLUs) are popular because they don't saturate for positive inputs.

**3. Batch Normalization:** Batch normalization normalizes the inputs to each layer, reducing internal covariate shift and helping gradients flow more consistently during training.

**4. Gradient Clipping:** This technique limits the magnitude of gradients during training to prevent them from becoming too large and causing the exploding gradient problem.

**5. Architecture Design:** Using skip connections or residual connections in deep networks can help gradients flow more easily through the network.

**6. Learning Rate Scheduling:** Adjusting the learning rate during training, such as using learning rate annealing or adaptive learning rate methods like Adam, can mitigate both problems.

These techniques collectively help mitigate the vanishing and exploding gradient problems and allow for the successful training of deep neural networks. The choice of which technique to use often depends on the specific network architecture and problem domain.

</details>

<details>

<summary>Perceptron</summary>

What is a perceptron?

**Answer**

* A **Perceptron** is a fundamental unit of a Neural Network that is also a single-layer Neural Network.
* Perceptron is a linear *classifier*. Since it uses already labeled data points, it is a *supervised learning algorithm*.
* The *activation function* applies a step rule (convert the numerical output into +1 or -1) to check if the output of the weighting function is greater than zero or not.

A **Perceptron** is shown in the figure below:

<img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FblnCynnO6nhaKb3fw98z%2Fimage.png?alt=media&amp;token=fc496a78-908e-4014-8d87-72d00639e9e0" alt="" data-size="original">

</details>

<details>

<summary>Classify email into spam or ham</summary>

How many neurons needed to classify and email spam vs ham?

**Answer**

To classify email into spam or ham, you just need one neuron in the output layer of a neural network. For example, to indicate the probability that the email is spam, you would typically use the **logistic activation** function in the output layer when estimating a probability.

</details>

<details>

<summary>Role of Activation function</summary>

Can you explain the role of Activation function in NN?

**Answer**

* **Activation Functions** help in keeping the value of the output from the neuron restricted to a certain limit as per the requirement. If the limit is not set then the output will reach very high magnitudes. Most activation functions convert the output to `-1` to `1` or to `0` to `1`.
* The most *important* role of the activation function is the ability to add **non-linearity** to the neural network. Most of the models in real-life is non-linear so the activation functions help to create a non-linear model.
* The *activation function* is responsible for deciding whether a neuron should be activated or not.

</details>

<details>

<summary>Detect if a model has Vanishing Gradient Problem</summary>

How do you tell if your model has a Vanishing Gradient Problem?

**Answer**

* The model will improve *very slowly* during the training phase and it is also possible that training stops *very early*, meaning that any further training does not improve the model.
* The weights closer to the output layer of the model would witness more of a change whereas the layers that occur closer to the input layer would not change much (if at all).
* Model weights *shrink exponentially* and become *very small* when training the model.
* The model weights become `0` in the training phase.

</details>

<details>

<summary>Detect if model has an Exploding Gradient Problem?</summary>

How do you tell if your model has an Exploding Gradient Problem?

**Answer**

There are some subtle signs that you may be *suffering from exploding gradients* during the training of your network, such as:

* The model is unable to get traction on your training data (e g. *poor loss*).
* The model is *unstable*, resulting in large changes in loss from update to update.
* The model loss goes to `NaN` during training.

If you have these types of problems, you can dig deeper to see if you have a problem with exploding gradients. There are some less subtle signs that you can use to confirm that you have exploding gradients:

* The model weights quickly become very large during training.
* The model weights go to `NaN` values during training.
* The error gradient values are consistently above `1.0` for each node and layer during training.

</details>

<details>

<summary>Importance of initialization</summary>

Why is the initialization process important in neural network and explain some common initialization techniques?

**Answer**

The initialization process in neural networks is crucial because it sets the initial values for the model's parameters (weights and biases). Proper initialization can significantly impact the training process and the performance of the neural network. Here's why initialization is important:

1. **Avoiding Symmetry**: In neural networks, each neuron's parameters should have different initial values to break symmetry. If all neurons start with the same weights, they will all learn the same features, making the network no better than a single neuron. Proper initialization helps break this symmetry, allowing neurons to specialize in different features.
2. **Preventing Vanishing and Exploding Gradients**: Initializing weights properly can help mitigate the vanishing and exploding gradient problems, which can affect the stability and convergence of training.
3. **Faster Convergence**: Proper initialization can lead to faster convergence during training. A well-initialized network often requires fewer epochs to achieve good performance.

Here are some common weight initialization techniques used in neural networks:

1. **Zero Initialization**: Initializing all weights to zero is generally not recommended because it results in symmetric weights, causing all neurons to compute the same output. However, biases can be initialized to zero without much concern.
2. **Random Initialization**: This is the most common initialization technique. Weights are initialized with small random values. Some variations of random initialization include:
   * **Random Uniform Initialization**: Weights are sampled from a uniform distribution within a small range, such as \[-0.1, 0.1].
   * **Random Normal Initialization**: Weights are sampled from a normal distribution with a mean of 0 and a small standard deviation.
3. **Xavier/Glorot Initialization**: This initialization method is designed to work well with activation functions like the hyperbolic tangent (tanh) and the logistic sigmoid. It scales the initial weights based on the number of input and output units of a layer to maintain a consistent variance of activations throughout the network. The formula for Xavier initialization is:
   * For a layer with `n_in` input units and `n_out` output units:
     * Initialize weights from a random distribution with mean 0 and variance `2 / (n_in + n_out)`.
4. **He Initialization**: This initialization method is suitable for networks that use Rectified Linear Units (ReLUs) as activation functions. It scales the initial weights based on the number of input units to maintain a consistent variance of activations. The formula for He initialization is:
   * For a layer with `n_in` input units:
     * Initialize weights from a random distribution with mean 0 and variance `2 / n_in`.
5. **LeCun Initialization**: This initialization method is designed for networks using the Leaky ReLU activation function. It scales the initial weights based on the number of input units to maintain a consistent variance of activations. The formula for LeCun initialization is:
   * For a layer with `n_in` input units:
     * Initialize weights from a random distribution with mean 0 and variance `1 / n_in`.

Choosing the right initialization method depends on the specific neural network architecture, the choice of activation functions, and the problem at hand. Proper initialization can improve the network's training stability and convergence, making it an essential component of building effective neural networks.

</details>


# Recurrent Neural Network

{% hint style="warning" %}
This page is a Work In Progress
{% endhint %}

Normal neural network is insufficient to train sequence data. Some examples of sequence data are:

* Time series
* Music
* Videos
* Text

Sequential data contains multiple entities​ and the order in which these entities are present is important.

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FqrehdlzlpbSJhGXnjCeB%2Fimage2.png?alt=media&amp;token=0863a143-470c-40d7-ba5b-a2964506b023" alt=""><figcaption><p>In RNN each activation is dependent on two things: the activation in the previous layer <span class="math">l-1</span> at the current timestep <span class="math">t</span>, and the activation in the same layer <span class="math">l</span> at the previous timestep <span class="math">t-1</span></p></figcaption></figure>

## Types of RNN

* **Many to One:** This architecture involves a sequence as an input and a single entity as an output​. E.g.: next word prediction
* **Many to Many:** Model data which involves sequences in the input as well as the output​. Both the input and output sequences must have a one-to-one correspondence ​and therefore the input and output sequences are equal in length​. E.g.: POS tagging of sentences
* **Encoder-decoder:** Here the length of the input and the output sequence is not equal​. It is used in problems such as language translation and document summarization. Here the errors are backpropagated from the decoder to the encoder. The encoder and decoder have a different set of weights and they are different RNNs altogether. The loss is calculated at each timestep which can either be backpropagated at each timestep, or the cumulative loss (sum of all the losses from all the timesteps of a sequence) can be backpropagated after the entire sequence is ingested. Generally, the errors are backpropagated once an RNN ingests an entire batch

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FKc9eUrO7pXVGtCXe1OWM%2Fimage3.png?alt=media&amp;token=16fc5b32-a998-43cd-b017-251c7c261941" alt=""><figcaption></figcaption></figure>

* **One to Many:** This type of architecture has a single entity as an input and a sequence as the output​. It is used for generation such as music generation, creating drawings, generating text, etc.

## Backpropagation Through Time (BPTT)

Any given term in an RNN depends not only on the current input but also on the input from previous timesteps​. The gradients not only flow back from the output layer to the input layer, but they also flow back in time from the last timestep to the first timestep. Hence the name backpropagation through time.

## Exploding and Vanishing gradient:

* **Vanishing:** The vanishing gradient problem describes a situation encountered in the training of neural networks where the gradients used to update the weights shrink exponentially. As a consequence, the weights are not updated anymore, and learning stalls.
* **Exploding:** The exploding gradient problem describes a situation in the training of neural networks where the gradients used to update the weights grow exponentially. This prevents the backpropagation algorithm from making reasonable updates to the weights, and learning becomes unstable.

### How to understand

* **Vanishing:**
  * The parameters of the higher layers change significantly whereas the parameters of lower layers would not change much (or not at all)
  * The model weights may become 0 during training
  * The model learns very slowly and perhaps the training stagnates at a very early stage just after a few iterations
* **Exploding:**
  * There is an exponential growth in the model parameters
  * The model weights may become NaN during training
  * The model experiences avalanche learning

### How to rectify

* Use the ReLU Activation Function over the sigmoid activation function as it is prone to creating vanishing gradients, especially when several of them are chained together. This is due to the fact that the sigmoid function saturates towards 0 for large negative or towards 1 for large positive values.
* Another way to address vanishing and exploding gradients is through weight initialization techniques. In a neural network, we initialize the weights randomly. Certain techniques such as He initialization and Xavier initialization ensure that the weights are close to 1.
* Another simple option is gradient clipping. This way, you just define an interval within which you expected the gradients to fall. If the gradients exceed the permissible maximum, you automatically set them to the maximum upper bound of your interval. Similarly, if they fall below the permissible minimum, you automatically set them to the lower bound.
* To get rid of the vanishing gradients problem, the researchers came up with another type of cell that can be used inside an RNN layer, called the LSTM cell. We will take a look into that later.

## Bidirectional RNNs

There are two kinds of problems in sequences:

* **Offline sequences​:** These allow for a lookahead in time. The entire sequence is present for your perusal. For example, a document present in your local drive where you have access to the entire document.
* **Online sequences​:** These don’t allow for a lookahead. For example, a person is speaking and you need to transcribe the speech. In this case, you don’t know what is going to come next.

You can make use of offline sequences by looking ahead. You can feed the offline sequences to an RNN in regular order as well as the reverse order to get better results in whatever task you’re doing. Such an RNN is called a bidirectional RNN​. In a bidirectional RNN, the input at each timestep is a concatenation of the entity present in regular order and the entity present in reverse order. For example, for a sentence of length $100$, the input at the first timestep will be a concatenation of the first word $x\_1$ and the last word $x\_{100}$. A bidirectional RNN has $2$x number of parameters​ than a vanilla RNN.


# Lexical Processing

In any large text document (say, a hundred thousand words), the word frequencies follow Zipf distribution:

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2Fgit-blob-f72516f67294f6cd37e7b993e4e7cabffc0d317d%2Fimage1.png?alt=media" alt=""><figcaption><p>Zipf distribution</p></figcaption></figure>

* Remove **STOP words** as they have high frequency and is most cases does not provide anything important
* **Tokenize** or break the corpus by words, sentences, etc.
* **Canonicalisation** or Reduce the words into their base forms:

  * **Stemming:** a rule-based technique that just chops off the suffix of a word to get its root form which is called the ‘stem’. For example converts driving, drive etc. to driv. But is not good for words like feet, drove etc.
  * **Lemmatization:** a more sophisticated technique, it doesn’t just chop off the suffix of a word. Instead, it takes an input word and searches for its base word by going recursively through all the variations of dictionary words. The base word, in this case, is called the ‘lemma’. It is more resource intensive and you need to pass the POS tag of the word along with the word to be lemmatized.
  * **Phonetic Hashing:** certain words which have different pronunciations in different languages. As a result, they end up being spelt differently. Example being Delhi and Dilli

  <figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2Fgit-blob-90e24099a8a29f3392b10c8f3ccbe39faba0f108%2Fimage2%20(1).png?alt=media" alt=""><figcaption><p>Soundex Algorithm for Phonetic Hashing (<a href="https://www.sqlservercentral.com/articles/soundex-experiments-with-sqlclr-part-2">Source</a>)</p></figcaption></figure>

  * **Edit Distance:** Edit distance is a way of quantifying how dissimilar two strings are to one another by counting the minimum number of operations required to transform one string into the other. There are different methodologies for doing it, e.g. Hamming distance,Levenshtein distance, Jaro-Winkler etc.
  * **Pointwise Mutual Information(PMI):** Words like "Massachusetts Institute of Technology" are essentially one word but tokenization reduces these into individual words which is not desireable. PMI is used to determine whether this term should be represented by a single token or not.
* Convert the data into tabular form:
  * **Bag of Words(BoW):** A table containing which word is present in which document, can be either count or binary
  * **tf-idf:** is a numerical statistic that is intended to reflect how important a word is to a document in a collection or corpus. The tf–idf value increases proportionally to the number of times a word appears in the document and is offset by the number of documents in the corpus that contain the word, which helps to adjust for the fact that some words appear more frequently in general. $$tf-idf = \frac{\text{freq of term 't' in doc 'd'}}{\text{total terms in 'd'}} \* log \frac{\text{total number of docs}}{\text{total number of docs having term 't'}}$$


# Syntactic Processing

Syntactic processing is widely used in applications such as question answering systems, information extraction, sentiment analysis, grammar checking etc. There are 3 broad levels of syntactic processing:

* (Parts-of-Speech) POS tagging
* Constituency parsing
* Dependency parsing

POS tagging is a crucial task in syntactic processing and is used as a preprocessing step in many NLP applications

## POS tagging

The four main techniques used for POS tagging:

* **Lexicon-based** approach uses the following simple statistical algorithm: for each word, it assigns the POS tag that most frequently occurs for that word in some training corpus. For example, it will assign the tag "verb" to any occurrence of the word "run" if "run" is used as a verb more often than any other tag.
* **Rule-based** taggers first assign the tag using the lexicon method and then apply predefined rules. Some examples of rules are: Change the tag to VBG for words ending with ‘-ing’, Changes the tag to VBD for words ending with ‘-ed’, etc.
* **Probabilistic (or stochastic)** techniques don't naively assign the highest frequency tag to each word, instead, they look at slightly longer parts of the sequence and often use the tag(s) and the word(s) appearing before the target word to be tagged. The commonly used probabilistic algorithm for POS tagging is Hidden Markov Model (HMM)
* **Deep-learning based** POS tagging: Recurrent Neural Networks (RNNs) are used for sequential modeling processes

### Hidden Markov Models

Markov processes are commonly used to model sequential data, such as text and speech. The first-order Markov assumption states that the probability of an event (or state) depends only on the previous state. The Hidden Markov Model, an extension to the Markov process, is used to model phenomena where the states are hidden, and they emit observations. The transition and the emission probabilities specify the probabilities of transition between states and emission of observations from states, respectively. In POS tagging, the states are the POS tags while the words are the observations. To summarise, a Hidden Markov Model is defined by the initial state, emission, and the transition probabilities.

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2Fgit-blob-08847dde48e4a625a92369e054fc7459885fe7ea%2Fimage3.png?alt=media" alt=""><figcaption></figcaption></figure>

The POS tag $$T\_i$$ for given word $$W\_i$$ depends on two things: POS tag of the previous word and the word itself.

$$P(T\_i| W\_i) = P(W\_i|T\_i) \* P(T\_i-1|T\_i)$$

So, the probability of a tag sequence $$(T\_1, T\_2, T\_3,$$ …$$, T\_n)$$ for a given the word sequence $$(W\_1, W\_2, W\_3,$$ …$$, W\_n)$$ can be defined as:

$$P(T|W) = (P(W\_1|T\_1) \* P(T\_1|start)) \* (P(W\_2|T\_2) \* P(T\_2|T\_1)) \* ...\* (P(W\_n|T\_n) \* P(T\_n|T\_{n-1}))$$

**For a sequence of** $$n$$ **words and** $$t$$ **tags, a total of** $$t\_n$$ **tag sequences are possible.**

### Viterbi Heuristic

Viterbi Heuristic can deal with this problem by taking a greedy approach. The basic idea of the Viterbi algorithm is as follows - given a list of observations (words) $$O\_1,O\_2....O\_n$$ to be tagged, rather than computing the probabilities of all possible tag sequences, you assign tags sequentially, i.e. assign the most likely tag to each word using the previous tag.

More formally, you assign the tag $$T\_i$$ to each word $$W\_i$$ such that it maximises the likelihood:

$$P(T\_i| W\_i) = P(W\_i|T\_i) \* P(T\_i-1|T\_i)$$

where $$T\_i-1$$ is the tag assigned to the previous word. The probability of a tag $$T\_i$$ is assumed to be dependent only on the previous tag $$T\_{i-1}$$, and hence the term $$P(T\_i|T\_{i-1})$$ - Markov Assumption.

**Viterbi algorithm is an example of a dynamic programming algorithm.** In general, algorithms which break down a complex problem into subproblems and solve each subproblem optimally are called dynamic programming algorithms.

### Learning HMM Model Parameters

The process of learning the probabilities from a tagged corpus is called **training an HMM model**. The emission and the transition probabilities can be learnt as follows:

* **Emission Probability** of a word $$w$$ for tag $$t$$: $$P(w|t)$$ = Number of times $$w$$ has been tagged $$t$$/Number of times $$t$$ appears

  Example: $$P(dog|N)$$ = Number of times 'dog' appears as Noun/ Number of times Noun is appearing
* **Transition Probability** of tag $$t\_1$$ followed by tag $$t\_2$$: $$P(t\_2|t\_1)$$ = Number of times $$t\_1$$ is followed by tag $$t\_2$$/ Number of times $$t\_1$$ appears

  Example: $$P(Noun|Adj)$$ = number of times adjective is followed by Noun/ Number of times Adjective is appearing

## Constituency Parsing

shallow parsing is not sufficient. Shallow parsing, as the name suggests, refers to fairly shallow levels of parsing such as POS tagging, chunking, etc. But such techniques would not be able to check the grammatical structure of the sentence, i.e. whether a sentence is grammatically correct, or understand the dependencies between words in a sentence.

Two most commonly used paradigms of parsing - constituency parsing and dependency parsing, which would help to check the grammatical structure of the sentence.

In constituency parsing, you learnt the basic idea of constituents as grammatically meaningful groups of words, or phrases, such as noun phrase, verb phrase etc. You also learnt the idea of context-free grammars or CFGs which specify a set of production rules. Any production rule can be written as A -> B C, where A is a non-terminal symbol (NP, VP, N etc.) and B and C are either non-terminals or terminal symbols (i.e. words in vocabulary such as flight, man etc.).

Example a CFG is:

* S -> NP VP
* NP -> DT N| N| N PP
* VP -> V| V NP
* N -> ‘woman’| ‘bear’
* V -> ‘ate’
* DT -> ‘the’| ‘a’

There are two broad approaches to constituency parsing:

* **Top-down parsing:** starts with the start symbol $$S$$ at the top and uses the production rules to parse each word one by one. And, you continue to parse until all the words have been allocated to some production rule.

Top-down parsers have a specific limitation- Left Recursion.

Example of a left recursion: VP -> VP NP. Whenever a top-down parser encounters such a rule, it runs into an infinite loop, thus no parse tree is obtained. Following is the illustration of top-down parse:

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2Fgit-blob-60b0573dffc243913620eb232daf7f597a4ae7ca%2Fimage4.png?alt=media" alt=""><figcaption><p>Top-down parse</p></figcaption></figure>

* **Bottom-up parsing:** reduces each terminal word to a production rule, i.e. reduces the right-hand-side of the grammar to the left-hand-side. It continues the reduction process until the entire sentence has been reduced to the start symbol S. Shift-Reduce Parser algorithm, which parses the words of the sentence one-by-one either by shifting a word to the stack or reducing the stack by using the production rules. Below is an example of bottom-up parse tree.

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2Fgit-blob-c7c8a52acb2a2b7a357619dd95afe255a357f97c%2Fimage5.png?alt=media" alt=""><figcaption><p>Bottom-up parse</p></figcaption></figure>

### Probabilistic CFG

Since natural languages are inherently ambiguous, there are often cases where multiple parse trees are possible. In such cases, we need a way to make the algorithms figure out the most likely parse tree. Probabilistic Context-Free Grammars (PCFGs) are used when we want to find the most probable parsed structure of the sentence. PCFGs are grammar rules, similar to what you have seen before, along with probabilities associated with each production rule. An example production rule is as follows:

NP -> Det N (0.5) | N (0.3) |N PP (0.2)

It means that the probability of an NP breaking down to a ‘Det N’ is 0.50, to an 'N' is 0.30 and to an ‘N PP’ is 0.20. Note that the sum of probabilities is 1.00. Overall probability for a parsed structure of the sentence is probabilities of all rules used in that parsed structure. The parsed tree with maximum probability is best possible interpretation of the sentence.

### Chomsky Normal Form

The Chomsky Normal Form (CNF), proposed by the linguist Noam Chomsky, is a normalized version of the CFG with a standard set of rules defining how production rule must be written. The three forms of CNF rules can be written:

* A -> B C
* A -> a
* S -> ε

A, B, C are non-terminals (POS tags), a is a terminal (term), S is the start symbol of the grammar and ε is the null string. The table below shows some examples for converting CFGs to the CNF:

\| CFG | VP -> V NP PP | VP -> V | | CNF | VP -> V (NP1) | VP -> V (VP1) | | | NP1 -> NP PP | VP1 -> ε |

## Dependency Parsing

In dependency grammar, constituencies (such as NP, VP etc.) do not form the basic elements of grammar, but rather dependencies are established between the words themselves.

Free word order languages such as Hindi, Bengali are difficult to parse using constituency parsing techniques. This is because, in such free-word-order languages, the order of words/constituents may change significantly while keeping the meaning exactly the same. It is thus difficult to fit the sentences into the finite set of production rules that CFGs offer. Dependencies in a sentence are defined using the elements Subject-Verb-Object (SVO). The following table shows SVO dependencies in three types of sentences - declarative, interrogative, and imperative

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2Fgit-blob-dc2fefba72e3a593d851d45affadc9011f1fff35%2Fimage6.png?alt=media" alt=""><figcaption></figcaption></figure>

Apart from dependencies defined in the form of subject-verb-object, there's a non-exhaustive list of dependency relationships, which are called **universal dependencies**.

Dependencies are represented as labelled arcs of the form $$h → d(l)$$ where '$$h$$' is called the “head” of the dependency, '$$d$$' is the “dependent” and $$l$$ is the “label” assigned to the arc. In a dependency parse, we start from the root of the sentence, which is often a verb. And then start to establish dependencies between root and other words.

## Information Extraction

Information Extraction (IE) system can extract entities relevant for booking flights (such as source and destination cities, time, date, budget constraints etc.) in a structured format from unstructured user-generated input. IE is used in many applications such as chatbots, extracting information from websites, etc.

A generic pipeline for Information Extraction is as follows:

* **Preprocessing:**
  * Sentence Tokenization: sequence segmentation of text
  * Word Tokenization: breaks down sentences into tokens
  * POS tagging: assigning Parts of Speech tags to the tokens. The POS tags can be helpful in defining what words could form an entity
* **Entity Recognition:**
  * Rule-based models
  * Probabilistic models

Most IE pipelines start with the usual text preprocessing steps - sentence segmentation, word tokenisation and POS tagging. After preprocessing, the common tasks are **Named Entity Recognition (NER)**, and optionally relation recognition and record linkage. NER is arguably the most important and non-trivial task in the pipeline. There are various techniques and models for building Named Entity Recognition (NER) system, which is a key component in information extraction systems:

* **Rule-based techniques**
  * Regular expression-based techniques
  * Chunking
* **Probabilistic models**
  * Unigram & Bigram models
  * Naive Bayes Classifier
  * Decision trees
  * Conditional Random Fields (CRFs)

IOB (or BIO) method tags each token in the sentence with one of the three labels: **I - inside (the entity), O- outside (the entity) and B - beginning (of entity).** You saw that IOB labeling is especially helpful if the entities contain multiple words. For example: words like ‘*Delta Airlines*’, ‘*New York'*, etc., are single entities.

### Rule-based method for NER

Chunking is a common shallow parsing technique used to chunk words that constitute some meaningful phrase in the sentence. A noun phrase chunk (NP chunk) is commonly used in NER tasks to identify groups of words that correspond to some 'entity'.

**Sentence:** She bought *a new car* from *the BMW showroom*.

**Noun phrase chunks:** *a new car*, *the BMW showroom*

The idea of chunking in the context of entity recognition is simple - most entities are nouns and noun phrases, so rules can be written to extract these noun phrases and hopefully extract a large number of named entities. Example of chunking done using regular expressions:

**Sentence:** John booked the hotel.

**Noun phrase chunks:** 'John', 'the hotel'

**Grammar:** $$\text{NP\_chunk: {<DT>?<NN>}}$$

**Probabilistic method for NER**

The following two probabilistic models to get the most probable IOB tags for word:

* **Unigram chunker** computes the unigram probabilities P(IOB label | pos) for each word and assigns the label that is most likely for the POS tag.
* **Bigram chunker** works similar to a unigram chunker, the only difference being that now the probability of a POS tag having an IOB label is computed using the current and the previous POS tags, i.e., P(label | pos, prev\_pos).

**Gazetteer Lookup**, another way to identify named entities (like cities and states) is to look up a dictionary or a gazetteer. A gazetteer is a geographical directory which stores data regarding the names of geographical entities (cities, states, countries) and some other features related to the geographies.

**Naive Bayes and Decision Tree classifier can also be used in NER.**

**Conditional Random Fields**

HMMs can be used for any sequence classification task, such as NER. However, many NER tasks and datasets are far more complex than tasks such as POS tagging, and therefore, more sophisticated sequence models have been developed and widely accepted in the NLP community. One of these models is Conditional Random Fields (CRFs).

CRFs are used in a wide variety of sequence labelling tasks across various domains - POS tagging, speech recognition, NER, and even in computational biology for modelling genetic patterns etc. CRFs model the **conditional probability** $$P(Y|X)$$, where $$Y$$ is the vector of output sequence (IOB labels here) and $$X$$ is the input sequence (words to be tagged), which are similar to Logistic Regression classifier. Broadly, there are two types of classifiers in ML:

* **Discriminative classifiers** learn the boundary between classes by modelling the conditional probability distribution $$P(Y|X)$$, where $$Y$$ is the vector of class labels and $$X$$ represents the input features. Examples are Logistic Regression, SVMs etc.
* **Generative classifiers** model the joint probability distribution $$P(Y|X)$$. Examples of generative classifiers are Naive Bayes, HMMs etc.

CRFs use ‘feature functions’ rather than the input word sequence $$x$$ itself. The idea is similar to how features are extracted for building the naive Bayes and decision tree classifiers in a previous section. Some example ‘word-features’ (each word has these features) are:

* Word and POS tag-based features: word\_is\_city, word\_is\_digit, pos, previous\_pos, etc.
* Label-based features: previous\_label

A feature function takes the following four inputs:

* The input sequence of words: $$x$$
* The position of a word in the sentence (whose features are to be extracted)
* The label $$y\_i$$ of the current word (the target label)
* The label $$y\_{i-1}$$ of the previous word

Let's see an example of a feature function:

A feature function $$f\_1$$ which returns $$1$$ if the word $$x\_i$$ is a city and the corresponding label $$y\_i$$ is ‘I-location’, else $$0$$. This can be represented as:

$$f\_{1}(x,i,y\_i,y\_{i-1})= \[\[x\_i \text{ is in city last name}] \text{ and } \[y\_i \text{ is I-location}]]$$

The feature function returns $$1$$ only if both the conditions are satisfied, i.e. when the word is a city name and is tagged as ‘I-location’ (e.g. Tokyo/I-location).

Every feature function $$f\_i$$ has a weight $$w\_i$$ associated with it, which represents the ‘importance’ of that feature function. This is almost exactly the same as logistic regression where coefficients of features represent their importance. Training a CRF means to compute the optimal weight vector $$w$$ which best represents the observed sequences $$y$$ for the given word sequences $$x$$. In other words, we want to find the set of weights $$w$$ which maximises $$P(y|x,w)$$.

In CRFs, the conditional probabilities $$P(y|x,w)$$ are modeled using a scoring function. If there are $$k$$ feature functions (and thus $$k$$ weights), for each word $$i$$ in the sequence $$x$$, a scoring function for a word is defined as follows:

$$score\_i = exp(w\_1.f\_1 + w\_2.f\_2 ... + w\_k.f\_k) = exp(w\.f(y\_i,x\_i,y\_{i-1},i))$$

and the overall sequence score for the sentence can be defined as:

$$\text{sequence-score}(y|x) = \prod\_{i=1}^n (exp(w\.f(y\_i,x\_i,y\_{i-1},i))) = exp(\sum\_1^n(w\.f(y\_i,x\_i,y\_{i-1},i)))$$

The probability of observing the label sequence $$y$$ given the input sequence $$x$$ is given by:

$$P(y|x,w) = exp(\sum\_1^n(w\.f(y\_i,x\_i,y\_{i-1},i)))/Z(x) = exp(w\.f(x,y))/Z(x)$$

where $$Z(x)$$ is sum of scores of all possible tag sequences $$N$$ $$= \sum\_1^N(exp(w\.f(x,y)))$$

Training a CRF model means to compute the optimal set of weights $$w$$ which best represents the observed sequences $$y$$ for the given word sequences $$x$$. In other words, we want to find the set of weights $$w$$ which maximises the conditional probability $$P(y|x,w)$$ for all the observed sequences $$(x,y)$$, by taking log and simplifying the equations and adding a regularization term to prevent overfitting, the final equation comes out as:

$$L(w) = \sum\_1^N\[(w\.f)-log(Z)] - \text{regularization term}$$

The inference task to assign the label sequence $$y^\*$$ to $$x$$ which maximises the score of the sequence, i.e.

$$y^\* = argmax(w\.f(x,y))$$

The naive way to get $$y^\*$$ is by calculating $$w\.f(x,y)$$ for every possible label sequence , and then choose the label sequence that has maximum $$(w\.f(x,y))$$ value. However, there are an exponential number of possible labels ($$t^n$$ for a tag set of size $$t$$ and a sentence of length $$n$$), and this task is computationally heavy.


# Transformers

A summary of transformers and why it makes it important in Data Science

The Transformer architecture is a breakthrough in natural language processing (NLP) that has had a profound impact on various fields, including data science. It was introduced in the paper ["Attention is All You Need" by Vaswani et al. in 2017](https://arxiv.org/abs/1706.03762) and has since become the foundation for many state-of-the-art models, including BERT, GPT, and more. Here's a simplified explanation and its importance:

**Explanation:**

At its core, the Transformer architecture is designed to handle sequences of data (like sentences or time series) while capturing contextual relationships effectively. It consists of two key components: self-attention mechanisms and feedforward neural networks.

1. **Self-Attention Mechanism**: Self-attention is a technique that allows each word/token in a sequence to consider the importance of other words/tokens in relation to itself. This enables the model to weigh the significance of different words within the context of the entire sequence. Self-attention helps capture long-range dependencies and relationships between words in a more efficient way compared to traditional recurrent or convolutional approaches.
2. **Positional Encodings**: Since the Transformer processes words in parallel rather than sequentially, it doesn't inherently understand the order of words in a sequence. Positional encodings are added to the word embeddings to provide information about their positions, ensuring the model understands the sequential nature of the input.
3. **Multi-Head Attention**: To capture different types of relationships and nuances, the Transformer employs multi-head attention, where multiple self-attention mechanisms operate in parallel. This allows the model to focus on different parts of the input and learn various contextual relationships simultaneously.
4. **Feedforward Neural Networks**: After the self-attention step, the model passes the information through feedforward neural networks to further process and refine the features captured during self-attention.

**Importance:**

The Transformer architecture is highly important in data science for several reasons:

1. **Highly Effective for Sequences**: The Transformer's self-attention mechanisms excel at capturing context within sequences. This makes it well-suited for a wide range of data types, including text, time series, and more.
2. **Reduced Long-Term Dependencies Issue**: Traditional recurrent neural networks suffer from the vanishing gradient problem, which makes it hard for them to capture long-term dependencies. Transformers overcome this challenge through self-attention, enabling them to consider long-range relationships effectively.
3. **Parallelization and Efficiency**: Transformers process words in parallel rather than sequentially, making them computationally efficient and suitable for modern hardware, such as GPUs.
4. **State-of-the-Art Performance**: Many NLP models built on the Transformer architecture, like BERT and GPT, have achieved remarkable results across a range of NLP tasks. This has influenced the development of new techniques and models in various domains, contributing to advancements in data science.

In summary, the Transformer architecture's ability to capture contextual relationships efficiently has revolutionized NLP and extended its influence to other data science domains, leading to improved performance and capabilities in understanding and processing sequences of data.


# Power BI

Overview of Power BI and its core components.

Power BI is a technology-driven business intelligence tool provided by Microsoft for analyzing and visualizing raw data to present actionable information. It combines business analytics, data visualization, and best practices that help an organization to make data-driven decisions. In February 2019, [Gartner](https://start.paloaltonetworks.com/2019-gartner-mq-for-firewalls) confirmed Microsoft as Leader in the "*2019 Gartner Magic Quadrant for Analytics and Business Intelligence Platform*" as a result of the capabilities of the Power BI platform.

### Components of Power BI <a href="#components_of_power_bi" id="components_of_power_bi"></a>

1. #### Power Query&#x20;

   Power Query is the data transformation and mash up the engine. It enables you to discover, connect, combine, and refine data sources to meet your analysis need. It can be downloaded as an add-in for Excel or can be used as part of the Power BI Desktop.
2. #### Power Pivot&#x20;

   Power Pivot is a data modeling technique that lets you create data models, establish relationships, and create calculations. It uses Data Analysis Expression (DAX) language to model simple and complex data.
3. #### Power View&#x20;

   Power View is a technology that is available in Excel, Sharepoint, SQL Server, and Power BI. It lets you create interactive charts, graphs, maps, and other visuals that bring your data to life. It can connect to data sources and filter data for each data visualization element or the entire report.
4. #### Power Map&#x20;

   Microsoft's Power Map for Excel and Power BI is a 3-D data visualization tool that lets you map your data and plot more than a million rows of data visually on Bing maps in 3-D format from an Excel table or Data Model in Excel. Power Map works with Bing maps to get the best visualization based on latitude, longitude, or country, state, city, and street address information.
5. #### Power BI Desktop&#x20;

   Power BI Desktop is a development tool for Power Query, Power Pivot, and Power View. With Power BI Desktop, you have everything under the same solution, and it is easier to develop BI and data analysis experience.
6. #### Power Q\&A

   The Q\&A feature in Power BI lets you explore your data in your own words. It is the fastest way to get an answer from your data using natural language.


# Charts

Common Power BI charts and when to use them.

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FT0KPcDps6ZM5RtFYk8UP%2Fimage.png?alt=media&amp;token=3ed3dbea-7e6c-4dfe-bc22-89d616a99af1" alt=""><figcaption></figcaption></figure>

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FTITJSMHQJQw6ETFEX45W%2Fimage.png?alt=media&amp;token=6531ccf1-ce57-4676-943d-31c2e498c5ce" alt=""><figcaption></figcaption></figure>

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FC1iBEqkIk8hyDYHqq3Pq%2Fimage.png?alt=media&amp;token=0cb15307-bcf2-4a06-a462-77496df427c5" alt=""><figcaption></figcaption></figure>

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FDEeAm1AjK2RKzP70VabK%2Fimage.png?alt=media&amp;token=3734e796-425b-4aa9-89c4-cd5808aa94e1" alt=""><figcaption></figcaption></figure>

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2F7oCak4OFi9jKlu1tkOPF%2Fimage.png?alt=media&amp;token=98c4c441-8fb5-4254-8310-4a1aa35c5125" alt="" width="375"><figcaption><p>(<a href="https://tanducits.com/files/Data%20Visualization%20Cheat%20Sheet.pdf">Source</a>)</p></figcaption></figure>


# Problems

<details>

<summary>Calculated Column vs Measure</summary>

Can you explain the difference between Calculated Column vs Measure?

**Answer**

A calculated column belongs to a single table, while a measure belongs to the whole data model. A calculated column is evaluated in a row context (row by row, like in an excel table), while a measure is evaluated in the filter context.

Measures result in lower file size compared to calculated columns.

</details>

<details>

<summary>Difference in Direct Query and Importing data in PowerBI</summary>

Can you explain the difference between Direct Query and Importing the data in Power BI?

**Answer**

When using the Direct Query method of connection, your dashboard will be directly querying the data source at run time. Every filter and interaction with the report will kick off further queries. No data is imported into Power BI so you are always querying the data that is present in the data source itself.

The Import method of connection means that Power BI will cache the data that you’re connected to creating a point in time snapshot of your data. All of the interactions and filters applied to your data will be done to this compressed cache source instead of the actual data source itself.&#x20;

Advantages of using Power BI Direct Query:

* Data is queried from the data source so you are getting the most up to date data. The report refreshes occur every 15 minutes.
* Since you are not caching your data when using Direct Query, your Power BI Desktop files are much smaller and easier to work with (faster saving, publishing etc.)

Disadvantages of using Power BI Direct Query:

* Because you’re querying the data source at run time, you might be competing with other users for bandwidth. You’re also not taking advantage of the compression of the Vertipaq performance engine.
* You are not able to use all of the normal Power Query transformation features. Particular DAX functions are not available in this method as well. So if your data is poorly structured or needing lots of transformation, sometimes Direct Query is not a viable option.

Advantages of using Power BI Import:

* When you cache your data you are able to take full advantage of the Vertipaq performance engine. Normally your report performance will be better using this method.
* Unlike in Direct Query, you are able to use all M and DAX functions (notably all time intelligence functions), format fields however you desire, and there are no limitations to data modeling.
* Using Import you are able to combine data sources from various data sources (data flows, databases, csv).

Disadvantages of using Power BI Import:

* You can schedule up to 8 refreshes a day ([Premium SKUs](https://www.phdata.io/blog/what-is-the-price-of-power-bi-premium-and-what-sku-should-you-choose/) allow more), but you also need to consider the amount of reports you’re maintaining and how big the data sets are that you’re refreshing.
* Import caches are limited to 1GB per dataset (can be increased in Premium). While the Vertipaq engine does a great job a compression, you will still need to consider this when choosing your connection method
* Crazy enough, once you’ve selected Import, you cannot switch back to using Direct Query. So make sure you want Import before making the switch, or else you’ll have more work ahead of you!

</details>


# Visualization

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FTDA8EwijantUvL9Uvcag%2Fimage1.jpg?alt=media&amp;token=cde24b0a-7042-4e81-942f-cf983893ac43" alt=""><figcaption><p><a href="https://experception.net/">Source</a></p></figcaption></figure>


# Theoretical

<details>

<summary>Built in Datatypes</summary>

What are the *built-in types* available In Python?

**Answer**

Common *immutable* type:

1. numbers: `int()`, `float()`, `complex()`
2. immutable sequences: `str()`, `tuple()`, `frozenset()`, `bytes()`

Common *mutable* type (almost everything else):

1. mutable sequences: `list()`, `bytearray()`
2. set type: `set()`
3. mapping type: `dict()`
4. classes, class instances
5. etc.

You have to understand that Python represents all its data as objects. Some of these objects like lists and dictionaries are mutable, meaning you can change their content without changing their identity. Other objects like integers, floats, strings and tuples are objects that can not be changed.

</details>

<details>

<summary>Lambda</summary>

What is *Lambda Functions* in Python?

**Answer**

A **Lambda Function** is a small anonymous function. A lambda function can take *any* number of arguments but can *only* have *one* expression.

```python
x = lambda a : a + 10
print(x(5)) # Output: 15
```

</details>

<details>

<summary><code>tuple</code> vs <code>list</code> vs <code>dictionary</code></summary>

When to use a `tuple` vs `list` vs `dictionary` in Python?

**Answer**

* Use a `tuple` to store a sequence of items that *will not change*.
* Use a `list` to store a sequence of items that *may change*.
* Use a `dictionary` when you want to associate *pairs* of two items.

</details>

<details>

<summary>local vs global</summary>

What are the rules for local and global variables in Python?

**Answer**

While in many or most other programming languages variables are treated as global if not declared otherwise, Python deals with variables the other way around. They are local, if not otherwise declared.

* In Python, variables that are only referenced inside a function are implicitly *global*.
* If a variable is assigned a value anywhere within the function’s body, it’s assumed to be a *local* unless explicitly declared as global.

Requiring global for assigned variables provides a bar against unintended side-effects.

</details>

<details>

<summary>Descriptors</summary>

What are descriptors?

**Answer**

Descriptors were introduced to Python way back in version 2.2. They provide the developer with the ability to add managed attributes to objects. The methods needed to create a descriptor are `__get__`, `__set__` and `__delete__`. If you define any of these methods, then you have created a descriptor.

Descriptors power a lot of the magic of Python’s internals. They are what make properties, methods and even the super function work. They are also used to implement the new style classes that were also introduced in Python 2.2.

</details>

<details>

<summary>switch case</summary>

Does Python have a *switch-case* statement?

**Answer**

In Python before 3.10, we **do not hav**e a switch-case statement. Here, you may write a switch function to use. Else, you may use a set of if-elif-else statements. To implement a function for this, we may use a dictionary.

```python
def switch_demo(argument):
    switcher = {
        1: "January",
        2: "February",
        3: "March",
        4: "April",
        5: "May",
        6: "June",
        7: "July",
        8: "August",
        9: "September",
        10: "October",
        11: "November",
        12: "December"
    }
    print switcher.get(argument, "Invalid month")
```

Python 3.10 (2021) introduced the [`match`-`case`](https://www.python.org/dev/peps/pep-0634/) statement which provides a first-class implementation of a "switch" for Python. For example:

For example:

```python
def f(x):
    match x:
        case 'a':
            return 1
        case 'b':
            return 2
```

The `match`-`case` statement is considerably more powerful than this simple example.

</details>

<details>

<summary>static methods</summary>

Is it possible to have static methods in Python?

**Answer** ([Source](https://www.digitalocean.com/community/tutorials/python-static-method))

Static methods in Python are extremely similar to [python class](https://www.digitalocean.com/community/tutorials/python-classes-objects) level methods, the difference being that a static method is bound to a class rather than the objects for that class. This means that a static method can be called without an object for that class. This also means that static methods cannot modify the state of an object as they are not bound to it. Let’s see how we can create static methods in Python.

Static methods have a very clear use-case. When we need some functionality not w\.r.t an Object but w\.r.t the complete class, we make a method static. This is pretty much advantageous when we need to create Utility methods as they aren’t tied to an object lifecycle usually. Finally, note that in a static method, we don’t need the `self` to be passed as the first argument.

</details>

<details>

<summary><code>range</code> vs <code>xrange</code></summary>

What is the difference between `range` and `xrange` functions in Python?

**Answer (**[**Source**](https://www.geeksforgeeks.org/range-vs-xrange-in-python/)**)**

* **range()** – This returns a range object (a type of iterable).
* **xrange()** – This function returns the **generator object** that can be used to display numbers only by looping. The only particular range is displayed on demand and hence called “**lazy evaluation**“.

</details>

<details>

<summary>Pickling and Unpickling</summary>

What is Pickling and Unpickling?

**Answer**

*“Pickling”* is the process whereby a Python object hierarchy is converted into a byte stream, and *“unpickling”* is the inverse operation, whereby a byte stream (from a binary file or bytes-like object) is converted back into an object hierarchy.

</details>

<details>

<summary><code>*args vs</code> <code>**kwargs</code></summary>

What does this stuff mean: `*args`, `**kwargs`? Why would we use it?

**Answer (**[**Source**](https://www.geeksforgeeks.org/args-kwargs-python/)**)**

Special Symbols Used for passing arguments:-

* \*args (Non-Keyword Arguments)
* \*\*kwargs (Keyword Arguments)

The special syntax *\*args* in function definitions in python is used to pass a variable number of arguments to a function. It is used to pass a non-key worded, variable-length argument list.&#x20;

The special syntax *\*\*kwargs* in function definitions in python is used to pass a keyworded, variable-length argument list. We use the name *kwargs* with the double star. The reason is that the double star allows us to pass through keyword arguments (and any number of them).

* A keyword argument is where you provide a name to the variable as you pass it into the function.
* One can think of the *kwargs* as being a dictionary that maps each keyword to the value that we pass alongside it. That is why when we iterate over the *kwargs* there doesn’t seem to be any order in which they were printed out.

&#x20;&#x20;

</details>

<details>

<summary>eggs vs wheels</summary>

What are the Wheels and Eggs? What is the difference?

**Answer**

Wheel and Egg are both packaging formats that aim to support the use case of needing an install artifact that **doesn’t require building or compilation**, which can be costly in testing and production workflows.

The Egg format was introduced by setuptools in 2004, whereas the Wheel format was introduced by PEP 427 in 2012.

Egg packages are an older standard, you should ignore them nowadays. Use `pip install .` instead of `./setup.py install` to prevent creating them. (addendum: They are also `.zip`s in disguise, from which Python reads package data — not exactly the most performant solution)

Wheel packages, on the other hand, are the new standard. They allow for creation of portable binary packages for Windows, macOS, and Linux [(yes, Linux!)](https://www.python.org/dev/peps/pep-0513/). Nowadays, you can just do `pip install PyQt5` (as an example) and it will just work, no C++ compiler and Qt libraries required on the system. Everything is pre-compiled and included in the wheel. Non-binary packages also benefit, because it’s safer not to run `setup.py` (all the metadata is in the wheel). (addendum: those are also `.zip`s, but they are unpacked when installed)

</details>

<details>

<summary><code>meshgrid</code></summary>

What is `meshgrid` in Python?

**Answer** ([Source](https://www.educba.com/numpy-meshgrid/))

In python, meshgrid is a function that creates a rectangular grid out of 2 given 1-dimensional arrays that denotes the Matrix or Cartesian indexing. It is inspired from MATLAB. This meshgrid function is provided by the module numpy. Coordinate matrices are returned from the coordinate vectors.

</details>

<details>

<summary>metaclass</summary>

What is a metaclass in Python?

**Answer (**[**Source**](https://www.datacamp.com/tutorial/python-metaclasses)**)**

A metaclass in Python is a class of a class that defines how a class behaves. A class is itself an instance of a metaclass. A class in Python defines how the instance of the class will behave. In order to understand metaclasses well, one needs to have prior experience working with Python classes.

</details>


# Basics

This page deals with Basic Python Questions

There are some common programming techniques that you must be familiar with if you want to be comfortable enough to solve Python questions live during an interview. Here we have summarized a few of such techniques but do hone your skills using platforms like Leet code to ensure that you perform well during interviews.

We have collated some resources from the internet as a starting point to help you prepare on this:

* [14 Coding Interview Patterns](https://hackernoon.com/14-patterns-to-ace-any-coding-interview-question-c5bb3357f6ed)
* [Sliding Window patterns](https://leetcode.com/problems/frequency-of-the-most-frequent-element/solutions/1175088/C++-Maximum-Sliding-Window-Cheatsheet-Template/)
* [Two Pointers Patterns](https://leetcode.com/discuss/study-guide/1688903/Solved-all-two-pointers-problems-in-100-days)
* [Substring Problem Patterns](https://leetcode.com/problems/minimum-window-substring/solutions/26808/Here-is-a-10-line-template-that-can-solve-most-'substring'-problems/)
* [Dynamic Programming Patterns](https://leetcode.com/discuss/study-guide/458695/Dynamic-Programming-Patterns)
* [Binary Search Patterns](https://leetcode.com/discuss/study-guide/786126/Python-Powerful-Ultimate-Binary-Search-Template.-Solved-many-problems)
* [Backtracking Patterns](https://leetcode.com/problems/permutations/solutions/18239/A-general-approach-to-backtracking-questions-in-Java-\(Subsets-Permutations-Combination-Sum-Palindrome-Partioning\)/)
* [Tree Patterns](https://leetcode.com/discuss/study-guide/937307/Iterative-or-Recursive-or-DFS-and-BFS-Tree-Traversal-or-In-Pre-Post-and-LevelOrder-or-Views)
  * [Tree Iterative Traversal](https://medium.com/leetcode-patterns/leetcode-pattern-0-iterative-traversals-on-trees-d373568eb0ec)
  * [Tree Question Pattern](https://leetcode.com/discuss/study-guide/2879240/TREE-QUESTION-PATTERN-2023-oror-TREE-STUDY-GUIDE)
* [Graph Patterns](https://leetcode.com/discuss/study-guide/655708/Graph-For-Beginners-Problems-or-Pattern-or-Sample-Solutions)
* [Monotonic Stack Patterns](https://leetcode.com/discuss/study-guide/2347639/A-comprehensive-guide-and-template-for-monotonic-stack-based-problems)
* [Bit Manipulation Patterns](https://leetcode.com/discuss/study-guide/3901862/All-Types-of-Patterns-for-Bits-Manipulations-and-How-to-use-it)
* [String Question Patterns](https://leetcode.com/discuss/study-guide/2001789/Collections-of-Important-String-questions-Pattern)
* [DFS + BFS Patterns (1)](https://medium.com/leetcode-patterns/leetcode-pattern-2-dfs-bfs-25-of-the-problems-part-2-a5b269597f52)
* [DFS + BFS Patterns (2)](https://medium.com/leetcode-patterns/leetcode-pattern-2-dfs-bfs-25-of-the-problems-part-2-a5b269597f52)

### Questions

<details>

<summary>[SNAPCHAT] Palindrome Checker</summary>

Given a string, determine whether any permutation of it is a palindrome.

For example, *carerac* should return *true*, since it can be rearranged to form *racecar*, which is a palindrome. *sunset* should return *false*, since there’s no rearrangement that can form a palindrome.

**Answer**

<pre class="language-python"><code class="lang-python"># A string can be a palindrome only if it has even pair of characters and at max 1 odd character
<strong>def palindrome(x):
</strong>    char_dict = {}
    for i in x:
# we will check if the element exists else we will add it to the dictionary
        try:
            char_dict[i] = char_dict[i] + 1
        except:
            char_dict.update({i:1})
 # next we will create a list of element counts and use list comprehension 
 # to check if odd element count is 1 and rest all even           
    check = list(char_dict.values())
    if (len([temp for temp in check if temp%2==1]) == 1 and len([temp for temp in check if temp%2==0]) > 1):
        return "palindrome"
    else:
        return "not palindrome"
    
print(palindrome("carerac"))
print(palindrome("abc"))
</code></pre>

</details>

<details>

<summary>Pattern Generation</summary>

Write a function to generate this pattern:\
1\
2 3\
4 5 6

Now change the code to output\
1\
1 2\
1 2 3

**Answer**

```python
def pyramid(n):
    counter = 1
    max = 1
    while(max<n):
        for i in range(max, max+counter): # for the second pattern change max to 1
            print(i, end=" ")
            max=max+1
        counter=counter+1
        print(" ")
    
pyramid(6)
```

</details>

<details>

<summary>[UBER] <a href="https://leetcode.com/problems/combination-sum/description/">Sum to N</a></summary>

Given a list of positive integers, find all combinations that equal the value N.

Example:

*integers = \[2,3,5], target = 8,*

*output = \[\[2,2,2,2],\[2,3,3],\[3,5]]*

**Answer**

We will solve it in 2 ways, one using itertools the other one using recursion:

```python
def combinationSum(candidates, target):
    ans = []                                        # for adding all the answers
    def traverse(candid, arr, sm):                  # arr : an array that contains the accused combination; sm : is the sum of all elements of arr 
        if sm == target: ans.append(arr)            # If sum is equal to target then add it to the final list
        if sm >= target: return                     # If sum is greater than target then no need to move further.
        for i in range(len(candid)):                # we will traverse each element from the array.
            traverse(candid[i:], arr + [candid[i]], sm+candid[i])   #most important, splice the array including the current index, splicing in order to handle the duplicates.
    traverse(candidates,[], 0)
    return ans

combinationSum([2, 3, 5], 8)
```

```python
import itertools
integers = [2,3,5]
target = 8
final = []
# output = [[2,2,2,2],[2,3,3],[3,5]]

max = target//min(integers)
for i in range(1,max+1):
  a = list(itertools.combinations_with_replacement(integers,i))
  for k in a:
    
    if sum(list(k)) == target:
      final.append(k)

print(final)
```

</details>

<details>

<summary>[STARBUCKS] Max Profit</summary>

Given a list of stock prices in ascending order by datetime, write a function that outputs the max profit by buying and selling at a specific interval.

*Example:*

*stock\_prices = \[10,5,20,32,25,12]*

*buy –> 5 sell –> 32*

**Answer**

```python
li = [10,5,20,32,25,12]
diff = 0
for i,j in enumerate(li):
    try:
        if(diff<(max(li[i+1:])-j)):
            diff = max(li[i+1:])-j
            buy = j
            sell = max(li[i+1:])
        print(j, max(li[i+1:]))
    except:
        print("end")
    
print(buy, sell, diff)
```

</details>

<details>

<summary>[IBM] Isomorphic string check</summary>

Write a function which will check if each character of string1 can be mapped to a unique character of string2.

*Example: string1 = ‘donut’ string2 = ‘fatty’*

*string\_map(string1, string2) == False # as n and u both get mapped to t*

*string1 = ‘enemy’ string2 = ‘enemy’*

*string\_map(string1, string2) == True # as e’s get mapped to e even though there is two e*

*string1 = ‘enemy’ string2 = ‘yneme’*

*string\_map(string1, string2) == False # as e’s dont get mapped uniquely*

**Answer**

```python
def string_map(string1, string2):    
    if(string1==string2):
        status = True
    elif(len(string1)!=len(string2)):
        status = False
    else:
        tempstore = {}
        for i,j in enumerate(string1):
            if(j in tempstore):               
                if(tempstore[j] != string2[i]):
                    status = False
                    break
            elif(string2[i] in tempstore.values()):
                    status = False
                    break
            else:
                tempstore[j] = string2[i]
                status = True
    return status

print(string_map('enemy', 'enemy'))
print(string_map('enemy', 'yneme'))
print(string_map('cat', 'ftt'))
print(string_map('ctt', 'fat'))
print(string_map('cat', 'fat'))
```

</details>

<details>

<summary>[WORKDAY] Sorted String merge</summary>

Given two sorted lists, write a function to merge them into one sorted list.

What’s the time complexity?

**Answer**

{% code overflow="wrap" %}

```python
def mergeArrays(arr1, arr2):
    
    n1 = len(arr1)
    n2 = len(arr2)
    arr3 = [None] * (n1 + n2)
    i = 0
    j = 0
    k = 0
    # Traverse both array
    while i < n1 and j < n2:     
        # Check if current element of first array is smaller than current element of second array. 
        # If yes, store first array element and increment first array index. Otherwise do same with second array
        
        if arr1[i] < arr2[j]:
            arr3[k] = arr1[i]
            k = k + 1
            i = i + 1
        else:
            arr3[k] = arr2[j]
            k = k + 1
            j = j + 1
     
    # Store remaining elements
    # of first array
    while i < n1:
        arr3[k] = arr1[i];
        k = k + 1
        i = i + 1
 
    # Store remaining elements
    # of second array
    while j < n2:
        arr3[k] = arr2[j];
        k = k + 1
        j = j + 1
    print("Array after merging")
    for i in range(n1 + n2):
        print(str(arr3[i]), end = " ")
 

arr1 = [1, 3, 5, 7]
arr2 = [2, 4, 6, 8]
mergeArrays(arr1, arr2)

# Next coming to the time complexity it is linear as the execution time of the algorithm grows in direct proportion to the size of the data set it is processing.
# For merging two arrays, we are always going to iterate through both of them no matter what, 
# so the number of iterations will always be m+n and the time complexity being O(m+n) where m = len(arr1) and n = len(arr2)
```

{% endcode %}

</details>

<details>

<summary>[POSTMATES] Weekly Aggregation</summary>

Given a list of timestamps in sequential order, return a list of lists grouped by week (7 days) using the first timestamp as the starting point.

*Example:*

*ts = \[ ‘2019-01-01’, ‘2019-01-02’, ‘2019-01-08’, ‘2019-02-01’, ‘2019-02-02’, ‘2019-02-05’, ]*

*output = \[ \[‘2019-01-01’, ‘2019-01-02’], \[‘2019-01-08’], \[‘2019-02-01’, ‘2019-02-02’], \[‘2019-02-05’] ]*

**Answer**

{% code overflow="wrap" %}

```python
ts = [
    '2019-01-01', 
    '2019-01-02',
    '2019-01-08', 
    '2019-02-01', 
    '2019-02-02',
    '2019-02-05',
]

from datetime import datetime as dt
from itertools import groupby

first = dt.strptime(inp[0], "%Y-%m-%d")
out = []

for k, g in groupby(ts, key=lambda d: (dt.strptime(d, "%Y-%m-%d") - first).days // 7 ):
    out.append(list(g))

print(out)
```

{% endcode %}

</details>

<details>

<summary>[MICROSOFT] Find the missing number</summary>

You have an array of integers of length n spanning 0 to n with one missing. Write a function that returns the missing number in the array

*Example:*

*nums = \[0,1,2,4,5] missingNumber(nums) -> 3*

Complexity of O(N) required.

**Answer**

```python
def missingNumber(nums):

  diff = 0
  miss_num = []
  for i, j in enumerate(nums):
    try:
      t_diff = nums[i+1]-nums[i]
      if t_diff>0:
        for k in range(i+1,i+t_diff):
          # print(k)
          miss_num.append(k)
    except:
      pass
  return miss_num

nums = [0,1,2,5,6]
missingNumber(nums)
```

</details>

<details>

<summary>[SQUARE] Book Combinations</summary>

You have store credit of N dollars. However, you don’t want to walk a long distance with heavy books, but you want to spend all of your store credit.

Let’s say we have a list of books in the format of tuples where the first value is the price and the second value is the weight of the book -> (price,weight).

Write a function *optimal\_books* to retrieve the combination allows you to spend all of your store credit while getting at least two books at the lowest weight.

*Note: you should spend all your credit and getting at least 2 books, If no such condition satisfied just return empty list.*

*Example:*

```
N = 18
books = [(17,8), (9,4), (18,5), (11,9), (1,2), (13,7), (7,5), (3,6), (10,8)]

def optimal_books(N, books) -> [(17,8),(1,2)]
```

**Answer**

{% code overflow="wrap" %}

```python
# Let's take this step by step
import itertools

def optimal_books(N, books):
    print("(Price,Weight) details of books: ",books)
    print("Store Credit: ",N)
    final_books = [] # empty list to store the final books
    # sorting the books by weight as we need the lightest books
    sorted_books = sorted(books, key = lambda x:x[1]) 
    price = [i[0] for i in sorted_books] #list of prices sorted by weight
    
    for i in range(2,len(price)+1):
        templist = (list(itertools.combinations(price,i))) # generating all combinations of price
        res = [sum(j) for j in templist] # summing individual combination to get total price of each combination
        if N in res: # if the result matches traceback traceback and append the combination         
            tempbooks = (templist[res.index(N)])            
            for k in tempbooks:
                final_books.append(sorted_books[price.index(k)])            
            break
            
    return final_books
        
N = 18
books = [(17,8), (9,4), (18,5), (11,9), (1,2), (13,7), (7,5), (3,6), (10,8)]
print("Best Combination: ",optimal_books(N,books))
```

{% endcode %}

</details>

<details>

<summary>[WISH] Intersecting Lines</summary>

Say you are given a list of tuples where the first element is the slope of a line and the second element is the y-intercept of a line.

Write a function find\_intersecting to find which lines, if any, intersect with any of the others in the given x\_range.

*Example*

`tuple_list = [(2, 3), (-3, 5), (4, 6), (5, 7)] x_range = (0, 1)`

*Output*

`def find_intersecting(tuple_list, x_range) -> [(2,3), (-3,5)]`

**Answer**

```python
# for 2 lines to intersect the formulas used here are:
# y = mx + c
# x = (c2-c1)/(m1-m2)
# https://www.cuemath.com/geometry/intersection-of-two-lines/ Check this link for details of the formula

def intersectinglines(tuple_list,x_range):
    output=[]
    for i in range(len(tuple_list)):
        for j in range(i+1,len(tuple_list)):

            x = (tuple_list[j][1]-tuple_list[i][1])/(tuple_list[i][0]-tuple_list[j][0])
            y = tuple_list[j][1]*x+tuple_list[j][0]

            if x>=x_range[0] and x<=x_range[1]:
                output.extend([tuple_list[i],tuple_list[j]])
    return output

tuple_list = [(2, 3), (-3, 5), (4, 6), (5, 7)]
x_range = (0, 1)

intersectinglines(tuple_list, x_range)
```

</details>

<details>

<summary>Find the majority element in a list.</summary>

a = \[2,3,4,6, 6, 2,2] answer --> 2

**Answer**

```python
a = [2,3,4,6, 6, 2,2]

def major_ele(x):
    counter_dict = {}
    for i in a:
        if i not in counter_dict:
           counter_dict.update({i:1})
        else:
           counter_dict[i] =  counter_dict[i] + 1 
    for i,j in counter_dict.items():
        if j == max(list(counter_dict.values())):
            return i

print(major_ele(a))
```

</details>

<details>

<summary>[INTUIT] Iterator</summary>

Implement an iterator function which takes three iterators as the input and sorts them.

**Answer**

2 Solutions are provided below:

```python
import heapq

def sorted_merge(*iterators):
    # Use heapq.merge to merge and sort the input iterators
    sorted_iterator = heapq.merge(*iterators)
    
    # Return the sorted iterator
    return sorted_iterator

# Example usage:
if __name__ == "__main__":
    # Create three sorted iterators (lists in this case)
    iterator1 = iter([1, 3, 5, 7])
    iterator2 = iter([2, 4, 6, 8])
    iterator3 = iter([0, 9, 10])
    
    # Merge and sort the iterators
    sorted_iterator = sorted_merge(iterator1, iterator2, iterator3)
    
    # Iterate through the sorted values
    for value in sorted_iterator:
        print(value)

```

```python
def sort_iterators(it1, it2, it3):
  """Sorts three iterators.

  Args:
    it1: The first iterator.
    it2: The second iterator.
    it3: The third iterator.

  Returns:
    An iterator that yields the sorted elements of the three iterators.
  """
  # Create a list to store the elements of the three iterators.
  elements = []

  # Iterate over the three iterators and add the elements to the list.
  for element in it1:
    elements.append(element)
  for element in it2:
    elements.append(element)
  for element in it3:
    elements.append(element)

  # Sort the list.
  elements.sort()

  # Create an iterator that yields the elements of the sorted list.
  return iter(elements)
```

</details>

<details>

<summary>[SPLUNK] Last Page Number</summary>

We're given a string of integers that represent page numbers.

Write a function to return the last page number in the string. If the string of integers is not in correct page order, return the last number in order.

```
input = '12345'
output = 5

input = '12345678910111213'
output = 13

input = '1235678'
output = 3
```

**Answer**

```python
def get_last_page(int_string):
    print(int_string)
    count = 0
    counter2 = 0
    for i in int_string:
        count = count+1+counter2//10
        counter2 = counter2+1
        if(str(counter2)==int_string[count-1:count+counter2//10]):
            pass
        else:
            return counter2-1
```

</details>

<details>

<summary>Python Recursion</summary>

Explain Python recursion with an example.

**Answer**

Recursion is a programming technique that allows a function to call itself. This can be useful for solving problems that involve self-similar structures, such as trees and graphs.

```python
def factorial(n): # 6! = 6*3*2*1
  if(n==0 or n==1): # define the base case
    return 1
  return factorial(n-1)*n # recursively call the func

factorial(6)
```

This function works by calling itself recursively to calculate the factorial. The base cases are when n is 0 or 1, in which case the function simply returns n.

Recursion can be a powerful tool, but it is important to use it carefully. If a recursive function is not designed carefully, it can easily lead to stack overflows.

</details>

<details>

<summary>[INTUIT] <a href="https://leetcode.com/problems/subarray-product-less-than-k/">Subarray Product Less Than K</a></summary>

This was asked in INTUIT Sr. Data Scientist initial round using Glider

Given an array of integers `nums` and an integer `k`, return *the number of contiguous subarrays where the product of all the elements in the subarray is strictly less than* `k`.

&#x20;**Example 1:**

<pre><code><strong>Input: nums = [10,5,2,6], k = 100
</strong><strong>Output: 8
</strong><strong>Explanation: The 8 subarrays that have product less than 100 are:
</strong>[10], [5], [2], [6], [10, 5], [5, 2], [2, 6], [5, 2, 6]
Note that [10, 5, 2] is not included as the product of 100 is not strictly less than k.
</code></pre>

**Example 2:**

<pre><code><strong>Input: nums = [1,2,3], k = 0
</strong><strong>Output: 0
</strong></code></pre>

&#x20;**Constraints:**

* `1 <= nums.length <= 3 * 104`
* `1 <= nums[i] <= 1000`
* `0 <= k <= 106`

**Answer**

You are not only required to solve the problem in a limited time frame (\~30mins) but the ask is also to ensure that the all test-cases pass and at least one of them fails if the code does not meet the required time complexity even if you get the required answer.

```python
class Solution:
    def numSubarrayProductLessThanK(self, nums: List[int], max_product: int) -> int:        
        result = 0
        for i in range(1, len(nums)):
            for j, l in enumerate(nums):
                temp = nums[j:j+i]
                res = 1
                for m in temp:
                    res = m*res
                if(res<max_product and len(temp)==i):
                    result +=1      
        return result
```

The above code fails some test cases as it has O(n^3) complexity. There is the two-pointer or inchworm approach given below which solves the problem along with taking care of the complexity. Click on the Leetcode link above and check the discussions if you want to understand it better:

```python
class Solution:
    def numSubarrayProductLessThanK(self, nums: List[int], max_product: int) -> int:        
        left = 0
        result = 0
        product = 1
        
        for right in range(len(nums)):
            product *= nums[right]
            
            if product >= max_product:
                while product >= max_product and left <= right:
                    product /= nums[left]
                    left += 1
            
            result += right - left + 1
        
        return result
```

</details>

<details>

<summary><a href="https://leetcode.com/problems/reverse-vowels-of-a-string/description/?envType=study-plan-v2&#x26;envId=leetcode-75">Reverse Vowels in a String</a></summary>

Given a string `s`, reverse only all the vowels in the string and return it.

The vowels are `'a'`, `'e'`, `'i'`, `'o'`, and `'u'`, and they can appear in both lower and upper cases, more than once.

**Example 1:**

<pre><code><strong>Input: s = "hello"
</strong><strong>Output: "holle"
</strong></code></pre>

**Example 2:**

<pre><code><strong>Input: s = "leetcode"
</strong><strong>Output: "leotcede"
</strong></code></pre>

**Answer**

This can be solved using the 2-pointer approach:

```python
class Solution:
    def reverseVowels(self, s: str) -> str:
        s = list(s)
        vow = "aeiouAEIOU"
        left = 0
        right = len(s)-1
        while left < right:
            if s[left] in vow and s[right] in vow:
                
                s[left], s[right] = s[right], s[left]
                
                left += 1; right -= 1
            
            elif s[left] not in vow:
                left += 1
            
            elif s[right] not in vow:
                right -= 1            
        return ''.join(s)
```

</details>

<details>

<summary><a href="https://leetcode.com/problems/move-zeroes/description/?envType=study-plan-v2&#x26;envId=leetcode-75">Swap numbers</a></summary>

Given an integer array `nums`, move all `0`'s to the end of it while maintaining the relative order of the non-zero elements.

**Note** that you must do this in-place without making a copy of the array.

**Example:**

<pre><code><strong>Input: nums = [0,1,0,3,12]
</strong><strong>Output: [1,3,12,0,0]
</strong></code></pre>

**Answer**

```python
class Solution:
    def moveZeroes(self, nums: list) -> None:
        slow = 0
        for fast in range(len(nums)):
            if nums[fast] != 0 and nums[slow] == 0:
                nums[slow], nums[fast] = nums[fast], nums[slow]

            # wait while we find a non-zero element to
            # swap with you
            if nums[slow] != 0:
                slow += 1
```

**Algorithm complexity:**\
*Time complexity: O(n)*. Our fast pointer does not visit the same spot twice.\
*Space complexity: O(1)*. All operations are made in-place

</details>

<details>

<summary>[SALESFORCE] Interpolation</summary>

![](https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FdQWmC6jiJI4b40bk2BIC%2Fimage.png?alt=media\&token=bfd0ea13-6583-4e4b-b5c8-48d93cd00ff0)

**Answer**

```python
x = [25, 50, 100]
y = [5.0, 4.0, 3.0]

def interporlate(n):
  if(n in x):
    return y[x.index(n)]
  elif(n > x[-1]):
    x_t = x[-1]-x[-2]
    y_t = y[-1]-y[-2]
    return y[-1] + (y_t / x_t * (n-x[-1]))
  elif(n < x[0]):
    x_t = x[0]-x[1]
    y_t = y[0]-y[1]
    return y[0] + (y_t / x_t * (n-x[0]))
  else:
    for i,j in enumerate(x):
      if n<j:
        break
    x_t = x[i-1]-x[i]
    y_t = y[i-1]-y[i]
    return y[i-1] + (y_t / x_t * (n-x[i-1]))


print(interporlate(50))
print(interporlate(150))
print(interporlate(25))
print(interporlate(75))
```

</details>

<details>

<summary>[SALESFORCE] Prison Problem</summary>

<img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2Fm8aeW147EwphXkkGse7e%2Fimage.png?alt=media&amp;token=6d5d4db5-2706-4b97-b83b-f31bf010772b" alt="" data-size="original">

**Answer**

```python
def prison(size, x, y):
  x_t = list(range(0,6,1))
  y_t = list(range(0,6,1))
  x_t = [a for a in x_t if a not in x]
  y_t = [a for a in y_t if a not in y]

  # accounting for the outer walls, we will add 0 as starting
  # and add +1 to all others
  x_t = list(map(lambda t: t + 1, x_t))
  y_t = list(map(lambda t: t + 1, y_t))
  x_t.insert(0,0)
  y_t.insert(0,0)

  max_x = 1
  max_y = 1

  for i,j in enumerate(x_t):
    try:
      if(max_x< x_t[i+1]-j):
        max_x = x_t[i+1] -j    
    except:
      pass

  for i,j in enumerate(y_t):
    try:
      if(max_y< y_t[i+1]-j):
        max_y = y_t[i+1] -j
    except:
      pass

  return "Max cell size: ", max_x*max_y


prison(5, [3, 2], [0, 1, 3])
```

</details>


# Data Manipulation

### Questions

<details>

<summary>[GOOGLE] Score Bucketization</summary>

Let’s say you’re given a list of standardized test scores from high schoolers from grades 9 to 12

Given the dataset, write code in Pandas to return the cumulative percentage of students that received scores within the buckets of <50, <75, <90, <100

Example Input:

```
|  user_id | grade | test score |
| -------- | ----- | ---------- |
| 1        | 10    | 85         |
| 2        | 10    | 60         |
| 3        | 11    | 90         |
| 4        | 10    | 30         |
| 5        | 11    | 99         |
```

Example Output:

```
| grade | test score | percentage |
| ----- | ---------- | ---------- |
| 10    | <50        | 30%        |
| 10    | <75        | 65%        |
| 10    | <90        | 96%        |
| 10    | <100       | 99%        |
| 11    | <50        | 15%        |
| 11    | <75        | 50%        |
```

**Answer**

```python
import pandas as pd
import numpy as np

df = pd.DataFrame([[1,10,85],[2,10,60],[3,11,90],[4,10,30],[5,11,99]], columns = ["user_id","grade","test score"])

df["<50"] = np.where(df["test score"]<50,1,0)
df["<75"] = np.where(df["test score"]<75,1,0)
df["<90"] = np.where(df["test score"]<90,1,0)
df["<100"] = np.where(df["test score"]<100,1,0)

df = df.groupby(["grade"])[["<50","<75","<90","<100"]].sum().reset_index()
df = df.melt(id_vars=["grade"],var_name="test score",value_name="count")

df["grp_ttl"] = df.groupby("grade")["count"].transform('max')
df["percentage"] = 100*df["count"]/df["grp_ttl"]

df = (df[["grade","test score","percentage"]].copy()).sort_values(["grade","percentage"],ascending=True)

df["percentage"] = df.percentage.astype(int).astype(str)
df["percentage"] = df["percentage"] + "%"

df.head(10)
```

</details>

<details>

<summary>[NEXTDOOR] Complete Addresses</summary>

You’re given two dataframes. One contains information about addresses and the other contains relationships between various cities and states:

df\_addresses

address

*4860 Sunset Boulevard, San Francisco, 94105 3055 Paradise Lane, Salt Lake City, 84103 682 Main Street, Detroit, 48204 9001 Cascade Road, Kansas City, 64102 5853 Leon Street, Tampa, 33605*

df\_cities

*city state Salt Lake City Utah Kansas City Missouri Detroit Michigan Tampa Florida San Francisco California*

Write a function complete\_address to create a single dataframe with complete addresses in the format of street, city, state, zipcode.

**Answer**

```python
import pandas as pd

addresses = {"address": ["4860 Sunset Boulevard, San Francisco, 94105", "3055 Paradise Lane, Salt Lake City, 84103", "682 Main Street, Detroit, 48204", "9001 Cascade Road, Kansas City, 64102", "5853 Leon Street, Tampa, 33605"]}

cities = {"city": ["Salt Lake City", "Kansas City", "Detroit", "Tampa", "San Francisco"], "state": ["Utah", "Missouri", "Michigan", "Florida", "California"]}

df_addresses = pd.DataFrame(addresses)
df_cities = pd.DataFrame(cities)


def complete_address(df_addresses,df_cities):
    temp = df_addresses['address'].str.split(", ", n = 4, expand = True)
    temp.columns = ['street','city','zip']
    temp = temp.merge(df_cities, on=["city"], how="inner")
    temp["final"] = temp[["street","city","state","zip"]].apply(lambda x: (", ").join(x), axis = 1)
    temp = temp[["final"]].copy()
    temp.columns = ["address"]
    return temp

complete_address(df_addresses,df_cities)
```

</details>

<details>

<summary>PANDAS vs SQL</summary>

Can you tell me what is approximately Windows function equivalent in Pandas?

**Answer**

Windows function in SQL brings row wise calculation capabilities. An approximate equivalent of it can be `transform` in pandas it brings row wise calculation capabilities in Python.

</details>

<details>

<summary>[Forbes] Most Profitable Companies</summary>

[Check this link to practice​.](https://platform.stratascratch.com/coding/10354-most-profitable-companies?code_type=2)

Find the 3 most profitable companies in the entire world. Output the result along with the corresponding company name. Sort the result based on profits in descending order.

**Answer**

```python
forbes_global_2010_2014.head()
t = forbes_global_2010_2014.sort_values('profits', ascending = False)
t.head(3)
```

</details>

<details>

<summary>[Amazon] [DoorDash] Workers With The Highest Salaries</summary>

[Check this link to practice​.](https://platform.stratascratch.com/coding/10353-workers-with-the-highest-salaries?code_type=2)

You have been asked to find the job titles of the highest-paid employees.

Your output should include the highest-paid title or multiple titles with the same salary.

**Answer**

```python
t = pd.merge(worker, title, left_on = 'worker_id', right_on = 'worker_ref_id', how='inner')
t.sort_values('salary', ascending = False, inplace = True)
t['rank'] = t['salary'].rank(method='dense', ascending= False)
t[t['rank']==1]
```

</details>

<details>

<summary>[Meta] Users By Average Session Time</summary>

[Check this link to practice​.](https://platform.stratascratch.com/coding/10352-users-by-avg-session-time?code_type=2)

Calculate each user's average session time. A session is defined as the time difference between a page\_load and page\_exit. For simplicity, assume a user has only 1 session per day and if there are multiple of the same events on that day, consider only the latest page\_load and earliest page\_exit, with an obvious restriction that load time event should happen before exit time event . Output the user\_id and their average session time.

**Answer**

```
# Import your libraries
import pandas as pd
import numpy as np

# Start writing code
entry = facebook_web_log[facebook_web_log['action'].isin(['page_load'])].copy()
exit = facebook_web_log[facebook_web_log['action'].isin(['page_exit'])].copy()
entry['day'] = entry['timestamp'].dt.date
exit['day'] = exit['timestamp'].dt.date
entry = entry.groupby(['user_id','day'], as_index=False).max()
exit = exit.groupby(['user_id','day'], as_index=False).max()

t =pd.merge(entry, exit, on=['user_id','day'], how='inner')
t['diff'] = t['timestamp_y'] - t['timestamp_x']
t.groupby(['user_id']).apply(np.mean)
```

</details>

<details>

<summary>[Google] Activity Rank</summary>

[Check this link to practice​.](https://platform.stratascratch.com/coding/10351-activity-rank?code_type=2)

Find the email activity rank for each user. Email activity rank is defined by the total number of emails sent. The user with the highest number of emails sent will have a rank of 1, and so on. Output the user, total emails, and their activity rank. Order records by the total emails in descending order. Sort users with the same number of emails in alphabetical order. In your rankings, return a unique value (i.e., a unique rank) even if multiple users have the same number of emails. For tie breaker use alphabetical order of the user usernames.

**Answer**

```python
import pandas as pd
import numpy as np

result = google_gmail_emails.groupby(
    ['from_user']).count().to_frame('total_emails').reset_index()
result['rank'] = result['total_emails'].rank(method='first', ascending=False)
result = result.sort_values(by=['total_emails', 'from_user'], ascending=[False, True])
```

</details>

<details>

<summary>[Amazon] Finding User Purchases</summary>

[Check this link to practice​.](https://platform.stratascratch.com/coding/10322-finding-user-purchases?code_type=2)

Write a query that'll identify returning active users. A returning active user is a user that has made a second purchase within 7 days of any other of their purchases. Output a list of user\_ids of these returning active users.

**Answer**

```python
import pandas as pd
import numpy as np
from datetime import datetime

amazon_transactions["created_at"] = pd.to_datetime(amazon_transactions["created_at"]).dt.strftime('%m-%d-%Y')
df = amazon_transactions.sort_values(by=['user_id', 'created_at'], ascending=[True, True])
df['prev_value'] = df.groupby('user_id')['created_at'].shift()
df['days'] = (pd.to_datetime(df['created_at']) - pd.to_datetime(df['prev_value'])).dt.days
result = df[df['days'] <= 7]['user_id'].unique()

```

</details>

<details>

<summary>[Amazon] Monthly Percentage Difference</summary>

[Check this link to practice​.](https://platform.stratascratch.com/coding/10319-monthly-percentage-difference/discussion?code_type=2)

Given a table of purchases by date, calculate the month-over-month percentage change in revenue. The output should include the year-month date (YYYY-MM) and percentage change, rounded to the 2nd decimal point, and sorted from the beginning of the year to the end of the year. The percentage change column will be populated from the 2nd month forward and can be calculated as ((this month's revenue - last month's revenue) / last month's revenue)\*100.

**Answer**

```python
# Import your libraries
import pandas as pd

# Start writing code
sf_transactions.head()
sf_transactions['created_at'] = pd.to_datetime(sf_transactions['created_at'], format='%b')

sf_transactions['year-m'] = sf_transactions['created_at'].dt.to_period('M').astype(str)

df = sf_transactions.groupby('year-m', as_index=False)['value'].sum().sort_values(by='year-m', ascending = True)
df['LM'] = df['value'].shift()
df['prcnt_change'] = (100*(df['value'] - df['LM'])/df['LM']).round(2)
df.head()
```

</details>

<details>

<summary>[Salesforce][Tesla] New Products</summary>

[Check this link to practice​.](https://platform.stratascratch.com/coding/10318-new-products?code_type=2)

You are given a table of product launches by company by year. Write a query to count the net difference between the number of products companies launched in 2020 with the number of products companies launched in the previous year. Output the name of the companies and a net difference of net products released for 2020 compared to the previous year.

**Answer**

```
import pandas as pd
import numpy as np
from datetime import datetime

df_2020 = car_launches[car_launches['year'].astype(str) == '2020']
df_2019 = car_launches[car_launches['year'].astype(str) == '2019']
df = pd.merge(df_2020, df_2019, how='outer', on=[
    'company_name'], suffixes=['_2020', '_2019']).fillna(0)
df = df[df['product_name_2020'] != df['product_name_2019']]
df = df.groupby(['company_name']).agg(
    {'product_name_2020': 'nunique', 'product_name_2019': 'nunique'}).reset_index()
df['net_new_products'] = df['product_name_2020'] - df['product_name_2019']
result = df[['company_name', 'net_new_products']]

```

</details>

<details>

<summary>[Google][Netflix] Top Percentile Fraud</summary>

[Check this link to practice​.](https://platform.stratascratch.com/coding/10303-top-percentile-fraud?code_type=2)

ABC Corp is a mid-sized insurer in the US and in the recent past their fraudulent claims have increased significantly for their personal auto insurance portfolio. They have developed a ML based predictive model to identify propensity of fraudulent claims. Now, they assign highly experienced claim adjusters for top 5 percentile of claims identified by the model. Your objective is to identify the top 5 percentile of claims from each state. Your output should be policy number, state, claim cost, and fraud score.

**Answer**

```python
import pandas as pd
import numpy as np

fraud_score["percentile"] = fraud_score.groupby('state')['fraud_score'].rank(pct=True)
df= fraud_score[fraud_score['percentile']>.95]
result = df[['policy_num','state','claim_cost','fraud_score']]
fraud_score.head()
```

</details>

<details>

<summary>[LinkedIn] Risky Projects</summary>

[Check this link to practice​.](https://platform.stratascratch.com/coding/10304-risky-projects?code_type=2)

Identify projects that are at risk for going overbudget. A project is considered to be overbudget if the cost of all employees assigned to the project is greater than the budget of the project.

You'll need to prorate the cost of the employees to the duration of the project. For example, if the budget for a project that takes half a year to complete is $10K, then the total half-year salary of all employees assigned to the project should not exceed $10K. Salary is defined on a yearly basis, so be careful how to calculate salaries for the projects that last less or more than one year.

Output a list of projects that are overbudget with their project name, project budget, and prorated total employee expense (rounded to the next dollar amount).

HINT: to make it simpler, consider that all years have 365 days. You don't need to think about the leap years.

**Answer**

```python
import pandas as pd
import numpy as np
from datetime import datetime

df = pd.merge(linkedin_projects, linkedin_emp_projects, how = 'inner',left_on = ['id'], right_on=['project_id'])
df1 = pd.merge(df, linkedin_employees, how = 'inner',left_on = ['emp_id'], right_on=['id'])
df1['project_duration'] = (pd.to_datetime(df1['end_date']) - pd.to_datetime(df1['start_date'])).dt.days
df_expense = df1.groupby('title')['salary'].sum().reset_index(name='expense')
df_budget_expense = pd.merge(df1, df_expense, how = 'left',left_on = ['title'], right_on=['title'])
df_budget_expense['prorated_expense'] = np.ceil(df_budget_expense['expense']*(df_budget_expense['project_duration'])/365)
df_budget_expense['budget_diff'] = df_budget_expense['prorated_expense'] - df_budget_expense['budget']
df_over_budget = df_budget_expense[df_budget_expense["budget_diff"] > 0]
result = df_over_budget[['title','budget','prorated_expense']]
result = result.drop_duplicates().sort_values('title')

```

</details>

<details>

<summary>[Microsoft] Premium vs Freemium</summary>

[Check this link to practice​.](https://platform.stratascratch.com/coding/10300-premium-vs-freemium?code_type=2)

Find the total number of downloads for paying and non-paying users by date. Include only records where non-paying customers have more downloads than paying customers. The output should contain 3 columns date, non-paying downloads, paying downloads.

**Answer**

```python
# Import your libraries
import pandas as pd
import numpy as np

# Start writing code
ms_acc_dimension.head()
paying_accs = (ms_acc_dimension[ms_acc_dimension['paying_customer']!='no'])['acc_id'].unique()
paying_cust = (ms_user_dimension[ms_user_dimension['acc_id'].isin(paying_accs)])['user_id'].unique()
# paying_cust
ms_download_facts['p_np'] = np.where(ms_download_facts['user_id'].isin(paying_cust), 'paying downloads', 'non-paying downloads')
ms_download_facts['date'] = ms_download_facts['date'].dt.date
ms_download_facts = ms_download_facts.groupby(['date', 'p_np'], as_index= False)['downloads'].sum()
ms_download_facts = ms_download_facts.pivot(index= 'date', columns = 'p_np', values='downloads')
ms_download_facts['date'] = ms_download_facts.index
ms_download_facts.head()

```

</details>

<details>

<summary>[Microsoft][Apple] Most Popular Client_Id</summary>

[Check this link to practice​.](https://platform.stratascratch.com/coding/2029-the-most-popular-client_id-among-users-using-video-and-voice-calls?code_type=2)

Select the client\_ids based on a count of the number of users who have at least 50% of their events from the following list: 'video call received', 'video call sent', 'voice call received', 'voice call sent'.

**Answer**

2 versions of the answer are given with slight difference

```python
import pandas as pd
import numpy as np

events_list = ['video call received', 'video call sent', 'voice call received', 'voice call sent']

fact_events['valid_even_count'] = np.where(fact_events['event_type'].isin(events_list), 1,0)

fact_events = fact_events.groupby(['client_id']).apply(lambda x: x['valid_even_count'].sum()/x['event_id'].count()).rename("abc").reset_index()
# rename will set a name to the column else it will be blank and reset_index will
# bring back the client_id as a column
fact_events.head()
```

```python
# Import your libraries
import pandas as pd
import numpy as np

# Start writing code

fact_events['event_cat'] = np.where(fact_events['event_type'].isin(['video call received', 'video call sent', 'voice call received', 'voice call sent']),"valid","others")
fact_events = fact_events.groupby(['user_id','event_cat'], as_index=False)['time_id'].count()
fact_events['prnct'] = fact_events.groupby(['user_id']).transform(lambda x:100*x/x.sum())
fact_events.head()
```

</details>


# Statistics

## Questions

<details>

<summary>JOIN Dataframes</summary>

Can you tell me the ways in which 2 pandas data frames can be joined?

**Answer**

* merge() is used to combine two (or more) dataframes on the basis of values of common columns (indices can also be used, use left\_index=True and/or right\_index=True)
* concat() is used to append one (or more) dataframes one below the other (or sideways, depending on whether the axis option is set to 0 or 1).
* join() is used to merge 2 dataframes on the basis of the index; instead of using merge() with the option left\_index=True we can use join().

</details>

<details>

<summary>[GOOGLE] Normal Distribution</summary>

Write a function to generate N samples from a normal distribution and plot the histogram.

**Answer**

```python
import numpy as np
import numpy as np
import matplotlib.pyplot as plt

def generate_samples(n, mean, std):
  """Generates N samples from a normal distribution with mean `mean` and standard deviation `std`."""
  return np.random.normal(mean, std, n)

def plot_histogram(samples):
  """Plots a histogram of the given samples."""
  plt.hist(samples, bins=100)
  plt.show()

# Generate 100 samples from a normal distribution with mean 0 and standard deviation 1.
samples = generate_samples(1000, 0, 1)

# Plot the histogram of the samples.
plot_histogram(samples)
```

</details>

<details>

<summary>[UBER] Bernoulli trial generator</summary>

Given a random Bernoulli trial generator, write a function to return a value sampled from a normal distribution.

**Answer**

```python
# *Solution recieved from the community via [merge request](https://github.com/dipranjan/dsinterviewqns/pull/5)*

import numpy as np
import pandas as pd
import matplotlib.pyplot as plt

# straightforward using the central limit theorem.

p = .5
n = 10000

# returns standard normal output via the central limit theorem
def standard_normal_output(p,n):
    bernoulli_mean = p
    bernoulli_variance = p*(1-p)
    bernoulli_std = abs(np.sqrt(bernoulli_variance))
    sample = np.random.binomial(size = n, n = 1, p = p)
    return (sample.mean() - bernoulli_mean)/(bernoulli_std/np.sqrt(n))

# now we plot this output 10000 times to indeed show it is a standard normal distribution
def plot_output():
    outputs=[]
    for i in range(0,n):
        outputs.append(standard_normal_output(p=p,n=n))
    num_bins = 20
    plt.hist(outputs, bins=num_bins, facecolor='blue', alpha=0.5)
    plt.show() 
plot_output()
```

</details>

<details>

<summary>[PINTEREST] Interquartile Distance</summary>

Given an array of unsorted random numbers (decimals) find the interquartile distance.

**Answer**

```python
# Interquartile distance is the difference between first and third quartile

# first let's generate a list of random numbers

import random
import numpy as np

li = [round(random.uniform(33.33, 66.66), 2) for i in range(50)]
print(li)

qtl_1 = np.quantile(li,.25)
qtl_3 = np.quantile(li,.75)

print("Interquartile distance: ", qtl_1 - qtl_3)
```

</details>

<details>

<summary>[GENENTECH] Imputing the median</summary>

Write a function cheese\_median to impute the median price of the selected California cheeses in place of the missing values. You may assume at least one cheese is not missing its price.

**Answer**

```python
import pandas as pd

cheeses = {"Name": ["Bohemian Goat", "Central Coast Bleu", "Cowgirl Mozzarella", "Cypress Grove Cheddar", "Oakdale Colby"], "Price" : [15.00, None, 30.00, None, 45.00]}

df_cheeses = pd.DataFrame(cheeses)
```

</details>

<details>

<summary>Show the Central Limit Theorem</summary>

In order to do this we will start with a non-normal distribution example the uniform distribution. Next, we will sample that distribution and get the mean of the sample, we will do this repeatedly. As per the central limit theorem the plot of the means will resemble a normal distribution.

**Answer**

```python
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt

def sampling(n):
    # Create sample from uniform distribution
    sample = np.random.uniform(size=n, low = 1, high = 6)
    return sample.mean() #3.5 subtract the population mean if you want mean=0 for the normal distribution

# now we sample this 10000 times to indeed show it is a standard normal distribution
def plot_output(n):
    outputs=[]
    for i in range(0,n):
        outputs.append(sampling(30))
    num_bins = 20
    plt.hist(outputs, bins=num_bins, facecolor='blue', alpha=0.5)
    plt.title("Sample")
    plt.show() 

plot_output(10000)

```

&#x20;![](https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FaENqAS4gI9LNPIOCxsbk%2Fimage.png?alt=media\&token=306d4c77-1526-4ba1-b834-ec6385b5d8e6)

</details>


# NLP

<details>

<summary>[OKCUPID] Term Frequency</summary>

Say you are given a text document in the form of a string with the following sentences:

**Input**

`document = "I have a nice car with a nice tires"`

**Output**

`{ "I":0.11, "have":0.11, "a":0.22, "nice":0.22, "car": 0.11, "with":0.11, "tires":0.11 }`&#x20;

Write a program in python to determine the TF (term frequency) values for each term of this document.

**Answer**

{% code overflow="wrap" %}

```python
# term freq = freq of term 't' in doc 'd' / total terms in 'd'

def tf(doc):
    temp_list = doc.split(" ") # split terms
    final_dict = {} # store final dictionary with term freq
    for i in temp_list:
        if i not in final_dict.keys(): # check if term is already present in dict
            final_dict[i] = round(temp_list.count(i)/len(temp_list),2) # calculating tf
    return final_dict

tf("I have a nice car with a nice tires")
```

{% endcode %}

</details>


# Algorithms from scratch

Often companies ask to code different Algorithms from scratch as a part of their craft demo round.

The general steps are as follows:

1. Get 3 things: The formula for the algorithm, the cost function, the derivatives to be used for gradient descent.&#x20;
2. Initialize the weights and bias.
3. In the fit function set the loop to update the weights in each iteration as per the cost function optimization algorithm.
4. Create a predict function to predict the values.
5. Create a test dataset to test if your function works properly.
6. Plot the results.


# Linear Regression

The formula: $$y^{\wedge} = wx+b$$

The cost function, $$MSE = J(w,b) = 1/N \sum\_{i=1}^{n}(y\_i-(wx\_i+b))^2$$

The derivatives for Gradient Descent:

&#x20;$$df/dw = 1/N \sum\_{i=1}^{n}-2x\_i(y\_i-(wx\_i+b))$$

$$df/db = 1/N \sum\_{i=1}^{n}-2(y\_i-(wx\_i+b))$$

```python
import numpy as np

# Calculate the R2 score
def r2_score(y_true, y_pred):
    corr_matrix = np.corrcoef(y_true, y_pred)
    corr = corr_matrix[0, 1]
    return corr ** 2


class LinearRegression:
    def __init__(self, learning_rate=0.001, n_iters=1000):
        self.lr = learning_rate
        self.n_iters = n_iters
        self.weights = None
        self.bias = None

    def fit(self, X, y):
        n_samples, n_features = X.shape

        # init parameters
        self.weights = np.zeros(n_features) # this can be random as well
        self.bias = 0

        # gradient descent
        for _ in range(self.n_iters):
            y_predicted = np.dot(X, self.weights) + self.bias
            # compute gradients
            dw = (1 / n_samples) * np.dot(X.T, (y_predicted - y))
            db = (1 / n_samples) * np.sum(y_predicted - y)

            # update parameters
            self.weights -= self.lr * dw
            self.bias -= self.lr * db

    def predict(self, X):
        y_approximated = np.dot(X, self.weights) + self.bias
        return y_approximated


# Testing
if __name__ == "__main__":
    # Imports
    import matplotlib.pyplot as plt
    from sklearn.model_selection import train_test_split
    from sklearn import datasets

    def mean_squared_error(y_true, y_pred):
        return np.mean((y_true - y_pred) ** 2)

    # We are making the dataset to test out algorithm
    X, y = datasets.make_regression(
        n_samples=100, n_features=1, noise=20, random_state=4
    )

    X_train, X_test, y_train, y_test = train_test_split(
        X, y, test_size=0.2, random_state=1234
    )

    regressor = LinearRegression(learning_rate=0.01, n_iters=1000)
    regressor.fit(X_train, y_train)
    predictions = regressor.predict(X_test)

    mse = mean_squared_error(y_test, predictions)
    print("MSE:", mse)

    accu = r2_score(y_test, predictions)
    print("Accuracy:", accu)

    y_pred_line = regressor.predict(X)
    cmap = plt.get_cmap("viridis")
    fig = plt.figure(figsize=(4, 3))
    m1 = plt.scatter(X_train, y_train, color=cmap(0.9), s=10)
    m2 = plt.scatter(X_test, y_test, color=cmap(0.5), s=10)
    plt.plot(X, y_pred_line, color="black", linewidth=2, label="Prediction")
    plt.show()
```

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FlKRvylWI4kfkhTxFTpHg%2Fimage.png?alt=media&amp;token=8b62bb63-7606-4dcd-b2a3-a44130a3691b" alt=""><figcaption></figcaption></figure>


# Logistic Regression

```python
import numpy as np
class LogisticRegression:
    def __init__(self, learning_rate=0.001, n_iters=1000):
        self.lr = learning_rate
        self.n_iters = n_iters
        self.weights = None
        self.bias = None

    def fit(self, X, y):
        n_samples, n_features = X.shape

        # init parameters
        self.weights = np.zeros(n_features)
        self.bias = 0

        # gradient descent
        for _ in range(self.n_iters):
            # approximate y with linear combination of weights and x, plus bias
            linear_model = np.dot(X, self.weights) + self.bias
            # apply sigmoid function
            y_predicted = self._sigmoid(linear_model)

            # compute gradients
            dw = (1 / n_samples) * np.dot(X.T, (y_predicted - y))
            db = (1 / n_samples) * np.sum(y_predicted - y)
            # update parameters
            self.weights -= self.lr * dw
            self.bias -= self.lr * db

    def predict(self, X):
        linear_model = np.dot(X, self.weights) + self.bias
        y_predicted = self._sigmoid(linear_model)
        y_predicted_cls = [1 if i > 0.5 else 0 for i in y_predicted]
        return np.array(y_predicted_cls)

    def _sigmoid(self, x):
        return 1 / (1 + np.exp(-x))


# Testing
if __name__ == "__main__":
    # Imports
    from sklearn.model_selection import train_test_split
    from sklearn import datasets

    def accuracy(y_true, y_pred):
        accuracy = np.sum(y_true == y_pred) / len(y_true)
        return accuracy

    bc = datasets.load_breast_cancer()
    X, y = bc.data, bc.target

    X_train, X_test, y_train, y_test = train_test_split(
        X, y, test_size=0.2, random_state=1234
    )

    regressor = LogisticRegression(learning_rate=0.0001, n_iters=1000)
    regressor.fit(X_train, y_train)
    predictions = regressor.predict(X_test)

    print("LR classification accuracy:", accuracy(y_test, predictions))
```


# PySpark

A brief overview of PySpark

#### What is PySpark?

PySpark is the Python API for Apache Spark, a powerful open-source engine designed for large-scale data processing. It allows you to leverage Spark’s capabilities using Python, making it easier to work with big data.

#### Why is PySpark Necessary?

PySpark is essential for several reasons:

1. **Handling Big Data**: Traditional tools struggle with large datasets, but PySpark processes them efficiently in a distributed computing environment.
2. **Speed and Performance**: PySpark’s in-memory processing makes it faster than disk-based frameworks like Hadoop MapReduce, crucial for real-time data analysis.
3. **Versatility**: It supports both structured and unstructured data from various sources.
4. **Advanced Analytics**: PySpark includes built-in libraries for machine learning and graph processing.
5. **Python Compatibility**: It allows Python users to easily transition and collaborate.

#### Differences Between PySpark and Pandas

* **Scale**: Pandas is ideal for smaller datasets that fit into memory on a single machine, while PySpark is designed for distributed computing, handling massive datasets across multiple machines.
* **Performance**: PySpark can process large-scale data faster due to its distributed nature, whereas Pandas is more efficient for smaller datasets.
* **API and Functionality**: Both offer DataFrame APIs, but PySpark’s API is built for distributed processing, providing scalability and parallelism.
* **Use Cases**: Use Pandas for data manipulation and analysis on smaller datasets. Use PySpark for big data analytics, machine learning, and real-time data processing.

#### When to Use PySpark vs. When Not to Use It

**Use PySpark When:**

* You need to process large datasets that exceed the memory capacity of a single machine.
* Real-time data processing and analysis are required.
* You need to leverage distributed computing for performance and scalability.
* Your team is familiar with Python and you want to integrate with the Python ecosystem.

**Avoid PySpark When:**

* Your datasets are small and can be handled efficiently by Pandas.
* You require maximum performance and are comfortable using Scala or Java, which can offer better optimization with Spark[4](https://www.sparkcodehub.com/pyspark-vs-spark-comparison).
* The overhead of Python-Java interoperability might impact performance for your specific use case[5](https://granulate.io/blog/understanding-pyspark-features-ecosystem-optimization/).

#### Operational Efficiency of PySpark

PySpark’s operational efficiency can be optimized through several techniques:

* **Data Serialization and Caching**: Using efficient data formats like Parquet and caching frequently accessed data.
* **Optimized Execution Plans**: Leveraging the Spark DataFrame API for automatic optimization.
* **Resource Management**: Properly allocating memory and CPU resources, and tuning configurations.
* **Avoiding Expensive Operations**: Minimizing shuffles and using efficient transformations.


# Overview

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2F1Yh5R4ApuGHymUrguHiK%2Fimage.png?alt=media&amp;token=de4db32f-d650-4fa2-bca1-b9bbefa557db" alt=""><figcaption><p>(<a href="https://neptune.ai/wp-content/uploads/2022/10/MLOps-tools-landscape-horizontal-upd.png">Source</a>)</p></figcaption></figure>


# GIT

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FDZfzpJ1qYPQ5Ur5HB6bg%2FGit%20cartoon%20(1).png?alt=media&amp;token=1541c320-49b5-4727-8f1f-4aa25e49cc67" alt=""><figcaption><p><a href="https://xkcd.com/1597/">XKCD</a></p></figcaption></figure>

<table data-header-hidden><thead><tr><th width="241"></th><th></th></tr></thead><tbody><tr><td><strong>Command</strong></td><td><strong>Function</strong></td></tr><tr><td>git init</td><td>Create empty Git repo in specified directory. Run with no arguments to initialize the current directory as a git repository.</td></tr><tr><td>git clone</td><td>Clone repo located at onto local machine. Original repo can be located on the local filesystem or on a remote machine via HTTP or SSH.</td></tr><tr><td>git config user. name</td><td>Define author name to be used for all commits in current repo. Devs commonly use --global flag to set config options for current user.</td></tr><tr><td>git add</td><td>Stage all changes in for the next commit. Replace with a to change a specific file.</td></tr><tr><td>git commit -m ""</td><td>Commit the staged snapshot, but instead of launching a text editor, use as the commit message.</td></tr><tr><td>git status</td><td>List which files are staged, unstaged, and untracked.</td></tr><tr><td>git log</td><td>Display the entire commit history using the default format. For customization see additional options.</td></tr><tr><td>git diff</td><td>Show unstaged changes between your index and working directory.</td></tr><tr><td>git remote add</td><td>Create a new connection to a remote repo. After adding a remote, you can use as a shortcut for in other commands.</td></tr><tr><td>git fetch</td><td>Fetches a specific , from the repo. Leave off to fetch all remote refs.</td></tr><tr><td>git pull</td><td>Fetch the specified remote's copy of current branch and immediately merge it into the local copy.</td></tr><tr><td>git push</td><td>Push the branch to , along with necessary commits and objects. Creates named branch in the remote repo if it doesn't exist.</td></tr><tr><td>git revert</td><td>Create new commit that undoes all of the changes made in , then apply it to the current branch.</td></tr><tr><td>git reset</td><td>Remove from the staging area, but leave the working directory unchanged. This unstages a file without overwriting any changes.</td></tr><tr><td>git clean -n</td><td>Shows which files would be removed from working directory. Use the -f flag in place of the -n flag to execute the clean.</td></tr><tr><td>git diff HEAD</td><td>Show difference between working directory and last commit.</td></tr><tr><td>git diff --cached</td><td>Show difference between staged changes and last commit</td></tr><tr><td><strong>GIT RESET</strong></td><td></td></tr><tr><td>git reset</td><td>Reset staging area to match most recent commit, but leave the working directory unchanged.</td></tr><tr><td>git reset --hard</td><td>Reset staging area and working directory to match most recent commit and overwrites all changes in the working directory.</td></tr><tr><td>git reset</td><td>Move the current branch tip backward to , reset the staging area to match, but leave the working directory alone.</td></tr><tr><td>git reset --hard</td><td>Same as previous, but resets both the staging area &#x26; working directory to match. Deletes uncommitted changes, and all commits after .</td></tr><tr><td><strong>GIT REBASE</strong></td><td></td></tr><tr><td>git rebase -i</td><td>Interactively rebase current branch onto . Launches editor to enter commands for how each commit will be transferred to the new base.</td></tr><tr><td><strong>GIT PULL</strong></td><td></td></tr><tr><td>git pull --rebase</td><td>Fetch the remote's copy of current branch and rebases it into the local copy. Uses git rebase instead of merge to integrate the branches.</td></tr><tr><td><strong>GIT PUSH</strong></td><td></td></tr><tr><td>git push --force</td><td>Forces the git push even if it results in a non-fast-forward merge. Do not use the --force flag unless you're absolutely sure you know what you're doing.</td></tr><tr><td>git push --all</td><td>Push all of your local branches to the specified remote.</td></tr><tr><td>git push --tags</td><td>Tags aren't automatically pushed when you push a branch or use the --all flag. The --tags flag sends all of your local tags to the remote repo.</td></tr></tbody></table>


# Feature Store

Many machine learning and AI models work best on summaries of raw data called *features*. These features structure information into a form that makes it easier to train algorithms.

A simple feature might involve transforming a raw date into a weekday or weekend, both of which might be better predictors of behavior than a raw date number. Other kinds of features can be more complex and require intricate calculations across many data streams. A feature store provides a place to organize the most popular features so they can be reused across projects rather than redone from scratch every time they're used.

A feature store can increase automation, improve productivity by promoting sharing and reuse, reduce technical debt in software code, ensure consistency in calculations and provide governance, auditability and lineage for regulatory compliance, according to David Sweenor, senior director of product marketing at data science tools company Alteryx. However, a feature store isn't ideal for every company. Smaller ones may struggle with the overhead required to create and maintain a feature store. Companies may also struggle with reusing features across departments.

### What are the benefits of a feature store?

A feature, as it relates to data science, is any variable that can be used for analytics. Simple examples include name, age, sex, zip code and amount. These raw variables are transformed through a process known as *feature engineering* to yield better predictions. For example, a date could be transformed into a day of the week, a day of the year or a holiday.

A feature store enables a data scientist to create this transformation once rather than having each data scientist recreate the same features repeatedly. This ensures consistency since everyone is using the exact same transformation as part of their models. It also reduces the need to insert the same algorithm within code. If a company decides to change a complex feature, a feature store enables them to change it once and propagate it across all models that use it. Otherwise, someone would have to manually edit all the models using that feature.

Since processing these data is very expensive, and these data are slow-changing, it makes sense to process them once every hour or day and store the features into a feature store for hundreds of teams to use \[machine learning] ML to solve their business problems.


# Basics

{% hint style="info" %}
Since solving any reasonable SQL problem requires a combination of all the topics covered here, hence it becomes difficult to seggregate problems based on one topic alone. So for SQL we are creating a dedicated [**Problems**](/sql/problems) section. Theoritical and Basic questions will still be under their dedicated sections.
{% endhint %}

If you are new to SQL, the video playlist below[️](https://www.youtube.com/watch?v=7GVFYt6_ZFM\&list=PL08903FB7ACA1C2FB) is one of the most comprehensive materials of SQL that is available on the internet. Please go through the portions of interest to you.

{% embed url="<https://www.youtube.com/watch?list=PL08903FB7ACA1C2FB&v=7GVFYt6_ZFM>" %}

## DDL

DDL stands for Data Definition Language. These commands are used to change the structure of a database and database objects. Some common commands are as follows:

* **CREATE:** is used to create the database or its objects (like table, index, function, views, store procedure and triggers).

  The generic format of the create table is as follows:

  ```sql
  CREATE DATABASE databasename;

  CREATE TABLE table_name (
  column1 datatype CONSTRAINT,
  column2 datatype CONSTRAINT,
  column3 datatype CONSTRAINT); 
  ```

  Using constraints, we can specify the limit on the type of data that can be stored in a particular column in a table. Some of the constraints are:

  * **NOT NULL:** This constraint tells that the value of a column cannot be null.
  * **UNIQUE:** This constraint tells that the values in any row of a column must not be repeated.
  * **PRIMARY KEY:** A primary key is a field which can uniquely identify each row in a table. We can say that PRIMARY KEY is combination of NOT NULL and UNIQUE constraints. A table can have only one field as primary key.
  * **FOREIGN KEY:** Foreign Key is a field in a table which uniquely identifies each row of another table. This field points to primary key of another table. This usually creates a relation between the tables.
  * **CHECK:** This constraint helps to ensure that the value stored in a column meets a specific condition.
  * **DEFAULT:** This constraint specifies a default value for the column when no value is specified by the user.
* **DROP:** is used to delete databases or tables.
* **ALTER:** is used to alter the structure of the database. You can use this to add/modify/delete a column.
* **TRUNCATE:** The TRUNCATE TABLE statement is used to delete the data inside a table, *but not the table itself.*
* **RENAME:** is used to rename an object existing in the database.

## Referential Integrity

Referential integrity is a property of data stating references within it are valid. In the context of relational databases, it requires every value of one attribute (column) of a relation (table) to exist as a value of another attribute (column) in a different (or the same) relation (table). For referential integrity to hold in a relational database, any column in a base table that is declared a foreign key can contain either a null value, or only values from a parent table's primary key or a candidate key. In other words, when a foreign key value is used it must reference a valid, existing primary key in the parent table. For instance, deleting a record that contains a value referred to by a foreign key in another table would break referential integrity. Some relational database management systems (RDBMS) can enforce referential integrity, normally either by deleting the foreign key rows as well to maintain integrity, or by returning an error and not performing the delete. Which method is used may be determined by a referential integrity constraint defined in a data dictionary.

## UNION & UNION ALL

* UNION removes duplicate rows, UNION ALL does not
* UNION has to perform a distinct sort to remove duplicates, hence is a little slow
* UNION combines rows of two tables, JOIN combines columns

## NULL

There are 3 ways to take care of NULL values:

Let's take an example suppose we want to fill the Manager name as 'No Manager' if an employee does have a manager assigned in the manager column. This can be done using:

* `ISNULL(manager, 'No Manager')`
* `CASE WHEN manager IS NULL THEN 'No Manager' ELSE manager END as manager`
* The other option is `COALESCE` but it essentially takes the first non-null value out of the passed columns

### Best practices

* Use aliases to shorten and simplify names. This can make your code more readable and usable.&#x20;
* Avoid using the SELECT \* statement. This can cause unexpected results and performance issues.&#x20;
* Use indexes to speed up queries. Indexes allow the database to quickly find entries that match specific criteria.&#x20;
* Use EXIST() instead of COUNT(). EXIST() only runs until it finds the first entry for a record in the table. This can save time and computing power.&#x20;
* Avoid using SELECT DISTINCT for large tables. This clause removes duplicate entries, but it is computationally expensive.&#x20;
* Use WHERE instead of HAVING.&#x20;
* Use ORDER BY to ensure the ordering of your results.
* Use STARTS\_WITH instead of LIKE.&#x20;
* Avoid using multiple nested queries.
* Avoid using unnecessary subqueries.  Instead, rewrite queries with outer joins.

### NOLOCK

* The NOLOCK hint allows SQL to read data from tables by ignoring any locks and therefore not get blocked by other processes.
* This can improve query performance by removing the blocks, but introduces the possibility of *dirty reads*.

```sql
SELECT * FROM Person.Contact WITH (NOLOCK) WHERE ContactID < 20 
```

If you want to learn more check[ this link.](https://www.mssqltips.com/sqlservertip/2470/understanding-the-sql-server-nolock-hint/#:~:text=The%20NOLOCK%20hint%20allows%20SQL,the%20possibility%20of%20dirty%20reads.)

## Questions

<details>

<summary>[MICROSOFT] Having vs Where</summary>

Can you elborate on the differences between HAVING and WHERE clause in a SQL query?

**Answer**

The major differences between HAVING and WHERE are as follows:

* WHERE can be used with Select, Insert, Update, Delete statements. HAVING can only be used with Select statements
* WHERE filters rows before aggregation, HAVING filters after that

Performance wise there is not much of a difference, the best practice is to filter out unwanted rows as early as possible.

</details>

<details>

<summary>DROP vs TRUNCATE vs DELETE</summary>

The DELETE command deletes one or more existing records from the table in the database. The DROP Command drops the complete table from the database. The TRUNCATE Command deletes all the rows from the existing table, leaving the row with the column names.

</details>

Can you tell the difference between T-SQL and PL SQL

**Answer**

| \*\*T-SQL\*\*                                                                                                                                                        | \*\*P/L SQL\*\*                                                                                                            |
| -------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------- |
| T-SQL   was originally developed by Sybase and now owned by Microsoft. Hence, it   works with Microsoft SQL Server only.                                             | P/L SQL was developed   by Oracle and works with Oracle only.                                                              |
| All   database objects like Tables/Views/Procedures are internally organized by   database names. Users are allowed access to a specific database and its   objects. | All database objects   are organized in Schemas. Users are allowed access to certain schemas via   roles and permissions.  |
| More   focussed on Microsoft SQL Server and its functions only.                                                                                                      | P/L SQL is more   versatile and encompassing, E.g. you can send Emails and access Web   Pages.                             |
| Allows   B-Tree Indexes only.                                                                                                                                        | Allows B-Tree, Bitmap,   Domain, and Partitioned Indexes.                                                                  |
| Provides   a greater degree of control on how an application works in Microsoft SQL   Server and is easy to learn.                                                   | Enhances the power of   plain SQL, and is considered to be more powerful and holistic.                                     |
| Limited   and patchy support for use of Cursors.                                                                                                                     | Powerful support for   use of Cursors, the underlying file/data organization supports it.                                  |
| No   direct support for Object-Oriented programming. Only supported though 3rd   party ORM   tools.                                                                  | Full support for Object-Oriented programming via Function Overloading, Data Encapsulation, etc.                            |
| Explicit   error handling capabilities using Try-Catch blocks.                                                                                                       | Provides error handling via checking for Exceptions only.                                                                  |
| Provides bulk insert and data loading.                                                                                                                               | No explicit bulk inserts.                                                                                                  |
| Allows reading data from an external sequential file and you can use the BULK   statement to fine-tune the external reads.                                           | Indirectly allows reading data from an external sequential file but no specific language   constructs to ease the process. |
| No packages to ease re-use.                                                                                                                                          | Allows grouping of Procedures into Packages, which can be reused.                                                          |
| No   need for Subqueries and DELETE and UPDATE are improvised accordingly.                                                                                           | Needs Subqueries to read data from another table.                                                                          |
| Provides WAITFOR construct to wait for a certain period of elapsed time.                                                                                             | No such direct construct in P/L SQL.                                                                                       |
| No support for Arrays.                                                                                                                                               | Supports Arrays.                                                                                                           |


# Joins

{% hint style="info" %}
Since solving any reasonable SQL problem requires a combination of all the topics covered here, hence it becomes difficult to seggregate problems based on one topic alone. So for SQL we are creating a dedicated [**Problems**](/sql/problems) section. Theoritical and Basic questions will still be under their dedicated sections.
{% endhint %}

Joins are best explained using Venn diagrams

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FG3bSI2fXbOqvdzg9mwQS%2Fimage1.png?alt=media&amp;token=1a257769-3189-419e-b9f1-cfb036df3775" alt=""><figcaption><p>TSQL JOIN Types, <a href="https://stevestedman.com/2015/05/tsql-join-types-poster-version-4-1/">Reference</a></p></figcaption></figure>

Remember that in case of multiple joins each single join produces a **single derived table** that is then joined to the next table and so on.

## Questions

<details>

<summary>Address of People</summary>

**Reference -** [**Leetcode**](https://leetcode.com/problems/combine-two-tables/)

<pre><code><strong>Table: Person
</strong>
| Column Name | Type    |
|-------------|---------|
| PersonId    | int     |
| FirstName   | varchar |
| LastName    | varchar |

PersonId is the primary key column for this table.

Table: Address


| Column Name | Type    |
|-------------|---------|
| AddressId   | int     |
| PersonId    | int     |
| City        | varchar |
| State       | varchar |

AddressId is the primary key column for this table.
</code></pre>

Write a SQL query for a report that provides the following information for each person in the Person table, regardless of if there is an address for each of those people:

FirstName, LastName, City, State

**Answer**

```sql
select a.FirstName, a.LastName, b.City, b.State
from Person a
left join Address b
on a.PersonId = b.PersonID
```

</details>


# Temporary Datasets

A summary of Temp Table vs Table variable vs CTE

{% hint style="info" %}
Since solving any reasonable SQL problem requires a combination of all the topics covered here, hence it becomes difficult to seggregate problems based on one topic alone. So for SQL we are creating a dedicated [**Problems**](/sql/problems) section. Theoretical and Basic questions will still be under their dedicated sections.
{% endhint %}

**Reference:** [📖Explanation](https://www.red-gate.com/simple-talk/databases/sql-server/t-sql-programming-sql-server/sql-server-cte-basics/)

Sometimes in order to ease things out you need to create a temporary version of the data for either viewing it or for running some further calculations to it. Now if it is just that you want to run the same query daily and get the results you might as well save it as a `VIEW` . **You can consider view as a SQL statement saved with a name**.

While it is very important to know what a `VIEW` is, however in most cases the ask will be to solve some question which for which you need to create a temporary dataset and run queries on that. You can use Common Table Expression or CTE for that. It uses the `WITH` keyword.

### **CTE**&#x20;

CTE can be of 2 types:

* A **recursive CTE** is one that references itself within that CTE. The recursive CTE is useful when working with hierarchical data because the CTE continues to execute until the query returns the entire hierarchy
* A **non-recursive CTE** is one that does not reference itself within the CTE. Nonrecursive CTEs tend to be simpler than recursive CTEs

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FsGtRslUEZnSOgvz1qfgt%2Fimage4.png?alt=media&amp;token=49075f64-0655-43c7-8383-4bc49f64ec45" alt=""><figcaption><p>The above image shows a case where we need to use CTE to answer a question</p></figcaption></figure>

### Table Variable

This biggest difference is that a CTE can only be used in the current query scope whereas a temporary table or table variable can exist for the entire duration of the session allowing you to perform many different DML operations against them.

### Compare Temp Table, Table Variable and CTE

([Source](https://www.c-sharpcorner.com/UploadFile/ff2f08/tips-to-improve-sql-database-performance/))

```sql
-- CTE
WITH t (customerid, lastorderdate) AS 
 (SELECT customerid, max(orderdate) 
  FROM sales.SalesOrderHeader
  GROUP BY customerid)
SELECT * 
FROM sales.salesorderheader soh
INNER JOIN t ON soh.customerid=t.customerid AND soh.orderdate=t.lastorderdate
GO

-- Temporary table
CREATE TABLE #temptable (customerid [int] NOT NULL PRIMARY KEY, lastorderdate [datetime] NULL);

INSERT INTO #temptable
SELECT customerid, max(orderdate) as lastorderdate 
FROM sales.SalesOrderHeader
GROUP BY customerid;

SELECT * 
FROM sales.salesorderheader soh
INNER JOIN #temptable t ON soh.customerid=t.customerid AND soh.orderdate=t.lastorderdate

DROP TABLE #temptable
GO

-- Table variable
DECLARE @tablevariable TABLE (customerid [int] NOT NULL PRIMARY KEY, lastorderdate [datetime] NULL);

INSERT INTO @tablevariable
SELECT customerid, max(orderdate) as lastorderdate 
FROM sales.SalesOrderHeader
GROUP BY customerid;

SELECT * 
FROM sales.salesorderheader soh
INNER JOIN @tablevariable t ON soh.customerid=t.customerid AND soh.orderdate=t.lastorderdate
GO
```

Looking at [SQL Profiler](https://www.mssqltips.com/sql-server-tip-category/83/profiler-and-trace/) results from these queries (each were run 10 times and averages are below) we can see that the CTE just slightly outperforms both the temporary table and table variable queries when it comes to overall duration. The CTE also uses less CPU than the other two options and performs fewer reads (significant fewer reads that the table variable query).

| Query Type     | Reads  | Writes | CPU | Duration (ms) |
| -------------- | ------ | ------ | --- | ------------- |
| CTE            | 1378   | 0      | 47  | 497           |
| Temp table     | 2146   | 51     | 109 | 544           |
| Table variable | 133748 | 51     | 297 | 578           |

<br>


# Windows Functions

{% hint style="info" %}
Since solving any reasonable SQL problem requires a combination of all the topics covered here, hence it becomes difficult to seggregate problems based on one topic alone. So for SQL we are creating a dedicated [**Problems**](/sql/problems) section. Theoretical and Basic questions will still be under their dedicated sections.
{% endhint %}

**Reference:** [📖Explanation](https://www.red-gate.com/simple-talk/sql/t-sql-programming/introduction-to-t-sql-window-functions/), [🔫Playground](https://dbfiddle.uk/?rdbms=sqlserver_2017\&fiddle=6379904805d1f465cc0f6ea33fc3c0d6)

Window (also, windowing or windowed) functions perform a calculation over a set of rows. I like to think of “looking through the window” at the rows that are being returned and having one last chance to perform a calculation. The window is defined by the OVER clause which determines if the rows are partitioned into smaller sets and if they are ordered. They allow you to add your favourite aggregate function to a non-aggregate query. Similar to Transform is pandas group by clause.

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FIQh9c437gjTDcCLnMoic%2Fimage5.png?alt=media&amp;token=22c38379-89c7-4819-aaed-a5aee96eb1fe" alt=""><figcaption><p>GROUP vs WINDOW</p></figcaption></figure>

## Common Windows Functions

* **Ranking functions**
  * **ROW\_NUMBER:** is used to add unique row numbers to a partition or to the entire result set. It has the ability to turn non-unique rows into unique rows
  * **RANK:** it will give numbers same as row\_number just that same data will get same rank
  * **DENSE\_RANK:** doesnot skip and rank number
  * **NTILE:** It assigns bucket numbers to the rows instead of row numbers or ranks

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2F9Nmfvq5kn3hHhX9l2EwB%2Fimage2.png?alt=media&amp;token=2e5e5c69-4edc-4269-8963-88a07163bcea" alt=""><figcaption><p>Ranking functions, check playground to work with this</p></figcaption></figure>

* **Offset functions**
  * **LAG:** the function allows you to pull columns or expressions from a row before the current row
  * **LEAD:** the function allows you to pull columns or expressions from a row after the current row
  * **FIRST\_VALUE:** the functions allows you to return values from the first row of the partition
  * **LAST\_VALUE:** the functions allows you to return values from the last row of the partition
* **Statistical functions** – **PERCENT\_RANK:** returns the percentage of rows that rank lower than the current row, its formula is $$\frac{\text{Rank} -1}{\text{Row count} -1}$$
  * **CUME\_DIST:** cumulative distribution, returns the exact rank, its formula is $$\frac{\text{Rank}}{\text{Row count}}$$
  * **PERCENTILE\_DISC & PERCENTILE\_CONT:** these two work in the opposite way. Given a percent rank, find the value at that rank. They differ in that PERCENTILE\_DISC will return a value that exists in the set while PERCENTILE\_CONT will calculate an exact value if none of the values in the set falls precisely at that rank. You can use PERCENTILE\_CONT to calculate a median by supplying 0.5 as the percent rank. For example, which temperature ranks at 50% in St. Louis?

{% hint style="info" %}
One problem is you cannot add window functions to the WHERE clause. But certain Data Bases like TeraData, Snowflake, Databricks, Big Query, etc. have support for something called QUALIFY. Let's take an example to understand the difference:

In traditional SQL you will write the below SQL:

`select * from (SELECT Product, Region, Revenue, ROW_NUMBER() OVER (PARTITION BY Region ORDER BY Revenue DESC) as rn FROM Sales ) where rn=1`

Here's how you can use the QUALIFY keyword to achieve this:

`SELECT Product, Region, Revenue FROM Sales QUALIFY ROW_NUMBER() OVER (PARTITION BY Region ORDER BY Revenue DESC) = 1;`
{% endhint %}

{% hint style="info" %}
Remember you CAN include window function and GROUP BY in the same statement, in such case the WINDOW function is evaluated *after* GROUP BY. *As an example, below we can get both the user2 count at a user1 level along with total number of users in the table in the same query:*

<pre class="language-sql"><code class="lang-sql">select user1 
, 100*(<a data-footnote-ref href="#user-content-fn-1">CAST(count(distinct(user2)</a>) as FLOAT)/CAST (<a data-footnote-ref href="#user-content-fn-2">count(*) over () as FLOAT</a>)) as popularity_percent
from facebook_friends
group by user1
</code></pre>

{% endhint %}

[^1]: GROUP BY

[^2]: WINDOW function


# Time

{% hint style="info" %}
Since solving any reasonable SQL problem requires a combination of all the topics covered here, hence it becomes difficult to segregate problems based on one topic alone. So, for SQL we are creating a dedicated [**Problems**](/sql/problems) section. Theoretical and Basic questions will still be under their dedicated sections.
{% endhint %}

{% hint style="info" %}
The below commands are for SQL Server, a lot of databases tend to do the Time part of things a little differently, many of them has additional commands or does not support a few commands. Please be cautious as the commands might change depending on the DB you are using.
{% endhint %}

**Reference:** [📖Explanation](https://www.sqlshack.com/learn-sql-sql-server-date-and-time-functions/)

In most product companies like Google, Salesforce, Facebook, etc. in short any company that deals with user interaction will store data about user interaction and more often than not will ask questions which leverages the querying or manipulating time part of stored data. While time part is not very difficult to solve people face problems in that they donot remember the right command to do what they are trying, so here we will share the list of common commands along with their intended usage.

To get system values:

* `GETDATE()` or `CURRENT_TIMESTAMP` : return the datetime value from the server where SQL Server runs
* `SYSDATETIME()` : same as above but with more precision

Next we will see how can we extract specific feature from a date column:

* `YEAR('2021/07/06 14:08:52')` or `DATEPART(YEAR, '2021/07/06 14:08:52')` or `DATENAME(YEAR, '2021/07/06 14:08:52')` : will return the year
* Same command can be used for MONTH, DAY, HOUR, MINUTE, SECOND
* DATENAME for Month will return the name of the month

The next 3 functions are used to create date or modify/combine dates:

* `DATEFROMPARTS(year, month, day)` takes a year, month and day as integer values and creates 1 date out of them.
* `DATEADD(date_part, interval, date)` takes 3 arguments and returns a date that is interval (date\_parts) number of given units (date\_part) distant from the given date (date)
* `DATEDIFF(date_part, start_date, end_date)` returns the number of units (date\_part) between end\_date and start\_date (end\_date – start\_date).


# Functions & Stored Proc

{% hint style="info" %}
Since solving any reasonable SQL problem requires a combination of all the topics covered here, hence it becomes difficult to segregate problems based on one topic alone. So, for SQL we are creating a dedicated [**Problems**](/sql/problems) section. Theoretical and Basic questions will still be under their dedicated sections.
{% endhint %}

Functions are calculated values that cannot make lasting modifications to SQL Server's environment (i.e., no INSERT or UPDATE statements allowed).

If a function provides a scalar value, it may be used inline in SQL queries; if it returns a result set, it can be joined on.

Functions must return a value and cannot change the data they receive as parameters, as defined by computer science (the arguments). Functions can't modify anything and must have at least one parameter. They also have to return a result. Stored procedures don't need a parameter, may modify database objects, and don't have to return a result.

Stored procedures are used to connect SQL queries in a transaction and to communicate with the outside world.

```sql
CREATE PROCEDURE SelectAllCustomers @City nvarchar(30), @PostalCode nvarchar(10)
AS
SELECT * FROM Customers WHERE City = @City AND PostalCode = @PostalCode

-- To run the Stored Proc
EXEC SelectAllCustomers @City = 'London', @PostalCode = 'WA1 1DP';
```

<details>

<summary>N-th Highest Salary</summary>

**Reference -** [**Leetcode**](https://leetcode.com/problems/nth-highest-salary/) Write a SQL query to get the Nth highest salary from the Employee table.

```
| Id | Salary |
|----|--------|
| 1  | 100    |
| 2  | 200    |
| 3  | 300    |
```

For example, given the above Employee table, the query should return 200 as the second highest salary (N =2). If there is no second highest salary, then the query should return null.

**Answer**

Multiple solutions are possible only one approach is given below for reference

```sql
	CREATE FUNCTION getNthHighestSalary(N INT) RETURNS INT
	BEGIN
	      DECLARE temp INT;
	      SET temp = N-1;
	  RETURN (      
	      Select DISTINCT Salary from Employee
	      Order by Salary Desc
	      LIMIT 1 Offset temp      
	  );
	END
```

</details>


# Index

An index is a disk-based structure linked to a table or view that facilitates quicker row retrieval. A table or view’s table or view’s columns are used to create keys in an index. These keys are kept in a structure (B-tree) that enables SQL Server to quickly and effectively locate the row or rows that correspond to the key values.

| CLUSTERED INDEX                                                                                          | NON-CLUSTERED INDEX                                                                                                                                        |
| -------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------- |
| A clustered index is faster.                                                                             | A non-clustered index is slower.                                                                                                                           |
| The clustered index requires less memory for operations.                                                 | A non-Clustered index requires more memory for operations.                                                                                                 |
| In a clustered index, the clustered index is the main data.                                              | In the Non-Clustered index, the index is the copy of data.                                                                                                 |
| A table can have only one clustered index.                                                               | A table can have multiple non-clustered indexes.                                                                                                           |
| The clustered index has the inherent ability to store data on the disk.                                  | A non-Clustered index does not have the inherent ability to store data on the disk.                                                                        |
| Clustered index store pointers to block not data.                                                        | Non-Clustered index storescontainThe non-Clustered both value and a pointer to the actual row that holds data.                                             |
| In Clustered index leaf nodes are actual data itself.                                                    | In Non-Clustered index leaf nodes are not the actual data itself rather they only contain included columns.                                                |
| In a Clustered index, Clustered key defines the order of data within a table.                            | In a Non-Clustered index, the index key defines the order of data within the index.                                                                        |
| A Clustered index is a type of index in which table records are physically reordered to match the index. | A Non-Clustered index is a special type of index in which the logical order of the index does not match the physical stored order of the rows on the disk. |
| The size of The primary clustered index is large.                                                        | The size of the non-clustered index is compared relativelyThe composite is smaller.                                                                        |
| Primary Keys of the table by default are clustered indexes.                                              | The composite key when used with unique constraints of the table act as the non-clustered index.                                                           |


# Performance Tuning

([Source](https://www.c-sharpcorner.com/UploadFile/ff2f08/tips-to-improve-sql-database-performance/))

#### 1. Do not use \* with select statement

SQL Server converts \* into all column names of the table before query execution. Instead of doing this, pass the name of the columns that are really required in query results.

```sql
--Bad practice

SELECT * FROM Table1
```

```sql
--Good practice

SELECT Column1, Column2, Column3 FROM Table1
```

#### 2. Use EXISTS instead of IN

EXISTS returns true if a sub query contains any rows else returns false. EXISTS is faster than the IN Query because for the IN Query SQL first collects all the data of the sub query. With EXISTS, the sub query does not produce actual query results. EXISTS is faster than IN only when sub query result is very large. IN is faster than EXISTS, when sub query result is very small

```sql
--Bad practice
SELECT  Column1, Column2, Column3 FROM Table1 WHERE Column1 IN (SELECT Column1 FROM Table2)
```

```sql
--Good practice
SELECT  Column1, Column2, Column3 FROM Table1 WHERE EXISTS (SELECT Column1 FROM Table2 Where Table2.Column1 = Table1.Column1)
```

#### 3. Select Appropriate Data Type of table columns

The selection of appropriate data types may help us to improve the SQL query performance. For example, I have an Employee Table and it has a code field. The length of the code may vary from 3 to 8. In this case, instead of selecting CHAR (8) we can use the VARCHAR (8) data type. Try to choose the smallest data type that works for each column. Try choosing an appropriate data type of a column to avoid explicit and implicit conversions, because both are costly in terms of time to take for conversion.

#### 4. Use proper join type

Select proper join type and join order. An Outer join is a more expensive process than the inner join. An Outer join must do everything that an Inner join does plus extra work for null extending results.

#### 5. Use Indexed Views

When we have a SQL Query that contains multiple joins between tables that do not change frequently (example lookup table), we can define an indexed view for better performance. An Index view is a view stored physically like a table. The Index view is updated by SQL Server itself when a table is modified that is used as part of the index view.&#x20;

#### 6. Do not use Count (\*)

Do not use Count (\*) when you require a row count. Instead of this, you can use Count (1) or Count (ColumnName).

```sql
--Bad practice
SELECT  COUNT(*) FROM Table1
```

```sql
--Good practice
SELECT  COUNT(1) FROM Table1
SELECT  COUNT(Column1) FROM Table1
```

#### 7. Avoid use of cursors

A Cursor is used to perform a function row by row. The Cursor forces the database engine to repeatedly fetch the rows, managing the locks and transmit the results. Forward-only and read-only cursors are faster and uses the least resources. If there is a primary key on a table then we can use a while loop. Try to avoid the use of a cursor on temp tables.

#### 8. Use a Table variable or CTE (Common Table Expression) instead of Temp Table whenever possible

Temp tables are stored physically in TempDB and they are permanent tables that are deleted after the session ends. CTE and Temp variables are created within memory. Note that CTE is not a replacement of a Table variable and a Temp table.

#### 9. Use SET NOCOUNT ON in Stored Procedure

The SET NOCOUNT ON statement prevents SQL Server from sending a message after each statement in a Stored Procedure. These messages are sent using a DONE\_IN\_PROC token. Each message contains the number of the row affected the row by the last executed SQL statement.

```sql
--Example
CREATE PROCEDURE [dbo].[MyTestProcedure]
(
   @para1 INT
)
AS
BEGIN
   SET NOCOUNT ON
   -- Body of stored procedure.
END
```

#### 10. Do not use "SP\_" prefix with any user define Stored Procedure name

In SQL Server, the master database has a Stored Procedure with the "sp\_" prefix, so SQL Server always looks first in the master database. If we use the "sp\_" prefix for any user define Stored Procedure and put it in a database other than master, the master database is still checked first.

#### 11. Use Stored Procedure for Complex Query and frequently used query

A Stored Procedure in SQL Server is a group of T-SQL statements. By default a Stored Procedure is precompiled, in other words a stored procedure is compiled when it executes the first time and also creates an execution plan that for subsequent calls of the Stored Procedure because the query processor does not need to create a new plan hence it takes less time to execute.

#### 12. Use Parameterized Query

The SQL Server saves execution plan for parameterized queries. This allows it to be reused on later execution.

#### 13. Use Try...Catch Block whenever required

A TRY...CATCH block is used for error handling in T-SQL statements. Use TRY...CATCH blocks whenever you work with transactions because if an error occurs during a transaction then it may cause a deadlock.

```sql
--Example
BEGIN TRAN Test
BEGIN TRY
      -- Do something
      COMMIT TRAN Test
END TRY
BEGIN CATCH
      ROLLBACK TRAN Test
END CATCH
```

#### 14. Select appropriate life of transaction (Try to Avoid long running transaction)

Keep your transactions as short as possible because locks are held during the transaction. Do not forget to commit or roll back the transaction before session end.

#### 15. Use proper Isolation level

The Transaction isolation level defines the degree to which one transition must be isolated from a resource and data modification made by another transaction. Isolation level is responsible for deciding how long read locks are held on a data row. A lower isolation level increases the ability to access the same data at the same time for many users but it increases the chance of a dirty read or loss of updated data. A higher isolation level decreases this type of concurrency issues but requires more database resources. So it is necessary to choose the appropriate isolation level for our transactions.

#### 16. Do not use function in WHERE Clauses

When a function is used with a select statement, it is not a bad thing because it returns powerful data with each row. But a function used with a WHERE clause forces SQL Server to do a table scan to determine the correct data.

#### 17. Try to avoid Expensive operators such as "LIKE", "NOT LIKE", Not equal to (<> or! =)

The Like Operator used to determine whether a specific character or string matches a specified pattern. This pattern may include regular characters or wildcard characters. The "Like" Operator always causes a table scan. This type of table scan is very expensive. Operators such as <> or NOT LIKE are also very costly in terms of performance.

#### 18. Define Relationships and Constraints whenever required

Primary key and foreign key relationships help us to ensure that we write optimal queries. SQL Server creates an optimal execution plan if the primary key and foreign key constraints are defined in the database schema.

#### 19. Create Index on All Foreign Keys

Mostly Foreign keys are used in joins, so if an index creates on foreign keys always beneficial.

#### 20. Create Index whenever required

Creating Indexes are one of the best ways to improve the performance of a SQL query. A table scan is done when no index is available on a table to help a SQL query. Some points must be kept in mind when creating an index.

* Do not create an index on columns that contain a high number of null values.
* Do not create indexes on columns that are frequently manipulated.
* Keep an index key short.
* Create indexes with a minimum percentage of duplicated values.

#### 21. Avoid Table and Index Scans

A Table and Index scan is an expensive process and it becomes more expensive the data size is greater. Select tables for scans that have few rows and a cluster index scan may be an effective option for some queries.

#### 22. Use multiple small indexes rather than a few wide indexes

We can create multiple indexes per table in SQL Server. Small or Narrow indexes provide more options than a wide composite index. Note that in an index, statistics are only kept for the first column of a composite index, so multiple single-column indexes ensure statistics are also kept on that column.

#### 23. Create indexes on columns used in "WHERE", "ORDER BY", "GROUP BY" and "DISTINCT" clauses

Create an index on columns used in a WHERE clause and used in aggregate operations, such as GROUP BY, DISTINCT, ORDER BY, MIN, MAX and so on.

#### 24. Remove unused indexes

Remove all unused indexes. An index is needed to be maintained even if they are not used.

### 25. Use ITW (Index Tuning Wizard)

The Index Tuning Wizard is used to obtain guidance and tips on index options. This is good way to create efficient indexes.

#### 26. Avoid long actions in triggers

A Trigger is always a part of a DML statement (INSERT, UPDATE or DELETE) and calling the transaction. A long-running action may cause a lock to be held longer than intended hence the result is the blocking of other queries. We can use a message queue to do long-running actions (accomplish this task asynchronously).

#### 27. Estimate the Hash Joins

Hash joins are used for many types of set matching operations such as inner join, outer join (left, right and full), intersection and union. In the absence of an index, a Hash Join is the best option. Evaluate the execution plan if you have many of hash joins to determine CPU usage.


# Problems

{% hint style="info" %}
If you want to have some hands on practice without the hassle of installing and setting up the required softwares in your local machine 🔫DB Fiddle provides free SQL sandbox. In a lot of problems below prebuilt sandbox links are already provided to refer but it is always recommended that you setup your personal sandbox to play around.
{% endhint %}

<details>

<summary>[Leetcode] Second highest salary</summary>

*For a similar problem with different approach check Nth highest salary problem*

Write a SQL query to get the second highest salary from the Employee table.

```
| Id | Salary |
|----|--------|
| 1  | 100    |
| 2  | 200    |
| 3  | 300    |
```

For example, given the above Employee table, the query should return 200 as the second highest salary. If there is no second highest salary, then the query should return null.

**Answer**

Multiple solutions are possible two approaches are given below for reference

```sql
SELECT
(SELECT DISTINCT(Salary)
FROM Employee
ORDER BY Salary DESC
LIMIT 1 OFFSET 1) 
AS SecondHighestSalary
```

I feel the below solution is more complete as it gives you the ability to handle edge cases if id is also needed and there are multiple employees with same salary:

```sql
with cte(
select 
salary,
dense_rank() over(order by salary desc) as rank
from Employee)

select salary as SecondHighestSalary
from cte where rank = 2
```

</details>

<details>

<summary>[Leetcode] Rank Scores</summary>

**Reference -** [**Leetcode**](https://leetcode.com/problems/rank-scores/)

Write a SQL query to rank scores. If there is a tie between two scores, both should have the same ranking. Note that after a tie, the next ranking number should be the next consecutive integer value. In other words, there should be no "holes" between ranks.

```
| Id | Score |
|----|-------|
| 1  | 3.40  |
| 2  | 3.65  |
| 3  | 4.00  |
| 4  | 3.50  |
| 5  | 4.00  |
| 6  | 3.65  |
```

For example, given the above Scores table, your query should generate the following report (order by highest score):

```
| score | Rank    |
|-------|---------|
| 4.00  | 1       |
| 4.00  | 1       |
| 3.95  | 2       |
| 3.65  | 3       |
| 3.65  | 3       |
| 3.40  | 4       |
```

**Answer**

The tie resolving method which is being asked in the question is called Dense Rank, if we use Rank it will have "holes"

```sql
select 
Score, dense_rank() over(order by score desc) as Rank
from Scores
```

</details>

<details>

<summary>[CHEWY] 2nd Highest score</summary>

```
| Id | subject | marks |
|---:|---------|------:|
|  1 | Maths   |    30 |
|  1 | Phy     |    50 |
|  1 | Chem    |    85 |
|  2 | Maths   |    90 |
|  2 | Phy     |    50 |
|  2 | Chem    |    85 |
```

Select the second highest mark for each student.

**Answer**

{% code overflow="wrap" %}

```sql
with CTE as(
	select *, rank() over(partition by Id order by marks desc) as Rank from tablename
)
select Id, subject, marks from CTE where Rank = 1
```

{% endcode %}

</details>

<details>

<summary>[Leetcode] Consecutive Numbers</summary>

**Reference -** [**Leetcode**](https://leetcode.com/problems/consecutive-numbers/)

Write an SQL query to find all numbers that appear at least three times consecutively.

Return the result table in any order.

Input:

Logs table:

```
| Id | Num |
|----|-----|
| 1  | 1   |
| 2  | 1   |
| 3  | 1   |
| 4  | 2   |
| 5  | 1   |
| 6  | 2   |
| 7  | 2   |
```

Result table:

```
| ConsecutiveNums |
|-----------------|
| 1               |
```

1 is the only number that appears consecutively for at least three times.

**Answer**

Multiple solutions are possible, one of them is given below

{% code overflow="wrap" %}

```sql
with a(Num,NextNum,SecondNextNum ) as(

	SELECT   Num
	         , LEAD(Num, 1) OVER (ORDER BY Id) AS NextNum
	         , LEAD(Num, 2) OVER (ORDER BY Id) AS SecondNextNum
	      FROM Logs
	      
	)

	select distinct(Num) as ConsecutiveNums from a
	where
	Num = NextNum
	and Num = SecondNextNum
```

{% endcode %}

</details>

<details>

<summary>[SALESFORCE] User Growth</summary>

[**🔫Playground**](https://dbfiddle.uk/?rdbms=sqlserver_2017\&fiddle=ce4ded37fa37bf552365c18cb7840c3b)

Given you have user data for 2 accounts for 2 months. Calculate the growth rate of users in each account where growth rate is defined as unique users in month 2 divided by unique users in month 1.

```
| date_details | account_id | user_id |
|--------------|------------|---------|
| 2021-01-01   | U1         | A1      |
| 2021-01-01   | U1         | A2      |
| 2021-01-01   | U1         | A3      |
| 2021-01-01   | U1         | A4      |
| 2021-02-01   | U1         | A1      |
| 2021-02-01   | U1         | A2      |
| 2021-02-01   | U1         | A3      |
| 2021-02-01   | U1         | A4      |
| 2021-02-01   | U1         | A5      |
| 2021-01-01   | U2         | A1      |
| 2021-01-01   | U2         | A2      |
| 2021-01-01   | U2         | A3      |
| 2021-02-01   | U2         | A1      |
| 2021-02-01   | U2         | A2      |
```

**Answer**

{% code overflow="wrap" %}

```sql
with cte as (
	select account_id, count(distinct(user_id)) as unique_user, MONTH(date_details) as user_month from tablename
	group by account_id, MONTH(date_details)
	)

select a.account_id,month_2,month_1,
cast((month_2/month_1)as float) as growth  from 
(select account_id, unique_user as month_1
from cte where user_month = 1)a
left join
(select account_id, unique_user as month_2
from cte where user_month = 2)b
on (a.account_id = b.account_id)
```

{% endcode %}

</details>

<details>

<summary>[SALESFORCE] Month over Month Revenue</summary>

[**🔫Playground**](https://dbfiddle.uk/?rdbms=sqlserver_2017\&fiddle=72569e574e670b477d2f62fdfc4276ca)

You have 2 tables:

* transactions: date, prod\_id, quantity
* products: prod\_id, price

Calculate the month over month revenue, example month over month revenue for month2 is month2\_Revenue- month1\_Revenue

**Answer**

```sql
with cte as(
	select MONTH(a.date_details) as month, sum(b.price*a.qty) as Rev
	from transactions a
	inner join products b
	on a.prod_id = b.prod_id
	group by MONTH(a.date_details)
	),

	cte2 as(
	select month, Rev, lag(Rev,1) over(order by month) as prev_month
	from cte
	)

select month, (Rev-Prev_month) as extra_rev  from cte2
where 
prev_month is not null
```

</details>

<details>

<summary>[SALESFORCE] Retention Rate</summary>

([Source](https://platform.stratascratch.com/coding/2053-retention-rate?code_type=5\&utm_source=blog\&utm_medium=click\&utm_campaign=YT+description+link))

Find the monthly retention rate of users for each account separately for Dec 2020 and Jan 2021. Retention rate is the percentage of active users an account retains over a given period of time. In this case, assume the user is retained if he/she stays with the app in any future months. For example, if a user was active in Dec 2020 and has activity in any future month, consider them retained for Dec. You can assume all accounts are present in Dec 2020 and Jan 2021. Your output should have the account ID and the Jan 2021 retention rate divided by Dec 2020 retention rate.

*Note: I believe the official solution provided on the website is not correct as of 25-10-2023*

**Answer**

```sql
-- ret rate for jan and dec = active in future/ active in dec
-- ret rate of jan 21 ret rate/ dec 20 ret rate at acc_id level

with cte as(
select a.account_id, a.user_id,dec_active,jan_active,last_actv    from 
(select account_id, user_id
, case when date between '2020-12-01' and '2020-12-31' then 1 else 0 end as dec_active
, case when date between '2021-01-01' and '2021-01-31' then 1 else 0 end as jan_active
from sf_events)a

inner join
(
select user_id, max(date) as last_actv
from sf_events group by user_id) b
on a.user_id = b.user_id),

cte2 as(
select 
account_id
, sum(case when dec_active = 1 and last_actv > '2020-12-31' then 1 else 0 end)/sum(dec_active) *1.0 as dec_retn
, sum(case when jan_active = 1 and last_actv > '2021-01-01' then 1 else 0 end)/sum(jan_active) *1.0 as jan_retn
 from cte
 
group by account_id)

select account_id, jan_retn/nullif(dec_retn,0) * 1.0 as retention from cte2

```

</details>

<details>

<summary>[SALESFORCE] Employee earning more than their manager</summary>

**Reference -** [**Leetcode**](https://leetcode.com/problems/employees-earning-more-than-their-managers/)

Write an SQL query to find the employees who earn more than their managers.

```
| Id | Name  | Salary | ManagerId |
|---:|-------|-------:|----------:|
|  1 | Joe   |  70000 |         3 |
|  2 | Henry |  80000 |         4 |
|  3 | Sam   |  60000 |           |
|  4 | Max   |  90000 |           |
```

Output will be : Joe

**Answer**

{% code overflow="wrap" %}

```sql
with cte as(
	Select a.Name as Employee, b.Name as Manager, a.Salary as Emp_Sal, b.Salary as Man_Salary
	from Employee a
	inner join Employee b
	on a.ManagerId = b.id)

Select Employee from cte where Emp_Sal > Man_Salary
```

{% endcode %}

</details>

<details>

<summary>[Leetcode] Highest Salary in each Department</summary>

**Reference -** [**Leetcode**](https://leetcode.com/problems/department-highest-salary/)

Write an SQL query to find employees who have the highest salary in each of the departments.

(../SQL/images/image3.PNG)

**Answer**

```sql
with cte as(
	select Name, Salary, DepartmentId,
	RANK() over(Partition by DepartmentId order by salary desc) as Rank
	from Employee
)
		    
select b.Name as Department, a.Name as Employee, a.Salary
from cte a
inner join Department b
on a.DepartmentId = b.Id
where a.Rank = 1
```

</details>

<details>

<summary>[AMAZON] Cumulative Sum</summary>

Given a users table, write a query to get the cumulative number of new users added by day, with the total reset every month.

[🔫Playground](https://dbfiddle.uk/?rdbms=sqlserver_2017\&fiddle=516b59f188aaf8c5c1296143d1b13bcd)

**Answer**

{% code overflow="wrap" %}

```sql
Select Created_date
,SUM(Count(Id)) OVER(partition by month(Created_date) order by Created_date) as Total_users
from users
group by Created_date
```

{% endcode %}

</details>

<details>

<summary>Tree Structure Labeling</summary>

[🔫Playground](https://dbfiddle.uk/?rdbms=sqlserver_2019\&fiddle=922326a37527cc50e64fe896c6d70608) Input:

```
| node | parent |
|------|--------|
| 1    | 2      |
| 2    | 5      |
| 3    | 5      |
| 4    | 3      |
| 5    | NULL   |
```

Write SQL such that you label each node as a “leaf”, “inner” or “Root” node, such that for the nodes above the output is:

Output:

```
| node | label |
|------|-------|
| 1    | Leaf  |
| 2    | Inner |
| 3    | Inner |
| 4    | Leaf  |
| 5    | Root  |
```

**Answer**

{% code overflow="wrap" %}

```sql
select node,
case
when parent is null then 'Root'
when node not in (select parent from tree where parent is not null) then 'Leaf'
else 'Inner'
end as label
from tree
```

{% endcode %}

</details>

<details>

<summary>[FACEBOOK] Binning data</summary>

[🔫Playground](https://dbfiddle.uk/?rdbms=sqlserver_2019\&fiddle=813e35800a6e3d955f57ac7e6c7c2e91) Input:

```
| id | length |
|---:|-------:|
|  1 |      4 |
|  2 |      3 |
|  3 |      7 |
|  4 |      8 |
|  5 |      9 |
|  6 |    110 |
|  7 |    113 |
```

Bin the videos into groups of 5 secs each

Output:

```
| bucket  | count |
|---------|------:|
| 0-5     |     2 |
| 5-10    |     3 |
| 110-115 |     2 |
```

Another similar question was asked in Facebook but instead of video length the ask was to write a SQL query to create a histogram of number of comments per user in the month of January 2020. As the approach is similar hence not including it here.

**Answer**

{% code overflow="wrap" %}

```sql
with cte as(
select id, length, ((CAST((FLOOR(length)/5)*5 as varchar)) +'-'+(CAST((FLOOR(length)/5)*5+5 as varchar))) as bucket
from video_view_details
)
select bucket, count(id) as 'count'
from cte
group by bucket
order by len(bucket) asc, bucket desc
```

{% endcode %}

</details>

<details>

<summary>[DROPBOX] Closest SAT Scores</summary>

[🔫Playground](https://dbfiddle.uk/?rdbms=sqlserver_2019\&fiddle=b7d44c3f8ec7caaab70d11fd7502f65f)

Given a table of students and their SAT test scores, write a query to return the two students with the closest test scores with the score difference. Assume a random pick if there are multiple students with the same score difference.

Input:

```
| id | score |
|---:|------:|
|  1 |    40 |
|  2 |    35 |
|  3 |    70 |
|  4 |    80 |
```

Output:

```
| id | other_student | diff |
|---:|--------------:|-----:|
|  1 |             2 |    5 |
```

**Answer**

```sql
with cte as(
select id
  , LEAD(id, 1) over(order by score desc) as other_student
  , score - LEAD(score, 1) over(order by score desc) as diff
from 
score)

select top 1 * from cte
where diff is not null
order by diff
```

</details>

<details>

<summary>[AMAZON] Average Distance between Cities</summary>

[🔫Playground](https://dbfiddle.uk/?rdbms=sqlserver_2017\&fiddle=876bddb9c3e31a222ce95fbb8eef7a00)

You are given a table with varying distances from various cities. How do you find the average distance between each of the pairs of the cities?

```
| scity  | dcity  | distance |
|--------|--------|---------:|
| City A | City B |       30 |
| City A | City B |       32 |
| City B | City A |       29 |
| City A | City C |       40 |
| City C | City A |       41 |
```

Output:

```
| city1  | city2  |         distance |
|--------|--------|-----------------:|
| City A | City C |             40.5 |
| City A | City B | 30.3333333333333 |
```

Another variant of this question is

"Write a query to create a new table, named flight routes, that displays unique pairs of two locations?"

**Answer**

```sql
select 
  (case when scity < dcity then scity else dcity end) city1, 
  (case when scity < dcity then dcity else scity end) city2,
  avg(cast(distance as float)) distance
from tablename 
group by
  (case when scity < dcity then scity else dcity end), 
  (case when scity < dcity then dcity else scity end)
order by avg(cast(distance as float)) desc
```

</details>

<details>

<summary>[AMAZON] Duplicate Rows</summary>

Given a users table, write a query to return only its duplicate rows

**Answer**

Multiple solutions are possible only one approach is given below for reference

Let's assume there are 2 columns: id, name

```sql
SELECT *, COUNT(*) FROM userstable
GROUP BY id, name
HAVING COUNT(*) > 1
```

</details>

<details>

<summary>[INTUIT] Product Average</summary>

**transactions table**

```
|   column   |   type   |
|:----------:|:--------:|
| id         | integer  |
| user_id    | integer  |
| created_at | datetime |
| product_id | integer  |
| quantity   | integer  |
```

**products table**

```
| column |   type  |
|:------:|:-------:|
| id     | integer |
| name   | string  |
| price  | float   |
```

Given a table of transactions and products, write a query to return the product id, product price, and average transaction price of all products with price greater than the average transaction price.

**Answer**

[Source](https://www.interviewquery.com/questions/zipcode-average?ref=question_email)

```sql
with cte as (
select 
    t.product_id, 
    avg(t.quantity * p.price) as avg_trans_price
from transactions t
inner join products p
	on t.product_id = p.id
group by t.product_id
)

select 
    p.id as product_id, 
    p.price as product_price, 
    c.avg_trans_price as avg_price
from products p
inner join cte c
	on p.id = c.product_id
where p.price > c.avg_trans_price
```

</details>

<details>

<summary>[INTUIT] Data Analyst Interview Question</summary>

Given the following tables:

&#x20;<img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2F5BSFLekMCYRBkpvJvVtV%2Fimage.png?alt=media&amp;token=1ce2cef5-7862-49e9-bca9-26a137734f4a" alt="" data-size="original">&#x20;

&#x20;Where

Experiments is a table in which we store whether a user is part of an experiment and if so whether they are in test or control (assume there is only one test variant per experiment).  The fields are:

●     user\_id - There are many users each of whom can be in many experiments

●     assignment\_ts - timestamp of when the user was allocated to the experiment.  A user is only allocated once per experiment.

●     experiment\_id - An experiment has many users

●     experiment\_assignment - Whether the user is in test or control.  Assignments are immutable and there is only one assignment per user/experiment combo.

&#x20;Subscriptions is a table of subscription related events.  For each user, there will always be a trial start event however there will only be a subscription start event if the user subscribes.  Assume a given user can only have one trial start and at most one subscription start.  The subscription can start at any time after the trial start and times for either event type are captured in event\_ts.

**Questions**

Write queries to produce the following:

1. When did each experiment start?  Use the first instance of an experiment assignment to either test or control for an experiment to equate to when the experiment started.  Results should look like:

&#x20;![](https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FUeVC43SnKHb6OwMBH8Pf%2Fimage.png?alt=media\&token=e7ebd9df-0dd4-4307-b23c-c6cda870522b)

2. How long did each experiment last, expressed in days?  Assume the last instance of an experiment assignment to test or control for an experiment to equate to when the experiment ended.  Results should look like:

&#x20;![](https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FlMfk544R2m0WswPmxBYl%2Fimage.png?alt=media\&token=653003cd-3abe-4a8b-a559-a0fa278d5b04)

3. How many users are in test and control for each experiment?  Result should look like:

![](https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FSzklq6oW4bTRLnYLuIRV%2Fimage.png?alt=media\&token=34527e97-d317-42e9-ba76-6ef70613e62d)

4. What is the conversion rate by experiment assignment for each experiment?  A conversion is any user for whom there is a subscription start event in addition to the trial start event (all users have a trial start event).  If a user is in multiple experiments at the same time, it’s ok to count them towards the conversion rate of each experiment.  We also want to only return one row per experiment.  Result should look like:

![](https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2F1jVN79U0BTtEGC09zQsZ%2Fimage.png?alt=media\&token=892e64ee-768d-4db9-b216-1107b7a82012)

5\) For each experiment\_id, rank and list first 3 user\_ids who subscribed to the product. Output should look like:

&#x20;![](https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2F6YzTzmRaZwvgrV1jGWuP%2Fimage.png?alt=media\&token=dcfbc34c-f5d4-42c5-abd5-234c9eb0d4a2)

</details>

<details>

<summary>[INTUIT] Employer EINs</summary>

🔫[Playground](https://dbfiddle.uk/wXAiFvAe)

We're given a table called ***employers*** that consists of a *user\_id, year*, and *employer EIN* label. Users can have multiple employers dictated by the different EIN labels.

Write a query to add a flag to each user if they've added a new employer in the current year.

Example:&#x20;

```
# employer
# 
# user_id    year    employer_ein
# 34323      2018    A 
# 34323      2018    B
# 34323      2018    C
# 34323      2017    F
# 34323      2017    A
# 34323      2017    B
# 
# 86323      2018    A
# 86323      2018    B
# 86323      2018    C
# 86323      2017    B
# 
# 98787      2018    A
# 98787      2018    B
# 98787      2017    B
# 98787      2017    A

# Output
# user_id    year    new_ein_flag
# 34323      2018      1
# 86323      2018      1
# 98787      2018      0 
```

**Answer**

This problem is a little trickier than it looks at the outset

```sql
with cte as
  (
select user_id, ein,
case when year = max(year) over() then 1 else 0 end as 'CY_flg',
case when year = max(year) over() -1 then 1 else 0 end as 'LY_flg'
from users
),

cte2 as(
select user_id, ein,  sum(CY_flg) as 'CY_flg', sum(LY_flg) as 'LY_flg' from cte
group by user_id, ein)

select b.user_id, isnull(newflg,0) as new_ein_count  from
  (select user_id, isnull(count(ein),0) as newflg from cte2
where CY_flg = 1 and Ly_flg = 0
  group by user_id)a
right join
(select distinct(user_id) from users) b
on a.user_id = b.user_id
```

</details>


# Excel Basics

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FhIwJYeXbrHFYGSK8dSJD%2Fimage1.png?alt=media&amp;token=6819dbe2-491a-429f-b9c6-d61c5cd62015" alt=""><figcaption></figcaption></figure>

***

## Questions

<details>

<summary>Find Revenue</summary>

From the source below can you find the Revenue the specified account?&#x20;

<img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2F6rIjsWcTfj9qdtfO9HFX%2Fimage2.png?alt=media&amp;token=323b2c86-deeb-4d6d-94c0-715ee5f2904b" alt="" data-size="original">

**Answer**

This can be solved with a simple VLOOKUP

```
=VLOOKUP(F3,B2:D12,3,FALSE)
```

</details>

<details>

<summary>Find Customer Number</summary>

From the source below can you find the Customer Number corresponding to the Account Name?

&#x20;![](https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FKkcX6YMPK8hlJzsWWPzp%2Fimage3.png?alt=media\&token=c828173c-d744-496c-b9ed-142f1f49b3bb)

**Answer**

VLOOKUP won't work as Customer Num is to the LEFT of the Account Name. We need INDEX MATCH [📖Explanation](https://exceljet.net/index-and-match)

```
=INDEX(A2:D12,MATCH(F7,B2:B12),1)
```

</details>

<details>

<summary>Total Revenue</summary>

From the source below can you find the total revenue per sales rep? ![](https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2F4UsQ6AsIEHOYLcOsTrWR%2Fimage4.png?alt=media\&token=de91d6d7-cd80-4a5a-b9b4-a0f3af8066aa)

**Answer**

This can be solved with a simple SUMIF. For the first row the answer is given below. It will be similar for other rows.

```
=SUMIF(C3:C12,"="&F11,D3:D12)
```

</details>


# Data Manipulation

{% hint style="info" %}
Excel, as a product, always remains under active development from Microsoft. With new releases, new features are brought in whereas old ones are discarded. Due to this there might be changes to the solutions mentioned below depending on the version that you are using.
{% endhint %}

<details>

<summary>SUM of digits</summary>

Can you write a formula to generate the SUM of all digits in a cell?

**Answer**

![](https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2F2j1qXa7IFsmtZs38f6ts%2Fimage5.png?alt=media\&token=f96a3fd1-85b2-49ec-aa0a-9c1a045899ff)

To use when you are sure that there are only digits in the column:

`=SUMPRODUCT(--MID(B2,ROW(INDIRECT("1:"&LEN(B2))),1))`

But if there are other characters too use this:

`=SUMPRODUCT((LEN(B3)-LEN(SUBSTITUTE(B3,ROW(1:9),"")))*ROW(1:9))`

</details>

<details>

<summary>DISTINCT &#x26; Duplicates</summary>

Given the data below, please answer the following questions

This is a 3-part question:

* Given a table of data how do you tell if it has duplicates?
* Create a table with distinct values from this
* Can you do a conditional duplicate check on this table?

```markup
| Region | ID |
|--------|----|
| A      | 1  |
| B      | 2  |
| C      | 3  |
| C      | 4  |
| B      | 3  |
| C      | 4  |
```

**Answer**

![](https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FTdI6rvMuXrb29zC97p1s%2Fimage6.png?alt=media\&token=3da180be-18b5-43d9-8b6c-5cfbc9ff5722)

You can check for duplicates using:

`= COUNTIF($B$2:$B$7)` Rows with value > 1 has duplicates

In order to create a table with Unique values there are 2 ways:

* Select the table and click on remove duplicates
* If you want to keep the source table and create the unique value table, elsewhere use:

`=UNIQUE(A2:B7)`

Conditional check can be done using IF clause, for example if you want to check duplicates only for ID > 3 you can use something like:

`=IF(B2>3,COUNTIF($B$2:$B$7,B4),0)`

</details>


# Time and Date

{% hint style="info" %}
Excel, as a product, always remains under active development from Microsoft. With new releases, new features are brought in whereas old ones are discarded. Due to this there might be changes to the solutions mentioned below depending on the version that you are using.
{% endhint %}

<details>

<summary>Time Diff</summary>

In many cases you will get business problems where you need to add or subtract month/year from a given date. Can you show how to do that in excel?

**Answer**

This can be done easily using the `DATE` function. The below formula will deduct a month, similarly one can change the year and date.

&#x20;`=DATE(YEAR(A3),MONTH(A3)-1,DAY(A3))`

</details>


# Python in Excel

Anaconda and Microsoft announced a groundbreaking innovation: Python in Excel. This marks a transformation in how Excel users and Python practitioners approach their work.

### `All You Need Is =PY()`&#x20;

Using Python in Excel is as simple as typing “=PY(” in your Excel cell, followed by your Python code. The results of your Python calculations or visualizations will then appear in your Excel worksheet.

For instance, you can use Python code to easily join two complex datasets, right within Excel.&#x20;

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FgLcknyhOFCbvNcOHfxzQ%2Fimage.png?alt=media&amp;token=8285f925-2c2b-46af-8a02-c4c7aecd9860" alt=""><figcaption></figcaption></figure>

Leverage robust Python visualization libraries, such as Matplotlib and Seaborn, right in your Excel workbook for comprehensive data representation.

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FkzhgVIylw3QBoZYGl7k7%2Fimage.png?alt=media&amp;token=771314f1-7f30-4e03-ab99-df9bfb2a6dfc" alt=""><figcaption></figcaption></figure>

Elevate your analysis using Python’s powerful libraries such as pandas and statsmodels. Accomplish comprehensive statistical tasks directly within your Excel cells. You don’t need to be a data science expert—Anaconda’s curated Python libraries embedded in Excel make advanced analytics accessible to everyone. 

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FoMItV7IdJStlWQaJrVfp%2Fimage.png?alt=media&amp;token=812ba268-8352-4db2-b452-704fd6aa8c19" alt=""><figcaption></figcaption></figure>

This feature doesn’t just bring Python into Excel; it brings the rich ecosystem of Python libraries as well. Libraries like pandas for data manipulation, statsmodels for advanced statistical modeling, and Matplotlib and Seaborn for data visualization are all available in Excel, unlocking a universe of new possibilities for your spreadsheets.

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FxJA00FsBWXrzXRLFgJJy%2Fimage.png?alt=media&amp;token=f5fa761b-aa94-403b-ba87-9dac0838e87b" alt=""><figcaption></figcaption></figure>


# PyCaret

[📚 Source](https://pycaret.gitbook.io/docs/)

PyCaret is an open-source, low-code machine learning library in Python that automates machine learning workflows. It is an end-to-end machine learning and model management tool that exponentially speeds up the experiment cycle and makes you more productive.

Compared with the other open-source machine learning libraries, PyCaret is an alternate low-code library that can be used to replace hundreds of lines of code with few lines only. This makes experiments exponentially fast and efficient. PyCaret is essentially a Python wrapper around several machine learning libraries and frameworks such as scikit-learn, XGBoost, LightGBM, CatBoost, spaCy, Optuna, Hyperopt, Ray, and a few more.

The link shared above has got detailed documentation and is an excellent resource to learn about PyCaret. Some common functions of PyCaret are discussed below.

## Setup

This function initializes the training environment and creates the transformation pipeline. Setup function must be called before executing any other function. It takes two mandatory parameters: data and target. All the other parameters are optional. Some of the things that can be done in setup are as follows:

* **Variables:** Define which variables to numerical, categorical or which are the ones to be ignored
* **Missing Value Imputation**
* **One Hot Encoding**
* **Target Imbalance**
* **Outlier Treatment**
* **Scaling and Transformation**
* **Feature Engineering and Selection**

## Train & Optimize

`compare_models` trains and evaluates the performance of all estimators available in the model library using cross-validation. Post model selection you can tune the model by optimizing prpbability threshold, stacking and ensembling models. PyCaret also provides an easy way to analyze your model by using the `evaluate_model` command. A single command which show out Hyperparameters, Plots, Feature importance and many things more.

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2FrvtEDUhfb8u6hhTnWBsx%2Fimage1.png?alt=media&amp;token=7469c2f1-2f56-4cfc-a4ad-55d44befca63" alt=""><figcaption></figcaption></figure>

The `interpret_model` command analyzes the predictions generated from a trained model. Most plots in this function are implemented based on the SHAP (Shapley Additive exPlanations).

## Predict & Deploy

`predict_model` function generates the output using a trained model. You can also get the raw probability scores for classification, set probability threshold and also monitor data drift. Post finalizing the model you can also easily deploy it in GCP, AWS and Azure.

**As of now PyCaret supports Regression, Classification, Clustering, Anamoly detection, NLP, Association Rule Mining and Time series.**


# Tensorflow

{% hint style="warning" %}
This page is a Work In Progress
{% endhint %}


# Business Scenarios

<details>

<summary>[SALESFORCE] User Base</summary>

Suppose as a Netflix product Analyst you are speaking with a Product Manager of Netflix. You get the question as to "I would want to understand user base of Netflix". What would be your approach?

**Answer**

This type of questions are basically a coversation and tries to judge your skills in breaking down the problem statement and providing an analytical solution.

* Ask follow up questions, these problem statements are kept vague to encourage you to ask questions
* Define KPIs
* Speak about your assumptions
* Break down the problem statement into 2-3 pieces and tell how would you solve them

</details>

<details>

<summary>[SALESFORCE] User Growth</summary>

Suppose your company has launched a Survey Monkey competitor and the PM wants to understand user growth. There are 2 flavors of the product:

* Unlimited survey's for 3 for a pro license
* One survey per license for a normal license

**Answer**

This type of questions are basically a coversation and tries to judge your skills in breaking down the problem statement and providing an analytical solution.

* Ask follow up questions, these problem statements are kept vague to encourage you to ask questions
* Define KPIs
* Speak about your assumptions
* Break down the problem statement into 2-3 pieces and tell how would you solve them

</details>

<details>

<summary>Valuation</summary>

Let's say that you're working at Netflix.

The company executives are working to renew a deal with another TV network that grants Netflix exclusive licensing to stream their hit TV series (think something like Friends or The Office). One of the executives wants to know how to approach this deal.

We know that the TV show has been on Netflix for a year already.

How would you approach valuing the benefit of keeping this show on Netflix?

**Answer**

Some of the factors which you can consider are as follows:

* Has the viewership stayed consistent or waned?
* Has there been any Social media backlash?
* Is there a new season or reunion or movie coming up?
* How many people who viewed this on Netflix renewed their subscriptions?
* What percent of viewers who saw this completed the entire series?
* How many of the people are repeat viewers of this seris?

</details>

<details>

<summary>[SQUARE] New Product</summary>

Let’s say that you’re in charge of Square’s small business division.

The CEO wants to hire a customer success manager for help managing a new software product. Another executive thinks that it’s worth just instituting a free trial instead.

What would be your recommendation on utilizing a customer success manager versus just a free trial to get new or existing customers to use the new product?

**ANSWER PENDING**

</details>

<details>

<summary>[FACEBOOK] Promote Product</summary>

[Source](https://www.interviewquery.com/questions/promoting-instagram)

Let's say you work on the growth team at Facebook and are tasked with promoting Instagram from within the Facebook app.

Where and how could you promote Instagram through Facebook?

**Answer**

Goal: Increase awareness of Instragram through Facebook Hypothesis: Showing Instragram ads to users in their News Feed will increase the likelihood that they will login to Instragram by $$X%$$.

**Run A/B test:**

* Control group: no changes
* Variant group: will be shown Instagram ads as the first ad they see when scrolling through their News Feed.
* Randomly assign users to each group, making sure they’re not biased and are representative of the population.
* Set a significance level like 95%
* Set experiment time, how long the long experiement run
* Set the power,usually 80%
* Esimate intended effect size - 20%

**Metrics:**

* Number of Instagram logins after being exposed to the Instagram ad, 24 hours
* Instagram logins / number of users - Percent logging into Instagram after using Facebook
* Stop-gap metric: Ad revenue, CTR, revenue per session. Since we’re taking up ad space, we want to see how much these ads cost us
* The other idea is: Notifying ppl on Facebook when their friends join Instagram. We can do a regression of number of friends on Instagram vs % of those users who use Instagram

Another thing to add to this would be to look at including an Instagram reference/clickable when a user uploads a new photo/video to Facebook. It is a natural fit there since Instagram is a photo/video sharing paltform and it is a natural integration to bring those together for those users who do not have Instagram.

</details>

<details>

<summary>[LINKEDIN] Email Campaign</summary>

[Source](https://www.linkedin.com/feed/update/urn:li:activity:6984161706524442624/)

LinkedIn’s marketing team tested two versions of email campaigns, A and B, in San Francisco(SF) and New York City (NYC). Of the 100,000 emails sent, 80% of the emails were version A while the rest were version B. Address the two questions below to guide the marketing team.

* Given that the click-through-rate (CTR) of email A was 15% while that of email B was 30%, which version is more effective?
* In SF, the CTR of email A was 15% while that of email B was 12.5%. In NYC, the CTR of email A was 15% while that of email B was 41.7%.

The summary table below shows the number of emails per variation per city. Explain your assessment

```
| Variations | NYC    | SF     |
|------------|--------|--------|
| A          | 20,000 | 60,000 |
| B          | 12,000 | 8,000  |
```

**Answer**

**\[Candidate]** Before I address your question, I will first organize CTR numbers you presented in a summary table:

```
| Variations | Size   | CTR | Count  |
|------------|--------|-----|--------|
| A          | 80,000 | 15% | 12,000 |
| B          | 20,000 | 30% | 6,000  |
```

As you mentioned, a total of 100,000 emails were dispatched, 80% being email A while the remaining being email B. Given the 15% CTR for email A and 30% CTR for email B, the counts of conversions were 12,000 and 6,000, A and B respectively.

**\[Interviewer]** Very well. What is your assessment?

**\[Candidate]** Just based on CTR’s, email is more effective in conversion. However, the result might be due to chance.

**\[Interviewer]** Can you elaborate?

**\[Candidate]** To assess the CTR difference between the two emails, I would suggest a statistical test.

**\[Interviewer]** Interesting. How would you use a statistical test?

**\[Candidate]** First, I would state the hypothesis:

$$H\_o$$: The CTR’s of emails A and B are the same. $$H\_a$$: The CTR’s of emails A and B are different.

Second, I would use either a chi-square contingency test or a T-test for two-sample proportions. If p-value of a test is less than the significance level at 0.05, reject the null hypothesis and conclude there is statistical significance in the CTR’s of the two emails. Finally, I can conclude that email B brings higher CTR than email A.

**\[Interviewer]** Okay. Do you believe the unequal sample sizes between A and B pose an issue?

**\[Candidate]** Not necessarily. I could foresee how the test's power may decrease as a consequence of the unequal sample size. However, one of the statistical tests I proposed assume that the sample size being equal. Therefore, I do not see an issue.

Coming to the second question:

**\[Candidate]** The change in click-through-rates within each city is a well-known statistical phenomenon called Simpson’s Paradox, which states that trend changes or reverses when observations are grouped. In the previous problem, you stated that the CTR of email A was 15% while that of email B was 30%, which suggests that email B produced higher CTR than email A. However, within SF, email A performed better than email B.

**\[Interviewer]** Your observation that the reversal in CTR’s is Simpson’s Paradox is correct. What else can you tell me about the problem?

**\[Candidate]** As I had done in the previous problem, let me organize the numbers in a summary table, rounded:

```
|            | Total        |              | NYC    |               | SF     |               |
|------------|--------------|--------------|--------|---------------|--------|---------------|
| Variations | Size         | Conversions  | Size   | Conversions   | Size   | Conversions   |
| A          | 80,000 (80%) | 12,000 (15%) | 20,000 | 3,000 (15%)   | 60,000 | 9000 (15%)    |
| B          | 20,000 (20%) | 6,000 (30%)  | 12,000 | 5,000 (41.7%) | 8,000  | 1,000 (12.5%) |
```

I have generic observations about the table. I see that the distribution of emails are higher for version B than A. I also see that San Francisco received higher allocation of email A than New York while New York received higher allocation of email B than email A.

**\[Interviewer]** Okay, do you see any issues with unequal email allocations between the two cities?

**\[Candidate]** Not at all as the primary metric is proportion, not count.

**\[Interviewer]** Do you have a hypothesis on why CTR’s vary?

**\[Candidate]** I presume that the version A email could be a generic email that does not improve CTR in one city over another. Version B might resonate with NYC market; hence, the CTR is at least 3X higher than that of version A. Perhaps, version B includes jobs in a finance industry that’s more prevalent in NYC than SF. I want to emphasize that the observed CTR’s occurred due to chance alone.

**\[Interviewer]** What statistical method would you use to validate whether cities affect click-through-rate of emails?

**\[Candidate]** If we disregard the possibility of interaction effect between email variations and cities, then we can use T-test for sample proportions to test whether cities affect email conversions.

First, I would establish the following hypothesis: $$H\_o: CTR\_{NYC} = CTR\_{SF}$$ $$H\_a: CTR\_{NYC} ≠ CTR\_{SF}$$ Next, I would create the following numerical summary pooled across variations by city:

```
| Cities | Calculations                        | CTR's |
|--------|-------------------------------------|-------|
| NYC    | (3,000 + 5,000) / (20,000 + 12,000) | 25%   |
| SF     | (9,000 + 1,000)/6,000 + 60,000)     | -15%  |
```

Next, I would calculate the sample standard error, and calculate the T-statistic: $$T − Statistic = \frac{(CTR\_{NYC} − CTR\_{SF})}{SE}$$ If the corresponding p-value is less than the usual significance level at 0.05, then reject the null hypothesis and conclude that there is statistical significance in the difference between the CTRs in the two cities at the significance level of 0.05.

**\[Interviewer]** Could you have used a one-sample T-test?

**\[Candidate]** Let me think. I believe it is possible to use a one-sample T-test. Instead of comparing the difference of CTR’s between NYC and SF, I can compare if the difference of CTR’s against 0 such that the hypothesis statements is now the following: $$H\_o: (CTR\_{NYC} − CTR\_{SF}) = 0$$ $$H\_a: (CTR\_{NYC} − CTR\_{SF}) ≠ 0$$

Based on this change, the T-statistic is calculated as follows: $$T − Statistic = \[(CTR\_{NYC} − CTR\_{SF}) − 0] / SE$$

**\[Interviewer]** Suppose that we do care about the interaction effect between city and variation. How would you evaluate it?

**\[Candidate]** As I think about this problem more, I realize that I could have applied ANOVA with two main effects and one-interaction effect. The model could assess three null hypothesis at the same time:

* *Hypothesis 1 - There is no difference in CTRs between NYC and SF.*
* *Hypothesis 2 - There is no difference in CTRs between variations A and B.*
* *Hypothesis 3 - There is an interaction effect on CTRs between cities and variations.*

The ANOVA test would provde statistical significance of each of the three terms measured. If there is an interaction effect between the city and email variation, then the p-value of the interaction term will be less than the significance level.

**\[Interviewer]** Okay, what is your final recommendation to the marketing team upon learning that there is an interaction effect?

**\[Candidate]** I would explain that the statistical test result suggests that there is an effect on CTR given the variation across cities and emails. I would then advise that the study is expanded on more markets to assess this business hypothesis further. In addition, I would suggest that more variations of the email are tested to assess the following assumption - emails tailored to a city’s demography performs better than a generic version across markets.

```
| Assessments             | Rating | Comments                                                                                                                                                                                                                                                                                                                                                                                                                                                                                              |
|-------------------------|--------|-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
| Statistical Methodology | 4      | In both problems, the candidate's statistical know-how was fairly strong.  In problem #1, she devised a statistical framework that addressed the interview question.  Her approach of utilizing T-test was appropriate.   In problem #2, most of her responses were solid except on a follow-up question involving one-sample T-test.  Given that sample-size difference between the two groups, mathematically, the statistical test is not possible.                                                |
| Product Sense           | 5      | In both questions, she ensured that her responses are grounded in the marketing problem.  She devised a numerical summary, making sense of the CTR metrics.  This allowed her to suggests approaches that aligned with the problem.  Lastly, in her last problem, when asked about her recommendation, she offered sound suggestions.  The idea of testing email variations on more than two emails makes sense to assess whether emails  focused on target market perform better than a general one. |
| Communication           | 5      | The candidate ensured that she understood the problems, illustrating a numerical summary and explaining her analysis clearly.  Her explanation of statistical were easy to follow and comprehensive. Lastly, she clearly explained recommendations to a  marketing team, suggesting that she possesses fluidity in stakeholder engagement.                                                                                                                                                            |
```

</details>


# Industry Application

<details>

<summary>Zillow</summary>

Zillow's rise and fall is a classic case study on both what to do and what not to do with data science and machine learning.

* Zillow was the brainchild of Richard Barton, the 54-year-old American internet entrepreneur who also founded Expedia and Glassdoor
* Zestimate, which debuted in 2006, it was a proprietary algorithm based on a neural network model which used house facts, location, housing market trends, property values, data from county, tax assessor records, as well as direct feeds from hundreds of multiple listing services and brokerages to determine the price of a house
* It was very successful in driving conversation around property, so much so that browsing property valuations became a hobby for many. With shows like Saturday Night Live encouraging the habit, it became trendy to find the value of your neighboring houses to identify how posh of an area you live in.
* Zestimate has a median error rate of $1.9%$ for homes that are on-market and $6.9%$ for homes that are off-market. But this accuracy varies widely across cities and their corresponding housing markets. In Cincinnati, for instance, approximately $35%$ of Zestimates for off-market homes were within $5%$ of the eventual sales prices, and $82%$ were within $20%$ of the price. Comparatively, in Denver, $51%$ of Zestimates were within $5%$ of the sales price, and $94%$ were within $20%$ of the sales figure
* Zillow had recognized the need to improve the accuracy of its valuation model. In 2017, the company launched a contest with a $1M prize money, Zestimate was on average around $10K off of the actual sales price of a median-priced home, and that the information provided by the winning team would reduce that margin by around $1.3K, the prize went to a group of data scientists and engineers from three different countries: Chahhou Mohamed of Morocco, Jordan Meyer of the United States and Nima Shahbazi of Canada
* In February Zillow announced that the Zestimate would represent "an initial cash offer" for eligible homes spread across 20 cities, from Nashville, Houston, Phoenix, Miami to Denver and Los Angeles
* The problem started due to the pandemic, since summer of 2020 the housing market in USA was hot and quite some unforeseen trends had cropped up. For example, post pandemic there was a sudden increase in demand for bigger houses in the suburbs, which based on historical data was difficult to generate an estimate of. There were data issues too, Zillow missed out on having real time brokering data in new markets, post sales data also takes time to be updated hence there was always a lag as to what prices the recent houses have sold at. Other issues have also been pointed at, Zillow said it did not find enough workers to revamp and repair the houses it had bought in order to put them in the market. Zillow had tweaked its algorithms to price houses aggressively and at the end it had a purchased houses at a higher price that it could sell in a market that was cooling down
* By October 17th, there were signs of trouble: Zillow had paused new house offers for the year as it sorted through a backlog of properties under contract. On November 1st, Zillow marketed nearly 7,000 houses for a total of $2.8 Billion. The operation was shut down a day later by its board of directors

So, what went wrong, nothing that we already don't know of, extrapolating existing data in new pandemic era market, unforeseen market dynamics, over reliance on algorithmic outputs while not considering ground realities.

</details>


# Behavioral/Management

<details>

<summary>Disagreement with colleagues</summary>

Tell me a time when your colleagues did not agree with your approach. What did you do to bring them into the conversation and address their concerns?

**Answer**

The answer to this will obviously differ from person to person. Try to think of an example related to the projects that you have already discussed or showcased in your CV or in the interview. Focus on the following things:

* What was the conflict about?
* Show that you thought through both the viewpoints and not bulldozed your ideas
* Explain which idea was chosen finally and why did you choose it. Try to showcase that chosing the best method of the project was important rather than individual ideas

</details>

<details>

<summary>Exceed Expectations</summary>

Tell me about a time when you exceeded expectations during a project. What did you do, and how did you accomplish it?

**Answer**

Here showcase a project where you went the extra mile. To be honest not a very difficult example to come up with. Did business ask for a key driver analysis for customer churn and you not only did that but also created a easy to use Power BI dashboard where they can get a list of customers too with the highest chance of churning. Something along these lines.

</details>

<details>

<summary>Strength vs Weakness</summary>

When an interviewer asks a question along the lines of:

* What would your current manager say about you? What constructive criticisms might he give?
* What are your three biggest strengths and weaknesses you have identified in yourself?
* How would you respond?

**Answer**

Gone are those days where you could sell a weakness as a strength. Very few people gets sold on things like I am a workoholic where you are indirectly trying to sell your commitment to work at the cost of work life balance. Actually these could backfire beacuse a lot of companies do a culture fit of employees. Instead be honest, be humble, introspect find something which makes you human but also something which is a not huge handicap for the role you are applying for. You can tell things like need to work on story telling, though I have worked on it a lot etc etc.

There are no one size fits all answer to this question, it depends on which company, seniority of the role etc.

</details>


# Vector Database

With the rise of Foundational Models, Vector Databases skyrocketed in popularity. Vector Database is also useful outside of a Large Language Model context.

([Source](https://towardsdatascience.com/explaining-vector-databases-in-3-levels-of-difficulty-fc392e48ab78)) ([Source](https://www.linkedin.com/feed/update/urn:li:activity:7092611326891470849/))

The vector embedding takes a word or object and converts it into a bunch of numbers in a special way. This conversion is done in such a manner that similar words or objects end up having similar numbers in their embeddings. For instance, the embeddings of "cat" and "dog" would be close to each other because they share some similarities, both being animals. On the other hand, the embeddings of "cat" and "apple" would be far apart because they are very different things.

Vector embeddings help us understand relationships and similarities between different words or objects in a way that computers can easily handle and analyze. It's like giving words or objects a unique digital fingerprint, making them easier for computers to work with and compare. This concept is commonly used in natural language processing and various machine learning tasks.

The numerical representations enable us to apply mathematical calculations to objects, such as words, which are usually not suited for calculations. For example, the following calculation will not work unless you replace the words with their embeddings:

```
drink - food + hungry = thirsty
```

And because we are able to use the embeddings for calculations, we can also calculate the distances between a pair of embedded objects. The closer two embedded objects are to one another, the more similar they are.

### How does Vector Database work?

Vector databases are able to retrieve similar objects of a query quickly because they have already pre-calculated them. The underlying concept is called Approximate Nearest Neighbor (ANN) search, which uses different algorithms for indexing and calculating similarities.

As you can imagine, calculating the similarities between a query and every embedded object you have with a simple k-nearest neighbors (kNN) algorithm can become time-consuming when you have millions of embeddings. With ANN, you can trade in some accuracy in exchange for speed and retrieve the approximately most similar objects to a query.

**Indexing** — For this, a vector database **indexes** the vector embeddings. This step maps the vectors to a data structure that will enable faster searching. We will not go into the technical details of indexing algorithms, but if you are interested in further reading, you might want to start by looking up Hierarchical Navigable Small World (HNSW).

**Similarity Measures —** To find the nearest neighbors to the query from the indexed vectors, a vector database applies a similarity measure. Common similarity measures include cosine similarity, dot product, Euclidean distance, Manhattan distance, and Hamming distance.

### Database Operations

Let’s look into how one would interact with a Vector Database:

#### Writing/Updating Data

1. Choose a ML model to be used to generate Vector Embeddings.
2. Embed any type of information: text, images, audio, tabular. Choice of ML model used for embedding will depend on the type of data.
3. Get a Vector representation of your data by running it through the Embedding Model.
4. Store additional metadata together with the Vector Embedding. This data would later be used to pre-filter or post-filter ANN search results.
5. Vector DB indexes Vector Embedding and metadata separately. There are multiple methods that can be used for creating vector indexes, some of them: Random Projection, Product Quantization, Locality-sensitive Hashing.
6. Vector data is stored together with indexes for Vector Embeddings and metadata connected to the Embedded objects.

#### Reading Data

7. A query to be executed against a Vector Database will usually consist of two parts:

➡️ Data that will be used for ANN search. e.g. an image for which you want to find similar ones. ➡️ Metadata query to exclude Vectors that hold specific qualities known beforehand. E.g. given that you are looking for similar images of apartments - exclude apartments in a specific location.

8. You execute Metadata Query against the metadata index. It could be done before or after the ANN search procedure.
9. You embed the data into the Latent space with the same model that was used for writing the data to the Vector DB.
10. ANN search procedure is applied and a set of Vector embeddings are retrieved. Popular similarity measures for ANN search include: Cosine Similarity, Euclidean Distance, Dot Product.

Some popular Vector Databases: Qdrant, Pinecone, Weviate, Milvus, Faiss, Vespa.

<figure><img src="https://3998717274-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F2BHF6WoWT3fDmZodGjU2%2Fuploads%2F0kQ9066GIv2Vwl9FqJ8E%2Fimage.png?alt=media&amp;token=24ce0c54-cd11-488b-a558-95946f5702d3" alt=""><figcaption><p>(<a href="https://www.linkedin.com/feed/update/urn:li:activity:7092611326891470849/">Source</a>)</p></figcaption></figure>


# LLMs

An overview of large language models (LLMs), covering LLM prompting, LLM fine-tuning, and LLM application development.

{% hint style="info" %}
Check out these 2 brilliant resources [Source](https://www.kaggle.com/code/abireltaief/contemporary-large-language-models-llms) (more accessible) [Source](https://eugeneyan.com/writing/llm-patterns/)(detailed)
{% endhint %}

Large language models (LLMs) are AI models that are intended to comprehend and generate human language, code, and much more. They are typically derived from the Transformer architecture. This architecture was introduced by a team at Google Brain in 2017. Attention, transfer learning, and scaling up neural networks are the main principles that lay the groundwork for Transformers to flourish.

### Pretrained LLMs: Efficiency  <a href="#id-2.-pretrained-llms-efficiency" id="id-2.-pretrained-llms-efficiency"></a>

Pretrained LLMs, such as OpenAI's GPT-4 and GPT-3, Meta's LLaMA, Google's PaLM, Flacon, etc., are models with sizes of billions of parameters, and for some, the size exceeds 50 billion or even 100 billion parameters. This means we need a large memory capacity to train such models : tens of thousands of gigabytes. This is quite expensive and requires access to hundreds of GPUs (+ high carbon footprint).

Recent research studies have looked at the relationship between **model size** (number of parameters), **the dataset size** used to train the LLM model, and its **performance**. Is it true that the performance will be enhanced by increasing model parameters? How big is the dataset, exactly? What happens if the dataset size used to train the LLM is increased? What is the ideal ratio between these two key elements?

A team of academics from DeepMind conducted in-depth research on the performance of large language models with various model sizes (number of parameters) and training dataset sizes. The results were published in the paper "Training Compute-Optimal Large Language Models" in 2022 . The ideal model produced by the author's work is called "Chinchilla".

**The most important key lesson** from the Chinchilla research is that the best size of the training dataset for a given LLM model is about 20 times larger than the model's parameters number. The ideal training dataset (for the best compute optimal model Chinchilla) has 1.4 trillion tokens for 70 billion parameters, as depicted in Table-2, above.

**Important finding:**\
Just one year after the publication of the Chinchilla Paper, Meta AI makes available its state-of-the-art large language model, LLaMA-65B. This model was trained using a dataset that was roughly 1.4 trillion tokens in size, which is similar to Chinchilla's suggested number.

### Prompt Engineering

Sometimes using these state-of-the-art pretrained LLMs for inference leads to undesirable results right away. To get the model to behave in the desired manner, we might need to make multiple revisions to the language or format of our prompt. The process of creating and enhancing the prompt is referred to as prompt engineering.

#### Basic Methods: Zero-shot, one-shot & few-shot prompting

* **Zero-shot prompting:**

Zero-shot prompting refers to giving the model a prompt (task) that isn't part of the training data but can nonetheless lead to the desired output. For example, in the case of sentiment analysis, we are aware that the ideal method for classifying reviews is to train a machine learning model on labeled data in order to obtain the appropriate result for an as-yet-unseen review in inference. However, LLMs are not trained in this way or to perform this task, but if we ask them to perform classification tasks, they can still deliver the desired results. Here's an example of Zero-shot prompting:

> Classify this review:\
> I love red apples\
> Sentiment:

* **One-shot prompting & few-shot prompting**

Performance can be enhanced by one-shot prompting, where we can guide the LLM with an example of the desired behavior within the prompt.

> Classify this review:\
> I love red apples\
> Sentiment: positive
>
> Classify this review\
> I adore vegetables\
> Sentiment :

Sometimes providing one example is not enough. For the example of sentiment analysis, we can provide a prompt with two examples (or more): one positive and the other negative. This is few-shot prompting.

**Important finding**\
According to a study that was released in 2020, "Language Models are Few-Shot Learners", most large language models are quite good at inference with few-shot prompting. This paper examines the potential of few-shot prompting in LLMs.

#### Chain of Thought (CoT) <a href="#id-3.2-chain-of-thought-cot" id="id-3.2-chain-of-thought-cot"></a>

A prompting method used to assist LLMs in reasoning is Chain of Thought. Making the model think more like a human by breaking the task down into steps is one tactic that has shown some promise. When you use examples for one-shot or few-shot inference, it works by incorporating a number of intermediate reasoning steps.

The example below, which is derived from the paper "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models,2023", published by researchers at Google, demonstrates how the LLM's behavior was completely changed when a standard one-shot prompt (simply giving the answer) was replaced with a chain of thought prompt by including the reasoning steps that solve the problem. This helped the LLM succeed by implicitly asking him to mimic human behavior.

![Standard one-shot prompting vs Chain of Thought .Source: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models,2023](https://imgur.com/ZF7jBfK.png)

#### Reasoning & Acting (ReAct)

ReAct was created by researchers from Google and Princeton, in 2022. It was proposed in the paper "ReAct: Synergizing Reasoning and Acting in Language Models, 2022". It is a cutting-edge prompting technique that is being used by frameworks that communicate with LLMs, such as LangChain. It is an iterative process that requires calling the LLM to get a **"thought"** and an **"action"** (trying to solve the task asked by the user). This action would be executed by accessing external sources via an implemented tool that communicates between the LLM, the user App (where you write the prompt), and the real world (external data from the web, for example, or external APIs). It's usual to refer to this intermediary tool as a technology implemented in the **orchestration library**. The outcome of this first action is called **"observation"**. The ensemble of the first thought, the first action, and the first observation will construct the context for the second LLM's iteration (call). As a matter of fact, depending on this context (Thought1 + Action1 + Observation1), the LLM will generate a second thought and a second action trying to solve the task (asked by the user), then it executes this newly generated second action using the orchestration tool to collect a second observation. If the problem (task) is not resolved by this second observation, the process of coming up with thoughts and actions (and carrying them out) will go on until the problem (the task) is resolved.

As you have certainly noticed, ReAct is a technique based on the capacity of human intelligence to integrate verbal reasoning with task-oriented actions (exactly mimicking human behavior when solving a problem).

Here is an example that was taken from the original paper that introduced the ReAct paradigm, published in 2022 (mentioned above).

![An example of how ReAct prompts the LLM through thoughts and actions and how it makes decisions about the manner of interacting with external data to collect observations](https://imgur.com/UQHxeic.png)

### Fine-tuning LLMs: Low-rank Adaptation LoRA <a href="#id-4.-fine-tuning-llms-low-rank-adaptation-lora" id="id-4.-fine-tuning-llms-low-rank-adaptation-lora"></a>

When creating an LLM-powered app, working with pre-trained LLMs can help us save a lot of time. However, if our target area employs uncommon terminology, we might find it necessary to fine-tune LLMs based on specialized data. For instance, if we have to create a medical app powered by LLMs, typically, medical terminology uses many rare phrases to describe medical diseases and procedures that may not be regularly found in the training datasets (web scrapes and book texts) used to train LLMs. Thus, fine-tuning LLMs with specialized data will result in better models for our LLM-powered apps.

A group of methods known as **parameter-efficient fine tuning** trains only a limited subset of task-specific layers and parameters while maintaining most of the weights of the original LLM (or even all the original weights with LoRa). Since all or the majority of the pre-trained weights remain constant, these methods exhibit greater resistance against the well-known catastrophic forgetting phenomenon, which happens when a fully finetuned LLM loses its primary capabilities after updating its weights. There are several parameter-efficient fine-tuning methods.

**How does Low-rank Adaptation (LoRA) work?**

LoRa simply adds a minimal number of new parameters and fine-tunes them, leaving the existing original model weights untouched (frozen). LoRa injects two low rank decomposition matrices. These two new small matrices should be designed with dimensions such that the resulting matrix (their multiplication) has the same dimensions as the original weights matrix. Then, as previously stated, we just train these two low-rank matrices on the new data (the new task) while maintaining the original weights of the LLM frozen.

For inference, these two trained low-rank matrices are multiplied in order to produce a matrix with the same dimensions as the frozen weights. After that, we add this to the initial weights and update the model to reflect these new values (the sum of the original weights and the new trained weights). This is the LoRA-fine-tuned model that can complete our particular task. Depending on our new task and data, we can train as many LoRA matrices as we need: all we have to do is add these updated weights to the initial frozen weights each time and run the inference with the new model.

The new LoRa variant **QLoRA**, which is LoRa fine-tuning method but based on quantization, is introduced in the paper "QLoRA: Efficient Finetuning of Quantized LLMs", which was just released (May 2023). Quantization transforms the model's weights to a lower precision representation.

The authors named their best QLoRa fine-tuned model **Guanaco**. On the Vicuna benchmark, "it exceeds all prior available models, achieving 99.3% of ChatGPT's performance level with only 24 hours of fine-tuning on a single GPU." (this sentence was taken from the paper). Impressive findings.\
QLoRA is definitely the most effective way to fine-tune large language models on a single GPU.

### Fine-tuning LLMs with Human Feedback (RLHF)  <a href="#id-5.-fine-tuning-llms-with-human-feedback-rlhf" id="id-5.-fine-tuning-llms-with-human-feedback-rlhf"></a>

Making LLMs function in accordance with human preferences is the goal of the process known as "fine-tuning LLMs with Human Feedback". For example, we want LLMs to behave as helpful, harmless, and honest assistants. These three values harmlessness-helpfulness-honesty, are the key ones cited in most research articles when discussing fine-tuning with human feedback. As an illustration, the authors of the paper "Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback, 2022" \[10], discuss these 3 criteria and refer to them by the acronym "HHH".

Naturally, we may specify additional values that we want the LLMs to adhere to, in order to conform to human preferences (ensuring that the LLM generates completions that maximize or minimize any criteria set out by us). As an illustration, I would like the LLM to be particularly tolerant (not prejudiced or racist).

Reinforcement learning from human feedback (RLHF) is the most popular method for fine-tuning LLMs with human feedback.

**RLHF** is a process for fine-tuning LLMs to match human preferences. It is an iterative process (like other tuning processes). There are two central components in this process: the **reward model** and the **reinforcement learning algorithm**. The reward model outputs a reward value to the completions generated by the LLM and passes this value to the RL algorithm, which will iteratively adjust the weights of the LLM to maximize the reward obtained from human feedback, pushing the LLM to produce texts that are more aligned with the criteria we defined (for example, tolerance).

**How does RLH really work?**

* 1st step: Construct a dataset of many prompts and completions (generated by the LLM we want to fine-tune with RLHF).
* 2nd step: We collect feedback from humans labers that will score these completions according to the criterion we defined, let's say, tolerance, by classifying them as tolerant or non-tolerant (racist) generated texts. In fact, this step is more complicated (not straightforward, assigning many labers for each completion, etc.), and it is actually time-consuming and very expensive. At the end of this step, we get the dataset that will be used to build the reward model (one of the central components of the RLHF process), which will replace the human labers in the fine-tuning stage.
* 3rd step: Building the reward model (using a supervised learning model: a classifier).
* 4th step: The reinforcement learning finetuning process begins by passing prompts to the LLM, which generates completions. The reward model assigns a reward value to the completions (high for more tolerant text or low for less tolerant text). Then, it passes this value to the RL algorithm, which will update the weights of the LLM, pushing it to generate more aligned text with the criterion defined (tolerance in this case) to maximize the reward obtained from the reward model. This is a single iteration.
* 5th step : Step 4 will be repeated in an iterative manner, updating the weights of the LLM model at each iteration, until reaching a threshold already defined (for the criterion or number of iterations).

The final version of the fine-tuned LLM with Human feedback (using RL) is called **human-aligned LLM**: It is our ultimate goal to ensure that our LLM behaves in an aligned manner in deployment!

There are several different RL algorithms that can be used in the RLHF process. The most common one is proximal policy optimization (PPO).

Here is an illustration of the main steps of fine-tuning an LLM with human feedback using RL, derived from the paper "Learning to summarize from human feedback, 2022" \[12]

![Steps of fine-tuning an LLM with human feedback using RL (RLHF)Source: Learning to summarize from human feedback, 2022](https://imgur.com/OYGkdpn.png)

### Augmenting LLMs to create LLM-powered Apps : RAG <a href="#id-6.-augmenting-llms-to-create-llm-powered-apps-rag" id="id-6.-augmenting-llms-to-create-llm-powered-apps-rag"></a>

Regularly retraining LLMs with new data is ineffective as a solution. The significant carbon footprint caused by the enormous number of GPUs running continuously will make it not only incredibly expensive but also a burden for the environment and the planet.\
Implementing technologies that give LLMs access to external data during inference time is a more flexible and affordable technique to get around knowledge cutoffs. The best example of this, is **RAG** framework: **Retrieval Augmented Generation.**

Actually, RAG is a framework for **creating LLM-powered apps** that use outside data sources and get beyond the LLM's drawbacks, like information that is missing from the pre-training phase, known in some publications as the knowledge cutoff phenomenon. But not only that, RAG also gives access to external APIs and even confidential information kept in the private databases of companies and organizations (to restrict the access just to them). This is made possible through an orchestration library. This layer has the potential to enable certain potent technologies that will enhance LLM's performance.\
Simply said, RAG is an excellent approach to assist the LLM model in updating and broadening its view of the world !

RAG has many implementations, varying in complexity from simple to complicated. LangChain is one popular RAG implementation. The advanced prompting strategies, like ReAct, that I mentioned in paragraph-3, are effectively implemented in LangChain.

I'll be talking about the straightforward general architecture of RAG framework, that was suggested in the paper "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, 2021" \[13], One of the earliest papers published about RAG.

The **retriever** is the main component of this architecture. It is a combination of a query encoder and an external data source.

**How does it work ?**

* 1st step: The encoder takes the user's prompt and converts it into a format that the external data source (which could be a database, for example) can use. In the paper, it is a vector store (vector stores allow for a quick and effective similarity-based search and are also a good data format for the LLMs because they use the same format to generate text).
* 2nd step: Then, it searches for relevant information, in the external sources, that matches the query (the encoded user's prompt).
* 3rd step: When appropriate information is discovered in the external sources, the retriever adds the new data to the initial prompt.
* 4th: The new enhanced prompt (the original prompt + new extracted information) is passed to the LLM, which generates the final output (completion) for the user.

This general diagram (Figure-4, below) is an illustration of the straightforward architecture of the Retrieval Augmented Generation framework (RAG).

![Retrieval Augmented Generation framework (RAG)Source](https://imgur.com/yg8qK0p.jpg)

### Caching: To reduce latency and cost

Caching is a technique to store data that has been previously retrieved or computed. This way, future requests for the same data can be served faster. In the space of serving LLM generations, the popularized approach is to cache the LLM response keyed on the embedding of the input request. Then, for each new request, if a semantically similar request is received, we can serve the cached response.

For some practitioners, this sounds like “[a disaster waiting to happen.](https://twitter.com/HanchungLee/status/1681146845186363392)”

Caching can significantly reduce latency for responses that have been served before. In addition, by eliminating the need to compute a response for the same input again and again, we can reduce the number of LLM requests and thus save cost. Also, there are certain use cases that do not support latency on the order of seconds. Thus, pre-computing and caching may be the only way to serve those use cases.

An example of caching for LLMs is [GPTCache](https://github.com/zilliztech/GPTCache).

![Overview of GPTCache](https://eugeneyan.com/assets/gptcache.jpg)

When a new request is received:

* Embedding generator: This embeds the request via various models such as OpenAI’s `text-embedding-ada-002`, FastText, Sentence Transformers, and more.
* Similarity evaluator: This computes the similarity of the request via the vector store and then provides a distance metric. The vector store can either be local (FAISS, Hnswlib) or cloud-based. It can also compute similarity via a model.
* Cache storage: If the request is similar, the cached response is fetched and served.
* LLM: If the request isn’t similar enough, it gets passed to the LLM which then generates the result. Finally, the response is served and cached for future use.

Redis also shared a [similar example](https://www.youtube.com/live/9VgpXcfJYvw?feature=share\&t=1517), mentioning that some teams go as far as precomputing all the queries they anticipate receiving. Then, they set a similarity threshold on which queries are similar enough to warrant a cached response.

### Guardrails: To ensure output quality

In the context of LLMs, guardrails validate the output of LLMs, ensuring that the output doesn’t just sound good but is also syntactically correct, factual, and free from harmful content. It also includes guarding against adversarial input.

#### Why guardrails?

First, they help ensure that model outputs are reliable and consistent enough to use in production. For example, we may require output to be in a specific JSON schema so that it’s machine-readable, or we need code generated to be executable. Guardrails can help with such syntactic validation.

Second, they provide an additional layer of safety and maintain quality control over an LLM’s output. For example, to verify if the content generated is appropriate for serving, we may want to check that the output isn’t harmful, verify it for factual accuracy, or ensure coherence with the context provided.

There are many approaches to setting up guardrails, check the second source mentioned in the beginning of this page incase you are interested in learning more.


# NumPy

{% embed url="<https://assets.datacamp.com/blog_assets/Numpy_Python_Cheat_Sheet.pdf>" %}
([Source](https://assets.datacamp.com/blog_assets/Numpy_Python_Cheat_Sheet.pdf))
{% endembed %}


# Pandas

{% embed url="<https://pandas.pydata.org/Pandas_Cheat_Sheet.pdf>" %}
([Source](https://pandas.pydata.org/Pandas_Cheat_Sheet.pdf))
{% endembed %}


# Pyspark

{% embed url="<https://images.datacamp.com/image/upload/v1676303379/Marketing/Blog/PySpark_RDD_Cheat_Sheet.pdf>" %}
([Source](https://images.datacamp.com/image/upload/v1676303379/Marketing/Blog/PySpark_RDD_Cheat_Sheet.pdf))
{% endembed %}


# SQL

{% embed url="<https://learnsql.com/blog/sql-basics-cheat-sheet/sql-basics-cheat-sheet-letter.pdf>" %}


# Statistics

{% embed url="<https://static1.squarespace.com/static/54bf3241e4b0f0d81bf7ff36/t/55e9494fe4b011aed10e48e5/1441352015658/probability_cheatsheet.pdf>" %}
([Source](https://static1.squarespace.com/static/54bf3241e4b0f0d81bf7ff36/t/55e9494fe4b011aed10e48e5/1441352015658/probability_cheatsheet.pdf))
{% endembed %}


# RegEx

{% embed url="<https://res.cloudinary.com/dyd911kmh/image/upload/v1665049611/Marketing/Blog/Regular_Expressions_Cheat_Sheet.pdf>" %}
([Source](https://res.cloudinary.com/dyd911kmh/image/upload/v1665049611/Marketing/Blog/Regular_Expressions_Cheat_Sheet.pdf))
{% endembed %}


# Git

{% embed url="<https://education.github.com/git-cheat-sheet-education.pdf>" %}
([Source](https://education.github.com/git-cheat-sheet-education.pdf))
{% endembed %}


# Power BI

{% embed url="<https://s3.amazonaws.com/assets.datacamp.com/email/other/Power+BI_Cheat+Sheet.pdf>" %}
([Source](https://s3.amazonaws.com/assets.datacamp.com/email/other/Power+BI_Cheat+Sheet.pdf))
{% endembed %}


# Python Basics

{% embed url="<https://perso.limsi.fr/pointal/_media/python:cours:mementopython3-english.pdf>" %}
([Source](https://perso.limsi.fr/pointal/_media/python:cours:mementopython3-english.pdf))
{% endembed %}


# Keras

{% embed url="<https://s3.amazonaws.com/assets.datacamp.com/blog_assets/Keras_Cheat_Sheet_Python.pdf>" %}


# R Basics

{% embed url="<https://iqss.github.io/dss-workshops/R/Rintro/base-r-cheat-sheet.pdf>" %}
([Source](https://iqss.github.io/dss-workshops/R/Rintro/base-r-cheat-sheet.pdf))
{% endembed %}


# PRIVACY NOTICE

The short version is that we do not collect personal information. We use Google analytics for things like country or device our users are accessing from & which pages they are visiting. That's all.

This privacy notice for The Data Science Interview Project (doing business as The Data Science Interview Book) ("**The Data Science Interview Book**," "**we**," "**us**," or "**our**"), describes how and why we might collect, store, use, and/or share ("**process**") your information when you use our services ("**Services**"), such as when you:

* Visit our website at <https://book.thedatascienceinterviewproject.com>, or any website of ours that links to this privacy notice
* Engage with us in other related ways, including any sales, marketing, or events

**Questions or concerns?** Reading this privacy notice will help you understand your privacy rights and choices. If you do not agree with our policies and practices, please do not use our Services. If you still have any questions or concerns, please contact us at <thedatascienceinterviewbook@gmail.com>.\
\
**SUMMARY OF KEY POINTS**\
***This summary provides key points from our privacy notice, but you can find out more details about any of these topics by clicking the link following each key point or by using our table of contents below to find the section you are looking for. You can also click here to go directly to our table of contents.***\
**What personal information do we process?** When you visit, use, or navigate our Services, we may process personal information depending on how you interact with The Data Science Interview Book and the Services, the choices you make, and the products and features you use. Click here to learn more.\
**Do we process any sensitive personal information?** We do not process sensitive personal information.\
**Do we receive any information from third parties?** We do not receive any information from third parties.\
**How do we process your information?** We process your information to provide, improve, and administer our Services, communicate with you, for security and fraud prevention, and to comply with law. We may also process your information for other purposes with your consent. We process your information only when we have a valid legal reason to do so. Click here to learn more.\
**In what situations and with which parties do we share personal information?** We may share information in specific situations and with specific third parties. Click here to learn more.\
**How do we keep your information safe?** We have organizational and technical processes and procedures in place to protect your personal information. However, no electronic transmission over the internet or information storage technology can be guaranteed to be 100% secure, so we cannot promise or guarantee that hackers, cybercriminals, or other unauthorized third parties will not be able to defeat our security and improperly collect, access, steal, or modify your information. Click here to learn more.\
**What are your rights?** Depending on where you are located geographically, the applicable privacy law may mean you have certain rights regarding your personal information. Click here to learn more.\
**How do you exercise your rights?** The easiest way to exercise your rights is by filling out our data subject request form available [here](https://app.termly.io/notify/6c6ffbe2-67bd-4961-98f4-ff9a9600d8d5), or by contacting us. We will consider and act upon any request in accordance with applicable data protection laws.\
Want to learn more about what The Data Science Interview Book does with any information we collect? Click here to review the notice in full.\
\
**1. WHAT INFORMATION DO WE COLLECT?**\
**Personal information you disclose to us**\
***In Short:** We collect personal information that you provide to us.*\
We collect personal information that you voluntarily provide to us when you express an interest in obtaining information about us or our products and Services, when you participate in activities on the Services, or otherwise when you contact us.\
**Sensitive Information.** **We do not process sensitive information.**\
All personal information that you provide to us must be true, complete, and accurate, and you must notify us of any changes to such personal information.\
**Information automatically collected**\
***In Short:** Some information — such as your Internet Protocol (IP) address and/or browser and device characteristics — is collected automatically when you visit our Services.*\
We automatically collect certain information when you visit, use, or navigate the Services. This information does not reveal your specific identity (like your name or contact information) but may include device and usage information, such as your IP address, browser and device characteristics, operating system, language preferences, referring URLs, device name, country, location, information about how and when you use our Services, and other technical information. This information is primarily needed to maintain the security and operation of our Services, and for our internal analytics and reporting purposes.\
Like many businesses, we also collect information through cookies and similar technologies.\
The information we collect includes:

* *Device Data.* We collect device data such as information about your computer, phone, tablet, or other device you use to access the Services. Depending on the device used, this device data may include information such as your IP address (or proxy server), device and application identification numbers, location, browser type, hardware model, Internet service provider and/or mobile carrier, operating system, and system configuration information.
* *Location Data.* We collect location data such as information about your device's location, which can be either precise or imprecise. How much information we collect depends on the type and settings of the device you use to access the Services. For example, we may use GPS and other technologies to collect geolocation data that tells us your current location (based on your IP address). You can opt out of allowing us to collect this information either by refusing access to the information or by disabling your Location setting on your device. However, if you choose to opt out, you may not be able to use certain aspects of the Services.

**2. HOW DO WE PROCESS YOUR INFORMATION?**\
***In Short:** We process your information to provide, improve, and administer our Services, communicate with you, for security and fraud prevention, and to comply with law. We may also process your information for other purposes with your consent.*\
**We process your personal information for a variety of reasons, depending on how you interact with our Services, including:**

* **To save or protect an individual's vital interest.** We may process your information when necessary to save or protect an individual’s vital interest, such as to prevent harm.

\
**3. WHAT LEGAL BASES DO WE RELY ON TO PROCESS YOUR INFORMATION?**\
***In Short:** We only process your personal information when we believe it is necessary and we have a valid legal reason (i.e., legal basis) to do so under applicable law, like with your consent, to comply with laws, to provide you with services to enter into or fulfill our contractual obligations, to protect your rights, or to fulfill our legitimate business interests.*\
***If you are located in the EU or UK, this section applies to you.***\
The General Data Protection Regulation (GDPR) and UK GDPR require us to explain the valid legal bases we rely on in order to process your personal information. As such, we may rely on the following legal bases to process your personal information:

* **Consent.** We may process your information if you have given us permission (i.e., consent) to use your personal information for a specific purpose. You can withdraw your consent at any time. Click here to learn more.
* **Legal Obligations.** We may process your information where we believe it is necessary for compliance with our legal obligations, such as to cooperate with a law enforcement body or regulatory agency, exercise or defend our legal rights, or disclose your information as evidence in litigation in which we are involved.<br>
* **Vital Interests.** We may process your information where we believe it is necessary to protect your vital interests or the vital interests of a third party, such as situations involving potential threats to the safety of any person.

\
***If you are located in Canada, this section applies to you.***\
We may process your information if you have given us specific permission (i.e., express consent) to use your personal information for a specific purpose, or in situations where your permission can be inferred (i.e., implied consent). You can withdraw your consent at any time. Click here to learn more.\
In some exceptional cases, we may be legally permitted under applicable law to process your information without your consent, including, for example:

* If collection is clearly in the interests of an individual and consent cannot be obtained in a timely way
* For investigations and fraud detection and prevention
* For business transactions provided certain conditions are met
* If it is contained in a witness statement and the collection is necessary to assess, process, or settle an insurance claim
* For identifying injured, ill, or deceased persons and communicating with next of kin
* If we have reasonable grounds to believe an individual has been, is, or may be victim of financial abuse
* If it is reasonable to expect collection and use with consent would compromise the availability or the accuracy of the information and the collection is reasonable for purposes related to investigating a breach of an agreement or a contravention of the laws of Canada or a province
* If disclosure is required to comply with a subpoena, warrant, court order, or rules of the court relating to the production of records
* If it was produced by an individual in the course of their employment, business, or profession and the collection is consistent with the purposes for which the information was produced
* If the collection is solely for journalistic, artistic, or literary purposes
* If the information is publicly available and is specified by the regulations

\
**4. WHEN AND WITH WHOM DO WE SHARE YOUR PERSONAL INFORMATION?**\
***In Short:** We may share information in specific situations described in this section and/or with the following third parties.*\
**Vendors, Consultants, and Other Third-Party Service Providers.** We may share your data with third-party vendors, service providers, contractors, or agents ("**third parties**") who perform services for us or on our behalf and require access to such information to do that work. The third parties we may share personal information with are as follows:

* **Web and Mobile Analytics**

Google Analytics

\
**5. DO WE USE COOKIES AND OTHER TRACKING TECHNOLOGIES?**\
***In Short:** We may use cookies and other tracking technologies to collect and store your information.*\
We may use cookies and similar tracking technologies (like web beacons and pixels) to access or store information. Specific information about how we use such technologies and how you can refuse certain cookies is set out in our Cookie Notice.

\
**6. HOW LONG DO WE KEEP YOUR INFORMATION?**\
***In Short:** We keep your information for as long as necessary to fulfill the purposes outlined in this privacy notice unless otherwise required by law.*\
We will only keep your personal information for as long as it is necessary for the purposes set out in this privacy notice, unless a longer retention period is required or permitted by law (such as tax, accounting, or other legal requirements).\
When we have no ongoing legitimate business need to process your personal information, we will either delete or anonymize such information, or, if this is not possible (for example, because your personal information has been stored in backup archives), then we will securely store your personal information and isolate it from any further processing until deletion is possible.

\
**7. HOW DO WE KEEP YOUR INFORMATION SAFE?**\
***In Short:** We aim to protect your personal information through a system of organizational and technical security measures.*\
We have implemented appropriate and reasonable technical and organizational security measures designed to protect the security of any personal information we process. However, despite our safeguards and efforts to secure your information, no electronic transmission over the Internet or information storage technology can be guaranteed to be 100% secure, so we cannot promise or guarantee that hackers, cybercriminals, or other unauthorized third parties will not be able to defeat our security and improperly collect, access, steal, or modify your information. Although we will do our best to protect your personal information, transmission of personal information to and from our Services is at your own risk. You should only access the Services within a secure environment.

\
**8. WHAT ARE YOUR PRIVACY RIGHTS?**\
***In Short:** In some regions, such as the European Economic Area (EEA), United Kingdom (UK), and Canada, you have rights that allow you greater access to and control over your personal information. You may review, change, or terminate your account at any time.*\
In some regions (like the EEA, UK, and Canada), you have certain rights under applicable data protection laws. These may include the right (i) to request access and obtain a copy of your personal information, (ii) to request rectification or erasure; (iii) to restrict the processing of your personal information; and (iv) if applicable, to data portability. In certain circumstances, you may also have the right to object to the processing of your personal information. You can make such a request by contacting us by using the contact details provided in the section "HOW CAN YOU CONTACT US ABOUT THIS NOTICE?" below.\
We will consider and act upon any request in accordance with applicable data protection laws. If you are located in the EEA or UK and you believe we are unlawfully processing your personal information, you also have the right to complain to your local data protection supervisory authority. You can find their contact details here: <https://ec.europa.eu/justice/data-protection/bodies/authorities/index_en.htm>.\
If you are located in Switzerland, the contact details for the data protection authorities are available here: <https://www.edoeb.admin.ch/edoeb/en/home.html>.\
**Withdrawing your consent:** If we are relying on your consent to process your personal information, which may be express and/or implied consent depending on the applicable law, you have the right to withdraw your consent at any time. You can withdraw your consent at any time by contacting us by using the contact details provided in the section "HOW CAN YOU CONTACT US ABOUT THIS NOTICE?" below.\
However, please note that this will not affect the lawfulness of the processing before its withdrawal nor, when applicable law allows, will it affect the processing of your personal information conducted in reliance on lawful processing grounds other than consent.\
**Cookies and similar technologies:** Most Web browsers are set to accept cookies by default. If you prefer, you can usually choose to set your browser to remove cookies and to reject cookies. If you choose to remove cookies or reject cookies, this could affect certain features or services of our Services. To opt out of interest-based advertising by advertisers on our Services visit <http://www.aboutads.info/choices/>.

\
**9. CONTROLS FOR DO-NOT-TRACK FEATURES**\
Most web browsers and some mobile operating systems and mobile applications include a Do-Not-Track ("DNT") feature or setting you can activate to signal your privacy preference not to have data about your online browsing activities monitored and collected. At this stage no uniform technology standard for recognizing and implementing DNT signals has been finalized. As such, we do not currently respond to DNT browser signals or any other mechanism that automatically communicates your choice not to be tracked online. If a standard for online tracking is adopted that we must follow in the future, we will inform you about that practice in a revised version of this privacy notice.

\
**10. DO CALIFORNIA RESIDENTS HAVE SPECIFIC PRIVACY RIGHTS?**\
***In Short:** Yes, if you are a resident of California, you are granted specific rights regarding access to your personal information.*\
California Civil Code Section 1798.83, also known as the "Shine The Light" law, permits our users who are California residents to request and obtain from us, once a year and free of charge, information about categories of personal information (if any) we disclosed to third parties for direct marketing purposes and the names and addresses of all third parties with which we shared personal information in the immediately preceding calendar year. If you are a California resident and would like to make such a request, please submit your request in writing to us using the contact information provided below.\
If you are under 18 years of age, reside in California, and have a registered account with Services, you have the right to request removal of unwanted data that you publicly post on the Services. To request removal of such data, please contact us using the contact information provided below and include the email address associated with your account and a statement that you reside in California. We will make sure the data is not publicly displayed on the Services, but please be aware that the data may not be completely or comprehensively removed from all our systems (e.g., backups, etc.).

\
**11. DO WE MAKE UPDATES TO THIS NOTICE?**\
***In Short:** Yes, we will update this notice as necessary to stay compliant with relevant laws.*\
We may update this privacy notice from time to time. The updated version will be indicated by an updated "Revised" date and the updated version will be effective as soon as it is accessible. If we make material changes to this privacy notice, we may notify you either by prominently posting a notice of such changes or by directly sending you a notification. We encourage you to review this privacy notice frequently to be informed of how we are protecting your information.

\
**12. HOW CAN YOU CONTACT US ABOUT THIS NOTICE?**\
If you have questions or comments about this notice, you may email us at <thedatascienceinterviewbook@gmail.com>


