krit.club logo

Statistics and Probability - Correlation, qualitative handling

Grade 10IB

Review the key concepts, formulae, and examples before starting your quiz.

🔑Concepts

•

Bivariate Data: This involves the study of the relationship between two different variables, usually denoted as xx (independent variable) and yy (dependent variable).

•

Scatter Diagrams: A graphical representation where data points (x,y)(x, y) are plotted to visualize the relationship or correlation between two variables.

•

Positive Correlation: As the value of xx increases, the value of yy also tends to increase. The points cluster around a line with a positive gradient.

•

Negative Correlation: As the value of xx increases, the value of yy tends to decrease. The points cluster around a line with a negative gradient.

•

No Correlation: There is no apparent relationship between the variables; the points are scattered randomly on the diagram.

•

Strength of Correlation: Described qualitatively as 'Strong' (points are very close to a straight line), 'Moderate', or 'Weak' (points are widely spread but still show a trend).

•

Line of Best Fit (By Eye): A straight line drawn through the center of the data points that best represents the trend. It should ideally pass through the mean point (xˉ,yˉ)(\bar{x}, \bar{y}).

•

Mean Point: The point representing the average of all xx values and the average of all yy values, denoted as (xˉ,yˉ)(\bar{x}, \bar{y}).

•

Interpolation and Extrapolation: Interpolation is making a prediction within the range of the given data; Extrapolation is making a prediction outside the range of data, which is generally less reliable.

•

Outliers: Data points that lie significantly far away from the general trend of the rest of the data.

📐Formulae

xˉ=∑xn\bar{x} = \frac{\sum x}{n}

yˉ=∑yn\bar{y} = \frac{\sum y}{n}

y=mx+cy = mx + c

m=y2−y1x2−x1m = \frac{y_2 - y_1}{x_2 - x_1}

💡Examples

Problem 1:

A student measures the heights (xx) and shoe sizes (yy) of 5 classmates. The data is: (150,36),(160,38),(170,40),(180,42),(190,44)(150, 36), (160, 38), (170, 40), (180, 42), (190, 44). Calculate the mean point (xˉ,yˉ)(\bar{x}, \bar{y}).

Solution:

First, find the mean of xx: xˉ=150+160+170+180+1905=8505=170\bar{x} = \frac{150 + 160 + 170 + 180 + 190}{5} = \frac{850}{5} = 170 Next, find the mean of yy: yˉ=36+38+40+42+445=2005=40\bar{y} = \frac{36 + 38 + 40 + 42 + 44}{5} = \frac{200}{5} = 40 The mean point is (170,40)(170, 40).

Explanation:

The mean point is calculated by dividing the sum of the coordinates by the number of data points. Any line of best fit drawn for this data should pass through (170,40)(170, 40).

Problem 2:

Describe the correlation if a scatter plot shows that as the outside temperature increases, the sales of hot chocolate decrease, and the points are very close to a straight line.

Solution:

The correlation is Strong Negative Correlation.

Explanation:

It is 'Negative' because as one variable increases, the other decreases. It is 'Strong' because the points lie very close to a straight line.

Problem 3:

Given a line of best fit y=0.5x+10y = 0.5x + 10, where xx is the number of hours studied and yy is the test score. Predict the score for a student who studied for 1515 hours.

Solution:

Substitute x=15x = 15 into the equation: y=0.5(15)+10y = 0.5(15) + 10 y=7.5+10y = 7.5 + 10 y=17.5y = 17.5

Explanation:

We use the linear equation derived from the line of best fit to perform interpolation and estimate the dependent variable yy based on the independent variable xx.