krit.club logo

Statistics and Probability - Correlation, quantitative handling, using technology-extended

Grade 9IB

Review the key concepts, formulae, and examples before starting your quiz.

🔑Concepts

•

Bivariate Data: This involves the analysis of two variables, usually denoted as xx (the independent or explanatory variable) and yy (the dependent or response variable), to determine if a relationship exists between them.

•

Scatter Plots: A graphical representation where each data point is plotted as (x,y)(x, y). This is the first step in identifying patterns or trends.

•

Correlation: Describes the nature of the relationship. It can be positive (both variables increase together), negative (one increases as the other decreases), or zero (no apparent relationship).

•

Strength of Correlation: Quantified by how closely the points cluster around a line. It is categorized as strong, moderate, or weak.

•

Pearson’s Correlation Coefficient (rr): A numerical value between −1-1 and +1+1 that measures the strength and direction of a linear relationship. r=1r = 1 indicates a perfect positive correlation, r=−1r = -1 a perfect negative correlation, and r=0r = 0 no linear correlation.

•

Line of Best Fit (Trend Line): A line drawn through the data points that best represents the linear trend. Using technology (GDC), this is found via 'Linear Regression' in the form y=mx+cy = mx + c or y=ax+by = ax + b.

•

Mean Point: The line of best fit must always pass through the mean point (xˉ,yˉ)(\bar{x}, \bar{y}), where xˉ\bar{x} is the mean of the xx-values and yˉ\bar{y} is the mean of the yy-values.

•

Interpolation vs. Extrapolation: Interpolation is making a prediction within the range of the data (usually reliable). Extrapolation is predicting outside the data range (often unreliable as the trend may not continue).

•

Causation: A strong correlation does not necessarily imply that one variable causes the change in the other; there could be a third 'lurking' variable involved.

📐Formulae

xˉ=∑xn\bar{x} = \frac{\sum x}{n}

yˉ=∑yn\bar{y} = \frac{\sum y}{n}

y=mx+cy = mx + c

−1≤r≤1-1 \le r \le 1

💡Examples

Problem 1:

A researcher records the number of hours students study (xx) and their final exam scores (yy). The data for 5 students is: (2,45),(4,55),(6,75),(8,80),(10,95)(2, 45), (4, 55), (6, 75), (8, 80), (10, 95). Calculate the mean point (xˉ,yˉ)(\bar{x}, \bar{y}) and determine the equation of the line of best fit using technology. Predict the score for a student who studies for 77 hours.

Solution:

  1. Calculate xˉ\bar{x}: xˉ=2+4+6+8+105=305=6\bar{x} = \frac{2+4+6+8+10}{5} = \frac{30}{5} = 6
  2. Calculate yˉ\bar{y}: yˉ=45+55+75+80+955=3505=70\bar{y} = \frac{45+55+75+80+95}{5} = \frac{350}{5} = 70
  3. Mean point: (6,70)(6, 70).
  4. Using technology (GDC Linear Regression), the equation is approximately y=6.25x+32.5y = 6.25x + 32.5.
  5. Prediction for x=7x = 7: y=6.25(7)+32.5=43.75+32.5=76.25y = 6.25(7) + 32.5 = 43.75 + 32.5 = 76.25

Explanation:

The mean point acts as the anchor for the line of best fit. The linear regression equation y=6.25x+32.5y = 6.25x + 32.5 shows that for every hour studied, the score increases by 6.256.25 marks. Predicting for 77 hours is an example of interpolation.

Problem 2:

The correlation coefficient between the age of a car (xx in years) and its value (yy in dollars) is found to be r=−0.92r = -0.92. Interpret this value in context.

Solution:

r=−0.92r = -0.92 indicates a strong, negative linear correlation between the age of the car and its value.

Explanation:

Since the value is close to −1-1, the relationship is strong. The negative sign indicates that as the age of the car increases (xx), its value (yy) tends to decrease.

Problem 3:

A set of data has a mean xx-value of xˉ=15\bar{x} = 15 and a line of best fit given by y=0.8x+5y = 0.8x + 5. Calculate the mean yy-value.

Solution:

Since the line of best fit MUST pass through the mean point (xˉ,yˉ)(\bar{x}, \bar{y}), we substitute xˉ=15\bar{x} = 15 into the equation: yˉ=0.8(15)+5\bar{y} = 0.8(15) + 5 yˉ=12+5\bar{y} = 12 + 5 yˉ=17\bar{y} = 17

Explanation:

The fundamental property of the least squares regression line is that it contains the point representing the arithmetic means of both variables.