Review the key concepts, formulae, and examples before starting your quiz.
🔑Concepts
Bivariate data involves two variables, usually denoted as (the independent/explanatory variable) and (the dependent/response variable).
A scatter diagram is used to visualize the relationship between two variables. Patterns can be described by their direction (positive or negative), form (linear or non-linear), and strength (strong, moderate, or weak).
Pearson's product-moment correlation coefficient () measures the strength and direction of a linear relationship. Its value ranges from to , where is a perfect positive linear correlation and is a perfect negative linear correlation.
Spearman's rank correlation coefficient () is used when data is ranked or when the relationship is monotonic but not necessarily linear. It is calculated based on the ranks of the data points.
Correlation does not imply causation. Two variables may be highly correlated due to a third 'lurking' variable or pure coincidence.
The regression line of on (least squares regression line) is the line that minimizes the sum of the squares of the vertical distances from the data points to the line. It always passes through the mean point .
Interpolation is the process of predicting a value within the range of the given data set and is generally reliable. Extrapolation is predicting outside the range and is often unreliable.
📐Formulae
💡Examples
Problem 1:
Given the following paired data for and : . Calculate the mean point and the Pearson correlation coefficient .
Solution:
First, calculate the means: The mean point is . Using a GDC (Graphic Display Calculator), the value of is approximately .
Explanation:
The mean point is found by averaging the values and values separately. For IB AI, the correlation coefficient is typically found using the 'LinReg' function on the GDC.
Problem 2:
Calculate Spearman's rank correlation coefficient for the following data sets which have been ranked from to : Rank : Rank :
Solution:
First, calculate the differences and : Sum of : Using and :
Explanation:
Spearman's rank correlation is calculated by finding the difference in ranks for each pair, squaring those differences, and applying the formula. suggests a moderate positive monotonic relationship.
Problem 3:
A regression line is found to be . If the range of values in the original data was , predict the value of when and state if this prediction is reliable.
Solution:
Substitute into the equation: The prediction is reliable because lies within the range .
Explanation:
Predicting values within the observed data range is called interpolation. Since is between and , the model is expected to be valid.