Review the key concepts, formulae, and examples before starting your quiz.
πConcepts
Bivariate Data: This involves the study of two variables to determine the relationship between them. The independent variable () is the explanatory variable, and the dependent variable () is the response variable.
Scatter Diagrams: A graphical representation of bivariate data. It helps identify the correlation (positive, negative, or none) and the strength (weak, moderate, or strong) of the relationship.
Pearsonβs Product-Moment Correlation Coefficient (): A numerical measure of the linear relationship between two variables. It ranges from to , where is a perfect positive correlation, is a perfect negative correlation, and indicates no linear correlation.
The Least Squares Regression Line: The line of best fit that minimizes the sum of the squares of the vertical offsets (residuals). It is expressed in the form or .
The Mean Point: The regression line always passes through the centroid or mean point .
Interpolation vs. Extrapolation: Interpolation is making a prediction within the range of the original data set (usually reliable). Extrapolation is making a prediction outside the range of data (often unreliable and should be treated with caution).
πFormulae
π‘Examples
Problem 1:
A group of 5 students recorded the number of hours they studied () and their test scores (): . Calculate the equation of the regression line on and the Pearson correlation coefficient .
Solution:
Using a Graphic Display Calculator (GDC) to perform Linear Regression:
- Input values into List 1 and values into List 2.
- Perform 'Linear Reg ()'.
- Results: , , .
The equation is .
Explanation:
The value means that for every additional hour studied, the score is predicted to increase by points. The value indicates a very strong positive linear correlation.
Problem 2:
Using the regression line , predict the test score of a student who studies for hours and state whether this is interpolation or extrapolation.
Solution:
Substitute into the equation:
The predicted score is . Since lies within the original range of values ( to ), this is interpolation.
Explanation:
Interpolation is generally considered reliable because the prediction stays within the observed boundaries of the data collected.