Bivariate data, scatterplots and lines of fit: HSC Maths Standard 1 Year 12
“S3.2 Exploring and describing data arising from two quantitative variables: construct a scatterplot, identify the independent and dependent variables, describe the association (form, direction, strength), fit a line by eye and use it to make and assess predictions”
Plot bivariate data on a scatterplot with the independent variable horizontal, describe the association by form, direction and strength, draw a line of fit by eye and find its equation from two points on it, then predict. Interpolation is usually reasonable; extrapolation beyond the data can give unrealistic answers.
Jump to a section
What this dot point is asking
You need to explore and describe the relationship between two numerical variables: construct a scatterplot, describe the association, fit a line by eye, and use it to make predictions while judging their accuracy. NESA's topic guide suggests biometric data (height against arm span) and checking predictions against a person not in the original data.
The answer
Scatterplots
- The independent variable (explanatory) goes on the horizontal axis.
- The dependent variable (response) goes on the vertical axis.
- Each individual is one point. Digital tools (spreadsheets, graphing software) can construct scatterplots quickly.
Describing association
- Form: linear or non-linear.
- Direction: positive, negative or no association.
- Strength: strong (points close to a line), moderate, or weak.
- Mention any outliers.
Line of fit by eye
Draw a straight line that follows the trend, with about the same number of points above and below it. Choose two points on the line (not necessarily data points) to find its equation:
The gradient describes the average change in the dependent variable for each unit increase in the independent variable.
Making and assessing predictions
- Interpolation (within the data) is usually reasonable.
- Extrapolation (outside the data) is risky.
- Check a prediction against an actual measurement for someone not in the data set, and consider how scattered the points are.
Arm span (cm) and height (cm) for a class give a line of fit through (150, 152) and (180, 181).
- .
- (to the nearest whole number), so .
- Predict for arm span 170 cm: cm.
- A student not in the data with arm span 170 cm is 168 cm tall, so the prediction is 3 cm too high: reasonably accurate.
- Swapping the axes
- Independent on the horizontal axis.
- Joining the dots
- A line of fit is one straight line through the trend.
- Saying one variable causes the other
- Association is not causation.
Practice questions
Original practice questions graded from foundation to exam level, each with a full worked solution. Try them before revealing the solution.
foundation3 marksData on hours of sleep and test score for 10 students shows points rising from left to right, fairly close to a straight line. Identify the independent variable and describe the association.Show worked solution →
Independent variable: hours of sleep (it may help explain the test score).
Association: positive (more sleep tends to go with higher scores), linear, and moderate to strong (points are fairly close to a line).
Marking guide: 1 mark for the independent variable, 2 marks for describing direction, form and strength.
core4 marksA line of fit for age (a years, 11 to 16) and height (h cm) of boys passes through (11, 145) and (15, 169). Find its equation and predict the height of a 13-year-old.Show worked solution →
Gradient .
Using with (11, 145): , so .
Equation: .
Prediction: cm.
Marking guide: 1 mark for the gradient, 1 mark for the intercept, 1 mark for the equation, 1 mark for the prediction.
exam4 marksUsing the model h = 6a + 79 from the previous question, predict the height of a 30-year-old and comment on the result.Show worked solution →
cm.
A height of 2.59 m is unrealistic. The data only covered ages 11 to 16, when boys are growing quickly, so using the line for age 30 is extrapolation. Growth slows and stops in the late teens, so the linear model does not apply outside the data range.
Marking guide: 1 mark for the calculation, 1 mark for identifying the result as unrealistic, 2 marks for explaining extrapolation and the model's domain.