Residuals and residual analysis for VCE General Mathematics Unit 3 Data analysis
“Calculate residuals, construct a residual plot, and use residual analysis to test the assumption of linearity and check the quality of fit of a least squares line”
A residual is actual minus predicted: positive means the point is above the line, negative means below, and the residuals of a least squares line sum to zero. A residual plot graphs residuals against the explanatory variable; random scatter supports a linear model, while a clear pattern such as a curve shows the association is non-linear and calls for a transformation, no matter how high is.
What this dot point is asking
A least squares line is a model, and every model makes errors. A residual measures the error for one data point. VCAA wants you to calculate residuals, interpret their sign in context, construct a residual plot, and use the pattern in that plot to decide whether a linear model is appropriate. This is the study design's "residual analysis to test the assumption of linearity". If the residual plot shows a clear curve, the association is non-linear and the next step is a data transformation.
Residuals appear in almost every Examination 2 data analysis section. In 2025 students had to show that a missing residual was 27 984 and plot it; in the same question of students lost the mark for the correlation coefficient by forgetting its sign. Residual questions reward careful, fully written arithmetic.
The answer
What a residual is
For each data point, the least squares line makes a predicted value . The residual is the difference between what actually happened and what the line predicted:
This formula is on the VCAA formula sheet. Residuals are measured in the units of the response variable.
- Positive residual: the actual value is above the line; the model under-predicted.
- Negative residual: the actual value is below the line; the model over-predicted.
- Zero residual: the point lies exactly on the line.
For example, a kiosk's least squares line is drinks temperature. On a 30 °C day the line predicts drinks, but the kiosk sold 93. The residual is : the kiosk sold 4.2 fewer drinks than predicted, and the point sits 4.2 units below the line.
Rearranging the definition is also common in exam questions: actual predicted residual, and predicted actual residual.
Residual actual predicted. For a least squares line the residuals always sum to zero, so a residual plot is always centred on the horizontal axis. A random residual plot supports a linear model; a clear pattern (especially a curve) means the association is non-linear, whatever the value of .
The residual plot
A residual plot is a scatterplot with:
- the explanatory variable on the horizontal axis (the same as the original scatterplot), and
- the residual for each point on the vertical axis.
The horizontal line at residual represents the least squares line itself, "straightened out". Points above it are above the line; points below are below. What matters is the pattern:
| Residual plot | Meaning | Action |
|---|---|---|
| Random scatter above and below zero, no pattern | The linear model captures the trend; the assumption of linearity is supported | Use the linear model |
| Curved pattern (U-shape or inverted U) | The association is non-linear; the line is systematically wrong in places | Transform the data and refit |
| Fanning out (spread grows or shrinks with ) | The typical size of errors changes with | Predictions are less reliable where the spread is large |
| One isolated large residual | A possible outlier | Check the data point; comment on its influence |
The two plots above come from two data sets with almost the same coefficient of determination: for the kiosk data (left) and for the car values (right). On the left the residuals jump above and below zero with no pattern, so the straight line is a good model. On the right the residuals are positive at both ends and negative in the middle: the car's value falls quickly at first and then more slowly, which is a curve (it is reducing balance depreciation), and a straight line cuts across the curve. The line over-predicts in the middle years and under-predicts at both ends. A high could not have told you this.
Why the residuals sum to zero
The least squares line is the line that makes the sum of the squared residuals as small as possible (that is where "least squares" comes from). A consequence is that the line passes through and that the positive and negative residuals balance exactly, so
This gives you a quick check on a table of residuals, and it explains why a residual plot always has points on both sides of the zero line. For the kiosk data the residuals are , and they add to exactly 0.
Residuals and the coefficient of determination
The coefficient of determination and the residual plot answer different questions.
- answers "how much of the variation in the response variable is explained by the linear association with the explanatory variable?" The more tightly the points hug the line, the smaller the residuals and the larger .
- The residual plot answers "is a straight line the right shape for this association?"
You need both. A curved association can have a very high (the car data), and a genuinely linear association can have a low if there is a lot of scatter (the Melbourne house price data in the 2025 question has ). The study design wording, "use of residual analysis to check quality of fit" and "test the assumption of linearity", is about the residual plot, not about .
What to do when the residual plot is curved
A curved residual plot means you should try a transformation (squared, or reciprocal, applied to one variable only), refit the least squares line to the transformed data, and draw a new residual plot. If the new residual plot is random, the transformed model is appropriate. The shape of the curve suggests which transformation to try; the data transformation page covers the choices in detail.
Describe the pattern in words when you are asked to comment on a residual plot: "The residual plot shows a clear curved pattern (positive, then negative, then positive), which indicates the association is non-linear, so the linear model is not appropriate."
Residuals on the CAS
After fitting a linear regression on the CAS, the residuals are stored automatically (usually in a variable named something like resid or stat.resid). Plot them against the explanatory variable in a new scatterplot to get the residual plot. In an exam you should still be able to calculate a single residual by hand, with working, because "show that" questions require it.
How exam questions ask about residuals
- "Calculate the residual for ..." Substitute into the line to get the predicted value, then actual minus predicted. Show both lines.
- "Show that the residual is ..." The answer is given, so the mark is entirely for the working. Write the predicted value calculation and the subtraction.
- "Plot this residual on the residual plot." Horizontal position is the explanatory variable, vertical position is the residual.
- "Does the residual plot support the assumption of linearity? Explain." Random, so yes; or clear pattern, so no, with the pattern described.
- "What is the actual value if the residual is ...?" Actual predicted residual.
- "The residual is negative. What does this mean in context?" The actual value was less than the model predicted.
Calculating and interpreting a residual
The kiosk line is drinks temperature. On a 28 °C day, 95 drinks were sold. Find and interpret the residual.
- Predicted
- .
- Residual
- .
- Interpretation
- Positive: the kiosk sold 3.6 more drinks than the line predicted for a 28 °C day; the point is above the line.
Marker's note: interpret in context and with the right direction. "The model under-predicted by 3.6 drinks" earns the mark; "the residual is positive" alone does not.
Working backwards from a residual
The line has a residual of at . Find the actual value.
Predicted. .
Actual. .
Marker's note: rearrange the formula in words first (actual predicted residual) to avoid a sign error.
A residual analysis that rejects the linear model
Car values (thousands of dollars) against age (years): (1, 30), (2, 25.5), (3, 21.7), (4, 18.4), (5, 15.7), (6, 13.3), (7, 11.3), (8, 9.6). The least squares line is value age with .
Compute the residuals. Using the CAS (or by hand):
| Age | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 |
|---|---|---|---|---|---|---|---|---|
| Residual | 1.74 | 0.12 | 0.31 | 1.48 |
- Check
- They add to (allowing for rounding).
- Read the pattern
- Positive, then negative, then positive: a U-shape. The association is non-linear. The line under-predicts the value of new and old cars and over-predicts middle-aged cars.
- Conclusion
- Despite , a linear model is not appropriate. A transformation of value would be a sensible next step, because the values fall by roughly the same percentage each year.
Marker's note: this is the exact trap VCAA sets: a strong and a patterned residual plot. Always let the residual plot decide linearity.
The 2025 examination residual, step by step
The line is sale price distance. A home 15.5 km from the city sold for $1 250 000. Show that its residual is 27 984.
- Predicted
- .
- Residual
- .
- Plotting it
- On the residual plot, the point goes at distance 15.5 and residual 27 984, just above the zero line.
Marker's note: in the 2025 report only of students earned this mark; the examiners wanted "all working that led to the given residual value".
- Predicted minus actual
- The residual is actual minus predicted. Reversing it flips every sign and every interpretation.
- Plotting residuals against the response variable
- The residual plot's horizontal axis is the explanatory variable.
- Relying on to judge linearity
- A curved association can have above 0.95. Only a residual plot tests linearity.
- Calling a small pattern "random" by default
- Look at runs of signs: several positives, then several negatives, then several positives is a curve, not random scatter.
- Skipping working in a "show that" question
- The answer is given; the marks are for the predicted value and the subtraction.
- Forgetting units and context
- A residual of means 4.2 fewer drinks than predicted, not "4.2 below".
For any residual calculation, write two lines: "predicted " and "residual actual predicted ". When asked whether a residual plot supports linearity, use this template: "The residual plot shows [random scatter / a clear curved pattern], so the assumption of linearity is [supported / not supported] and a linear model is [appropriate / not appropriate]." If it is not supported, name a transformation to try.
A line of best fit makes a guess for every point, and a residual is how far off each guess was: the real value minus the guess. If you plot all those misses and they bounce up and down with no pattern, the straight line is doing a fair job. But if the misses form a smile or a frown shape, the line is consistently guessing too high in some places and too low in others, which means the data really follows a curve. That is why checking the misses matters more than just checking how close the points are to the line overall.
Exam-style questions
Questions in the style of VCAA exam questions on this dot point, each with a worked answer. They are written by ExamExplained unless tagged "Past paper"; the year shows the paper a question is modelled on.
2025 VCAA-style5 marksFor three-bedroom homes sold in Melbourne, the least squares line for sale price (in dollars) against distance from the city centre (in km) is: sale price distance from city centre, and the coefficient of determination is 0.0806. (a) Calculate the correlation coefficient , to three decimal places. (b) Predict the sale price of a home in the city centre. (c) A home 15.5 km from the city centre sold for $1 250 000; its residual is missing from the residual plot. Show that the residual is 27 984. (d) Describe the strength and direction of the linear association.
Show worked answer →
(a) . The slope of the line is negative, so is negative: . (1 mark) The examiners reported that of students scored zero here, mostly by leaving the answer as positive 0.284.
(b) In the city centre the distance is 0: sale price dollars. (1 mark)
(c) Predicted value: .
Residual actual predicted . (1 mark) Because the answer was given, the mark needed every line of working; only earned it.
(d) Weak (since ) and negative. (2 marks; earned both.)
Source: VCAA 2025 General Mathematics Examination 2, Question 4 (parts b, c, e and f.i), and the 2025 examination report.
Practice questions
Original practice questions graded from foundation to exam level, each with a full worked solution. Try them before revealing the solution.
foundation2 marksA least squares line for drinks sold at a kiosk against the maximum temperature (°C) is drinks temperature. On a 30 °C day the kiosk sold 93 drinks. (a) Find the predicted number of drinks. (b) Find the residual and interpret its sign.
Show worked solution →
(a) Predicted drinks. (1 mark)
(b) Residual actual predicted . The residual is negative, so the kiosk sold 4.2 fewer drinks than the line predicted: the point lies below the least squares line. (1 mark)
foundation2 marksFor a data point, the predicted value from a least squares line is 58.3 and the residual is 6.2. (a) What is the actual value? (b) Is the data point above or below the line?
Show worked solution →
(a) Residual actual predicted, so actual predicted residual . (1 mark)
(b) A positive residual means the actual value is larger than predicted, so the point is above the line. (1 mark)
foundation1 markA residual plot shows the points scattered randomly above and below the horizontal axis with no clear pattern. What does this indicate? A. The association is non-linear. B. A linear model is appropriate. C. The correlation coefficient is close to zero. D. The data contains an outlier.
Show worked solution →
B. A random scatter of residuals with no pattern supports the assumption that the association is linear, so the least squares line is an appropriate model. A random residual plot says nothing about the size of (option C), and a single outlier would show up as one isolated large residual (option D).
core3 marksThe value of a car (in thousands of dollars) is recorded against its age (years): age 1, 2, 3, 4, 5, 6, 7, 8; value 30, 25.5, 21.7, 18.4, 15.7, 13.3, 11.3, 9.6. The least squares line is value age, with . (a) Find the residuals for ages 1, 4 and 8, to two decimal places. (b) The full set of residuals, in order of age, is positive, positive, negative, negative, negative, negative, positive, positive. What does this pattern indicate?
Show worked solution →
(a) Using the rounded line:
- Age 1: predicted ; residual .
- Age 4: predicted ; residual .
- Age 8: predicted ; residual .
(2 marks for all three.)
(b) The residuals form a curved (U-shaped) pattern: positive at both ends and negative in the middle. This shows the association is non-linear, so a linear model is not appropriate even though is very high. A transformation should be considered. (1 mark)
core2 marksExplain why the residuals from a least squares line always add to zero, and why this makes the residual plot centred on the horizontal axis.
Show worked solution →
The least squares line is fitted so that it passes through the point ; as a result the positive residuals (points above the line) exactly balance the negative residuals (points below the line), so their sum, and therefore their mean, is zero. (1 mark)
Because the mean residual is zero, the residuals on a residual plot are always spread around the horizontal zero line, with points both above and below it. A residual plot with every point above the axis would indicate an arithmetic error. (1 mark)
core2 marksA student says: "The coefficient of determination is 0.98, so the linear model must be a good fit and I don't need a residual plot." Explain why the student is wrong.
Show worked solution →
measures how much of the variation in the response variable is explained by the linear model, but a curved association can still have a very high (the car values in the previous question have yet are clearly curved). (1 mark)
Only the residual plot tests the assumption of linearity: a clear pattern (such as a curve) means the linear model is systematically wrong, over-predicting in some places and under-predicting in others, whatever the value of . (1 mark)
exam3 marksA least squares line has equation . For the data point : (a) find the residual; (b) the residual for another point, , is ; find ; (c) a residual plot for all the data shows the residuals fanning out, getting larger in size as increases. State one limitation this places on predictions made with the line.
Show worked solution →
(a) Predicted . Residual . (1 mark)
(b) Predicted at : . Actual . (1 mark)
(c) The spread of the actual values about the line grows as increases, so predictions at large values of are much less reliable (have larger typical errors) than predictions at small values of . (1 mark)
exam2 marksThe scatterplot of (response) against (explanatory) shows a positive association. The residual plot for the least squares line shows residuals that are negative for small , positive for middle values of and negative again for large . Which transformation would you try, and why?
Show worked solution →
The residuals form an inverted-U, meaning the data curves downward relative to the line: rises quickly at first and then flattens out (increasing and concave down). (1 mark)
Suitable transformations are those that stretch the upper end of or compress the upper end of : try (squared transformation on the response) or (or ) on the explanatory variable, refit the line and check that the new residual plot is random. (1 mark) The choice between them is confirmed by which transformed residual plot has no pattern.