Data analysis and evaluating research (accuracy, precision, validity, errors, outliers): VCE Psychology Units 3 and 4
“The accuracy, precision, repeatability, reproducibility and validity of measurements; ways of organising, analysing and evaluating primary data to identify patterns and relationships, including sources of error and uncertainty; assumptions and limitations of investigation methodology and/or data generation and/or analysis methods; criteria used to evaluate the validity of measurements and psychological research; the nature of evidence that supports or refutes a hypothesis, model or theory”
Process data with percentages, percentage change (final minus initial, divided by initial, times 100), mean, median and mode, and treat standard deviation as a measure of spread. Put the IV on the x-axis; use bar charts for categorical IVs and line graphs for continuous ones. Accuracy is closeness to the true value (harmed by systematic error, which repetition does not fix); precision is agreement between measurements (harmed by random error, reduced by repetition and larger samples). Repeatability uses the same conditions, reproducibility changed ones. Internal validity means the study tests what it claims to; external validity means the results apply to similar people elsewhere. Account for outliers rather than deleting them, treat uncertainty qualitatively, and draw conclusions only as far as the design and sample allow.
What this dot point is asking
Every VCE Psychology exam includes data: a table, a graph or a described study. The key science skills ask you to process data (percentages, percentage change, mean, median, mode, and an understanding of standard deviation), present it appropriately, evaluate its quality (accuracy, precision, repeatability, reproducibility, validity, errors, uncertainty, outliers), and draw conclusions that the evidence actually supports. The Unit 4 Area of Study 3 key knowledge applies the same criteria to your own investigation and poster.
The study design defines these terms precisely, and VCAA marks against those definitions, so learn them in its words.
The answer
Types of data
- Primary data are collected by the researcher for their own investigation; secondary data were collected by someone else (for example, a literature review uses secondary data).
- Quantitative data are numerical (reaction times, ratings on a scale, counts); qualitative data are descriptive, in words or images (interview responses, diary descriptions).
- A questionnaire with a rating scale produces quantitative data even though it asks about feelings. Such self-report data are subjective, while physiological measures such as EEG readings are objective.
Processing quantitative data
- Percentage: the part divided by the whole, multiplied by 100.
- Percentage change measures the size of a change relative to where it started:
A negative result is a decrease.
- Mean: the sum of scores divided by the number of scores. It uses every score but is pulled by extreme scores.
- Median: the middle score when the scores are ordered (the average of the two middle scores if there is an even number). It is not affected by one extreme score.
- Mode: the most frequent score.
- Standard deviation: a measure of variability, showing how spread out the scores are around the mean. VCE requires an understanding, not a calculation: a larger standard deviation means scores are more spread out; a smaller one means they cluster closely around the mean.
Presenting data
- The IV goes on the horizontal (x) axis and the DV on the vertical (y) axis.
- Use a bar chart when the IV is categorical (groups or conditions), with bars that do not touch. Use a line graph when the IV is continuous (time, amount).
- Give every graph a title, labelled axes with units, and an appropriate scale. The 2024 examiners noted that many students drew a line graph for categorical data, which was not appropriate.
Evaluating the quality of data
| Term | Study design meaning | Linked error |
|---|---|---|
| Accuracy | How close a measurement is to the true value (not quantifiable: more or less accurate) | Reduced by systematic error |
| Precision | How closely a set of measurements agree with each other | Reduced by random error |
| Repeatability | Agreement between successive measurements under the same conditions (procedure, observer, instrument, location, short time) | |
| Reproducibility | Agreement between measurements under changed conditions (method, observer, instrument, location, time or culture) | |
| Validity | A measurement is valid if it measures what it is supposed to measure |
Internal validity: an investigation is internally valid if it investigates what it sets out or claims to investigate. It depends on the design, sampling and allocation, and on whether extraneous and confounding variables were controlled. If a study is not internally valid, its external validity is irrelevant.
External validity: the results can be applied to similar individuals in a different setting. It is increased by broad inclusion criteria and sampling techniques that produce a sample resembling the wider population.
Errors, uncertainty and outliers
- Random errors are unpredictable variations that produce a spread of readings, affecting precision. Reduce them by repeating measurements and calculating a new mean, increasing sample size, or refining the method.
- Systematic errors shift every reading from the true value by a consistent amount or proportion, affecting accuracy. Repeating the measurement does not fix them; better calibration and correct use of instruments does.
- Personal errors are mistakes, miscalculations and observer errors. They are not reported or analysed as error; the experiment is repeated correctly.
- Uncertainty is the lack of exact knowledge of the value being measured. It is present in all measurements and is often higher in psychology because many variables are psychological constructs. VCE requires only a qualitative treatment of uncertainty.
- Outliers are readings that lie a long way from the others. They must be analysed and accounted for, not automatically removed; repeating readings can help decide whether they reflect an error or a genuine result.
- Contradictory data (incorrect data) and incomplete data (missing answers or observations) should be identified, along with possible sources of bias.
From results to conclusions
- A conclusion is a statement based on the results of this particular study, and it says whether the evidence supports or refutes the hypothesis. The 2025 examiners distinguished a conclusion from an implication, which is a possible consequence or broader impact of the findings.
- Only a controlled experiment with good internal validity can support a cause-and-effect conclusion.
- Generalise only to the population the sample represents.
- Distinguish evidence (systematically collected data) from opinion (a personal view) and anecdote (an individual story).
The data
A student measures reaction times (ms) for seven classmates after a poor night's sleep: 250, 262, 245, 270, 255, 262, 410.
Step 1: descriptive statistics
Sum , so the mean ms. Ordered: 245, 250, 255, 262, 262, 270, 410, so the median is 262 ms and the mode is 262 ms.
Step 2: spot and handle the outlier
410 ms is far from the others. Before deciding what to do, check the logbook: was the participant distracted, or did the equipment lag? If a recording error is confirmed, repeat that measurement; if not, keep it and report the median as the typical value. Without 410, the mean would be ms, showing how strongly one outlier shifts the mean.
Step 3: evaluate
Seven classmates is a small convenience sample, which limits external validity, and there is no comparison with a rested condition, so no conclusion about the effect of sleep can be drawn.
Marker's note: show the working for every calculation and state units; a correct final value with no working often loses a mark.
- Swapping repeatability and reproducibility
- Same conditions (including the same participants) is repeatability; changed conditions, such as different participants, is reproducibility. In 2023, 60 per cent of students chose the reproducibility option on a repeatability question.
- Saying repeating measurements fixes systematic error
- It reduces random error only.
- Deleting outliers automatically
- The study design says outliers must be analysed and accounted for.
- Using "reliable" as a catch-all
- Use the study design's terms: accuracy, precision, repeatability, reproducibility, validity.
- Calling an implication a conclusion
- A conclusion answers the research question from this study's results; an implication looks beyond it.
- Drawing a line graph for categorical data
- Use a bar chart when the IV is a set of groups.
Exam-style questions
Questions in the style of VCAA exam questions on this dot point, each with a worked answer. They are written by ExamExplained unless tagged "Past paper"; the year shows the paper a question is modelled on.
2023 VCAA-style1 markIn a study, Dawes et al. compared people with aphantasia with control participants. If Dawes et al. wanted to test the repeatability of their results, they could conduct the investigation again using A. the same group of participants. B. a different visual imagery questionnaire. C. the same methodologies but with different participants. D. an additional group of people diagnosed with Alzheimer's disease.Show worked answer →
Answer: A. This is a 1 mark multiple-choice item.
Repeatability is the closeness of agreement between successive measurements made under the same conditions: same procedure, observer, instrument, location and participants, over a short time. Re-running the study with the same group keeps the conditions the same.
C was chosen by 60 per cent of students, but using different participants changes the conditions of measurement, so it tests reproducibility, not repeatability. B changes the instrument, and D changes the sample and the aim.
Source: VCAA 2023 VCE Psychology examination, Section A, Question 18 (first sentence of context added by us), and the 2023 examination report.
2023 VCAA-style1 markResearchers wanted to evaluate the impact of public health campaigns promoting strategies for mental wellbeing on the number of individuals using these strategies over time. What would be the most appropriate way for researchers to determine the impact of the campaigns? A. Calculate the percentage change in people using each strategy 12 months later. B. Calculate the mode of the number of individuals using a strategy at six-month intervals. C. Identify outliers of scores of the number of individuals using a strategy on a weekly basis. D. Assess the standard deviation of scores of the number of individuals using a strategy monthly.Show worked answer →
Answer: A. This is a 1 mark multiple-choice item.
The question is about change over time, and percentage change expresses the size of a change relative to the starting value, so it directly measures the campaigns' impact. The mode (B) only gives the most frequent value; outliers (C) are unusual scores, not a measure of impact; standard deviation (D) measures the spread of scores, not how much they changed.
Source: VCAA 2023 VCE Psychology examination, Section A, Question 9, and the 2023 examination report.
2025 VCAA-style1 markIn a study of partial sleep deprivation, participants' performance was measured in a driving simulator. The data from the driving simulator has been shown to be valid in several road safety studies. This means A. the data is precise. B. outliers have been removed from the dataset. C. the results of the research can be applied to all adults. D. the data is appropriate to assess the effects of sleep deprivation.Show worked answer →
Answer: D. This is a 1 mark multiple-choice item.
A valid measurement measures what it is supposed to measure, so valid simulator data are appropriate for assessing the effects of sleep deprivation on driving. Precision (A) is about how closely repeated measurements agree. Removing outliers (B) is a data-handling decision, not validity. C overstates external validity: results can only be generalised to people similar to the sample.
Source: VCAA 2025 VCE Psychology examination, Section A, Question 23 (stimulus summarised by us), and the 2025 examination report.
Practice questions
Original practice questions graded from foundation to exam level, each with a full worked solution. Try them before revealing the solution.
foundation2 marksA participant's anxiety rating fell from 40 before a mindfulness program to 30 after it. Calculate the percentage change, showing your working.Show worked solution →
1 mark correct substitution; 1 mark correct answer with sign and unit.
The anxiety rating decreased by 25 per cent.
core3 marksSeven participants' reaction times (in milliseconds) were 250, 262, 245, 270, 255, 262 and 410. Calculate the mean, median and mode, and state which best represents the typical reaction time.Show worked solution →
1 mark mean; 1 mark median and mode; 1 mark justified choice.
- Mean ms.
- Ordered: 245, 250, 255, 262, 262, 270, 410, so the median is 262 ms and the mode is 262 ms.
- The median (262 ms) best represents the typical time: the outlier of 410 ms pulls the mean upwards, while the median is not affected by one extreme score.
core3 marksDistinguish between random errors and systematic errors, and explain how each can be reduced.Show worked solution →
1 mark random errors: unpredictable variations that cause a spread of readings and affect precision; 1 mark systematic errors: shift all readings in one direction from the true value by a consistent amount or proportion and affect accuracy; 1 mark reduction: random errors by repeated measurements, a larger sample or a refined method; systematic errors by calibrating instruments and using them correctly, since repeating the measurement does not remove them.
exam4 marksIn a sleep study, the heart-rate monitor used for every participant was later found to read 5 beats per minute too high. Separately, three participants removed their sensors during the night, so their data are missing. Evaluate the effect of each problem on the quality of the data and suggest one improvement for each.Show worked solution →
Marks: 1 the monitor problem is a systematic error, making all readings inaccurate (shifted by the same amount) although they may still be precise; 1 improvement: calibrate the monitor against a reliable reference, or correct the data if the offset is known; 1 the missing sensor data are incomplete data, which reduce the sample size and may introduce bias if the participants who removed sensors differ from others; 1 improvement: secure sensors or check them during the night, and recruit a larger sample.
exam3 marksA study of 30 Year 12 students at one Melbourne school found that students who slept less than seven hours scored lower on a memory test. A newspaper reports: 'Lack of sleep causes poor memory in teenagers.' Evaluate this claim.Show worked solution →
1 mark the methodology appears correlational (sleep was not manipulated), so cause and effect cannot be concluded; a third variable such as stress or screen use could explain both; 1 mark a small sample from one school limits external validity, so the result cannot be generalised to all teenagers; 1 mark a judgement that the claim overstates the evidence, with what further evidence would help (for example, a controlled experiment with random allocation and a larger, stratified sample).
Practise this
Sources & how we know this
- 2023 VCE Psychology examination report (Section A, Questions 9 and 18) — VCAA
- 2024 VCE Psychology examination report (graphing feedback) — VCAA
- 2025 VCE Psychology examination report (Section A, Questions 8 and 23) — VCAA
- VCE Psychology Study Design (from 2023) — VCAA
- VCE Psychology: examination specifications, past examinations and reports — VCAA