Data sampling, active and passive collection, and data quality: HSC Enterprise Computing Data Science
“Investigate data sampling, including manual and computerised methods of active and passive data collection; assess the relevance, accuracy, validity and reliability of primary and secondary data; examine the impact of errors, uncertainty and limitations in data, including data sources, raw data versus processed data and data bias”
Data is collected manually or by computer, actively (people knowingly provide it) or passively (captured in the background). Judge primary and secondary data by relevance, accuracy, validity and reliability, and account for errors, uncertainty, processing and bias before drawing conclusions.
Jump to a section
What this dot point is asking
You need to describe how data is gathered (manual or computerised, active or passive), judge whether a data source can be trusted using four criteria, and explain how errors, uncertainty and bias limit the conclusions an enterprise can draw.
The answer
Sampling
A population is every item of interest; a sample is the subset actually collected. Enterprises sample because collecting everything is expensive or impossible. Common methods:
- Random: every member has an equal chance of selection. Reduces bias.
- Stratified: the population is split into groups (age bands, regions) and each group is sampled in proportion.
- Systematic: every nth record or customer.
- Convenience or self-selected: whoever is easiest to reach or chooses to respond. Cheap but often biased.
Manual and computerised, active and passive
| Active (person knowingly provides data) | Passive (captured in the background) | |
|---|---|---|
| Manual | Paper survey, interview, focus group | Staff tally of foot traffic, observation notes |
| Computerised | Online form, app rating, poll | Cookies, clickstream logs, GPS, IoT sensors, loyalty card scans |
Passive computerised collection produces huge volumes of accurate, time-stamped data cheaply, but raises privacy concerns because people may not realise it is happening.
Judging primary and secondary data
Primary data is collected first-hand for the current purpose. Secondary data was collected by others (ABS census data, industry reports, open government datasets).
- Relevance: does it answer this question, for this population and time?
- Accuracy: is it correct and free from errors?
- Validity: does it measure what it claims to, and does it meet the rules for the field (correct type, range, format)?
- Reliability: would the method give consistent results if repeated, and is the source trustworthy?
Primary data is usually more relevant but costs more; secondary data is cheaper and faster but may be out of date or collected for a different purpose.
Errors, uncertainty and limitations
- Data sources. Know who collected the data, how and why. A source with a commercial interest may present data selectively.
- Raw versus processed data. Raw data is the original, unprocessed records. Processed data has been cleaned, aggregated or summarised. Processing makes data usable but can introduce errors (a bad formula) and hides the detail needed to check it.
- Errors and uncertainty. Transcription errors, sensor faults, missing values and rounding all add uncertainty. Good practice records how much, for example a margin of error on a survey.
- Data bias. Sampling bias (unrepresentative sample), response bias (leading questions, social desirability), measurement bias (a faulty instrument) and historical bias (past discrimination built into records).
A school canteen wants to know which new menu items students would buy. It places a paper survey on the counter at lunchtime.
- Method: manual, active, self-selected sample.
- Bias: only students who already buy lunch at the canteen are sampled, so students who bring lunch from home (the group the canteen most wants to attract) are missed.
- Improvement: a short online form sent to a random sample of students from every year group, plus passive point-of-sale data on current purchases.
- Quality check: relevance improves (whole student population), reliability improves (random sampling can be repeated), and validity improves if questions ask about intended purchases at stated prices.
- Using "accurate" and "valid" as synonyms
- A measurement can be accurate to the gram and still invalid if it measures the wrong thing.
- Assuming secondary data is unreliable
- ABS data is highly reliable; the usual problem is relevance or age.
- Forgetting that passive data has bias too
- App analytics only describe people who use the app.
Practice questions
Original practice questions graded from foundation to exam level, each with a full worked solution. Try them before revealing the solution.
foundation3 marksA supermarket collects data through a loyalty card app and a printed customer survey at the checkout. Classify each method as manual or computerised and active or passive, with a reason.Show worked solution →
Loyalty card app: computerised and mostly passive. Each scan automatically records what was bought, when and where, without the customer entering anything.
Printed survey: manual and active. Customers knowingly write their answers on paper, which staff then enter into a system.
Marking guide: 1 mark per correct classification, 1 mark for a reason that refers to how the data is captured.
core4 marksA council wants to know how residents use a new bike path. It considers (A) a counter sensor on the path and (B) state government cycling statistics from 2019. Assess each source for relevance and reliability.Show worked solution →
(A) Sensor counter (primary data). Highly relevant because it measures use of this exact path. Reliability is good if the sensor is calibrated and maintained, but it cannot tell bikes from scooters and may double-count, so it should be checked against a manual count.
(B) 2019 state statistics (secondary data). Low relevance, because it describes cycling across the state before the path existed. It comes from an authoritative source, so it is reliable for its own purpose, but it is out of date and not specific to the path.
Marking guide: 1 mark each for relevance and reliability of each source (4 marks).
exam6 marksAn online retailer surveys customers by email after delivery and concludes that 92% are satisfied. Evaluate this conclusion by discussing sampling method, data bias and the difference between raw and processed data.Show worked solution →
- Sampling method
- An email survey is self-selected: only customers who choose to respond are counted. Customers whose order never arrived may be less likely to receive or open the email, so the sample is unlikely to represent all customers.
- Data bias
- Response bias is likely: very happy or very unhappy customers respond more often, and wording such as "How much did you enjoy your delivery?" can lead answers. Non-response bias means the 92% reflects respondents, not customers.
- Raw versus processed data
- "92% satisfied" is processed data. It hides how satisfaction was defined (4 and 5 on a 5-point scale?), how many people responded (92% of 50 or of 5,000?) and whether incomplete responses were removed. Without the raw data, the figure cannot be checked.
- Judgement
- The conclusion is weak evidence of overall satisfaction. The retailer should report the response rate, publish the question wording and compare the survey with passive data such as return rates and complaint logs.
Marking guide: 2 marks for sampling analysis, 2 marks for bias, 1 mark for raw versus processed data, 1 mark for a justified judgement.