Skip to main content

Machine learning and statistical modelling in data analytics: HSC Enterprise Computing Data Science

Syllabus dot point

“Explore how machine learning and statistical modelling are used in data analytics to analyse big data, and as a prediction tool”

HSCEnterprise ComputingData Science7 min read

Quick answer

Statistical modelling (regression, correlation, forecasting) and machine learning (supervised and unsupervised learning trained on large datasets) turn big data into predictions such as demand forecasts, fraud alerts and churn risk. They move analytics from describing the past to predicting and prescribing, but depend on data quality and can overfit or fail when conditions change.

Jump to a section
  1. What this dot point is asking
  2. The answer
  3. Practice questions

What this dot point is asking

You need to explain how enterprises use statistical modelling and machine learning (ML) to analyse big data and to predict. Know the basic methods, where each is used, and the limits of prediction.

The answer

From description to prediction

Type of analytics Question Example
Descriptive What happened? Sales fell 8% last quarter
Diagnostic Why did it happen? Sales fell most where a competitor opened
Predictive What will happen? Sales next quarter are forecast to fall 3%
Prescriptive What should we do? Run a promotion in the affected suburbs

Statistical modelling and ML move an enterprise from descriptive towards predictive and prescriptive analytics.

Statistical modelling

  • Measures of centre and spread summarise data.
  • Correlation measures how strongly two variables move together.
  • Regression fits a line or curve (sales = a + b times advertising spend) to estimate relationships and predict outputs.
  • Time series forecasting extends trends and seasonal patterns into the future.

Statistical models are transparent and explainable, which helps decision-makers trust them, but they assume a fixed form of relationship.

Machine learning

ML algorithms learn patterns from data rather than following hand-written rules.

  • Supervised learning uses labelled examples to predict a label (classification: fraud or not) or a number (regression: house price).
  • Unsupervised learning finds structure in unlabelled data (clustering customers into segments, detecting anomalies).
  • Training and testing: fit the model on training data, measure accuracy on unseen test data, then deploy and monitor.

ML suits big data because it handles many variables, large volumes and unstructured inputs (text, images) that simple models cannot.

Enterprise uses

  • Recommendation engines (streaming, online shopping)
  • Fraud detection and credit scoring
  • Demand forecasting and dynamic pricing
  • Predictive maintenance from sensor data
  • Customer churn prediction

Limits of prediction

  • Predictions are only as good as the data: bias in, bias out.
  • Overfitting makes models look accurate on training data but fail on new data.
  • Past patterns may not hold when conditions change.
  • Some models are "black boxes" that are hard to explain, which matters when decisions affect people.
Worked example

A streaming service wants to predict which subscribers will cancel next month.

  1. Data: viewing hours, days since last login, number of profiles, support tickets, plan type, and whether each past subscriber cancelled (the label).
  2. Model: supervised classification, trained on 80% of past subscribers and tested on the other 20%.
  3. Result: the model flags subscribers whose viewing has dropped sharply and who contacted support recently.
  4. Action (prescriptive): send those subscribers a personalised recommendation list and a discount offer.
  5. Check: compare churn among flagged subscribers who got the offer with a control group.
Common traps
Saying ML is always better than statistics
Simple models are often good enough and easier to explain.
Forgetting test data
Accuracy must be measured on data the model has not seen.
Treating predictions as certain
Always mention uncertainty and monitoring.

Practice questions

Original practice questions graded from foundation to exam level, each with a full worked solution. Try them before revealing the solution.

foundation2 marks
A bike shop's monthly sales rise by about 40 bikes each month. Explain how a simple statistical model could predict next month's sales.
Show worked solution →

Fit a linear trend (regression line) to past monthly sales, with month number as the input and sales as the output. The slope is about 40 bikes per month, so the model predicts next month's sales as this month's trend value plus about 40. It assumes the trend continues and ignores seasonal effects.

Marking guide: 1 mark for the model, 1 mark for how it predicts, with an assumption noted.

core4 marks
An online bank wants to flag fraudulent card transactions. Explain how supervised machine learning could be used, including the role of training and test data.
Show worked solution →

The bank gathers historical transactions labelled "fraud" or "genuine", with features such as amount, location, time, merchant type and distance from the last transaction. It trains a classification model on most of this data so it learns which feature patterns go with fraud.

It then tests the model on held-back labelled transactions it has not seen, measuring how many frauds it catches and how many genuine transactions it wrongly blocks. Once accurate enough, the model scores live transactions in real time, and high-risk ones are blocked or sent for checking. It is retrained as fraud patterns change.

Marking guide: 1 mark for labelled data and features, 1 mark for training, 1 mark for testing and accuracy measures, 1 mark for live use and retraining.

exam6 marks
A supermarket chain uses machine learning on big data to predict demand for fresh produce at each store. Evaluate this use of prediction.
Show worked solution →
How it works
The model combines years of sales history, weather forecasts, public holidays, local events and promotions to predict daily demand per product per store, then orders stock automatically.
Benefits
Big data captures patterns people would miss (a heatwave lifts watermelon sales in coastal stores), so predictions are more accurate than a manager's estimate. This reduces food waste and empty shelves and saves staff time.
Limitations
Predictions rely on past patterns, so unusual events (a pandemic, a supply shock) make them unreliable. Poor-quality or biased data (stores with missing records) produces poor predictions. Complex models can be hard to explain, so staff may not trust or be able to correct them.
Judgement
Machine learning is effective for routine demand forecasting at this scale, provided humans monitor its predictions, can override them for unusual events, and the model is regularly retrained and checked for accuracy.

Marking guide: 1 mark for how it works, 2 marks for benefits, 2 marks for limitations, 1 mark for a justified judgement.

Practise this

Sources & how we know this

ExamExplained