Skip to main content

Big data, data warehousing, data mining and data scale: HSC Enterprise Computing Data Science

Syllabus dot point

“Explore the use of big data and data warehousing, considering volume, variety and velocity; explore the risks and benefits of data mining; analyse the impact of data scale, including volume of raw data, storage, real-time and continuous streaming, opportunities for machine learning (ML), changes in human behaviour and ethical implications, including digital footprints”

HSCEnterprise ComputingData Science8 min read

Quick answer

Big data is defined by volume, variety and velocity. A data warehouse integrates cleaned historical data through ETL for analysis. Data mining finds patterns that bring benefits such as fraud detection but risks privacy and bias. Growing data scale drives storage and streaming demands, enables machine learning, changes behaviour and enlarges digital footprints.

Jump to a section
  1. What this dot point is asking
  2. The answer
  3. Practice questions

What this dot point is asking

This group of content points is about data at enterprise scale. You need to describe big data using volume, variety and velocity, explain data warehousing, weigh the benefits and risks of data mining, and analyse what happens as the scale of data grows.

The answer

Big data: volume, variety, velocity

Big data is data too large, varied or fast-moving for traditional tools to handle.

  • Volume: terabytes to petabytes, from transactions, logs, sensors and media.
  • Variety: structured tables, semi-structured JSON, unstructured text, images, audio and video, from many sources.
  • Velocity: data arrives continuously and often must be acted on in real time (fraud checks, live traffic).

Other sources add veracity (trustworthiness) and value, but the syllabus names the first three.

Data warehousing

A data warehouse is a central repository built for analysis. It is:

  • integrated: data from sales, stock, HR and marketing systems is combined;
  • subject-oriented: organised around subjects such as customers or products;
  • historical (time-variant): years of data are kept to show trends;
  • non-volatile: data is loaded and read, not constantly updated.

Data gets there through ETL: extract from source systems, transform (clean, remove duplicates, convert formats and codes) and load. Warehouses separate heavy analysis from the operational databases that run the business.

Data mining: benefits and risks

Data mining uses algorithms to find hidden patterns in large datasets: classification (will this customer leave?), clustering (groups of similar customers), association rules (bought X, often buys Y) and anomaly detection (unusual transactions).

Benefits Risks
Fraud and anomaly detection Invasion of privacy, including re-identifying "anonymous" people
Targeted products and marketing Discrimination when patterns reflect biased data
Demand forecasting and efficiency Spurious correlations mistaken for causes
Medical and scientific discovery Function creep: data used beyond its original purpose

The impact of data scale

  • Volume of raw data: more data can mean better insight, but also more noise and more cleaning.
  • Storage: cost and energy rise; enterprises use cloud storage, compression, tiering and retention policies.
  • Real-time and continuous streaming: systems must process data as it arrives (stream processing, edge computing) rather than in overnight batches.
  • Opportunities for machine learning: large labelled datasets make ML models more accurate and enable new services (recommendations, predictive maintenance).
  • Changes in human behaviour: people change how they act when data is collected (fitness streaks, driving more carefully under monitoring) and platforms use data to shape behaviour.
  • Ethical implications and digital footprints: every click, purchase and location adds to a person's digital footprint, which can be combined to profile them, sometimes without meaningful consent.
Worked example

A supermarket mines five years of loyalty card data in its data warehouse.

  1. Pattern found: customers who start buying nappies also increase purchases of ready-made meals.
  2. Benefit: the store sends relevant offers and stocks meals near the baby aisle.
  3. Risk: the model can infer sensitive facts (a pregnancy) before a customer has told anyone, and targeted mail could reveal it to others in the household.
  4. Response: set rules on sensitive inferences, give customers control over personalised offers, and review models for unintended profiling.
Common traps
Confusing a data warehouse with a backup
A warehouse is for analysis of integrated historical data, not for restoring systems.
Listing Vs without examples
Always apply volume, variety and velocity to the scenario.
Treating correlation as cause
Data mining finds associations; it does not prove why they occur.

Practice questions

Original practice questions graded from foundation to exam level, each with a full worked solution. Try them before revealing the solution.

foundation3 marks
Using a ride-share company as an example, describe each of the three Vs of big data.
Show worked solution →
  • Volume: millions of trips, each with GPS points every few seconds, fares and ratings.
  • Variety: structured trip records, GPS streams, driver photos, in-app messages and payment data.
  • Velocity: location updates arrive continuously and must be processed in real time to match riders with nearby drivers and set surge prices.

Marking guide: 1 mark per V with a relevant example.

core4 marks
Explain why a retail chain would load its sales, stock and loyalty data into a data warehouse rather than analysing each operational database directly.
Show worked solution →

Each operational database is designed for fast day-to-day transactions and holds only current data in its own format. Running heavy analysis on them would slow down checkouts.

A data warehouse extracts data from all three systems, transforms it into consistent formats (the same product codes and dates) and loads it into one subject-oriented store that keeps years of history. Analysts can then query sales against stock and loyalty data together, find long-term trends and run reports without affecting the live systems.

Marking guide: 1 mark for limits of operational databases, 1 mark for ETL and integration, 1 mark for historical data, 1 mark for performance separation.

exam7 marks
Analyse the impact of data scale on a city's use of smart traffic cameras and sensors. Refer to storage, real-time streaming, machine learning, changes in human behaviour and ethical implications.
Show worked solution →
Storage
Thousands of cameras recording continuously generate petabytes of video. The city must decide what to keep (for example, only counts and number plates of flagged vehicles) and use tiered storage, with recent data on fast storage and older data archived or deleted, to control cost.
Real-time streaming
Sensor data arrives continuously and must be processed within seconds to change signal timing or send alerts about crashes. This needs edge processing near the cameras and high-bandwidth links.
Machine learning
The scale of data allows models to learn traffic patterns and predict congestion, adjusting signals before jams form. Large volumes of labelled examples improve accuracy.
Changes in human behaviour
Drivers who know they are monitored may slow down or change routes, which is a benefit, but constant surveillance can also create a chilling effect on lawful movement and protest.
Ethical implications
Number-plate and face data create detailed location histories (digital footprints) of residents who never consented individually. Risks include function creep (using data for purposes beyond traffic), data breaches and unequal enforcement. Clear retention limits, de-identification and public transparency are needed.
Overall
Data scale makes smarter, safer traffic management possible, but the same scale multiplies privacy and security risks, so the benefits depend on strong governance.

Marking guide: 1 mark per aspect analysed with the scenario (5 marks), 2 marks for linking scale to both benefits and risks with an overall conclusion.

Practise this

Sources & how we know this

ExamExplained