Data integrity, data warehousing and big data in visualisation: HSC Enterprise Computing Data Visualisation
“Assess data integrity in the development of a data visualisation, including ownership, source, validation and risk; explain the impact of enterprise data warehousing on data visualisation, including analysis and use of historical data trends and patterns, correlation with current data and data refinement/optimisation; explain how big data affects the design and development of data visualisation, including scope of visible information and types and depth of insight provided by the data”
Assess the integrity of data behind a visualisation by its ownership, source, validation and risk. An enterprise data warehouse provides cleaned historical data for trends, correlation with current data and fast, refined summaries. Big data forces designers to aggregate and filter what is visible, while enabling deeper and more varied insight.
Jump to a section
What this dot point is asking
A visualisation is only as trustworthy as its data. You need to assess data integrity using four named factors, explain what an enterprise data warehouse contributes, and explain how the sheer size of big data changes what a visualisation can show and how it must be designed.
The answer
Assessing data integrity
Data integrity means data is accurate, complete, consistent and trustworthy throughout its life. Assess it using the four factors NESA names:
- Ownership: who owns the data, who may use and publish it, and under what licence or privacy obligations?
- Source: where did it come from, how was it collected, and is the source authoritative, current and unbiased?
- Validation: has the data been checked against rules (data type, range, format, presence, consistency with other records)? Are missing values and duplicates handled?
- Risk: what could go wrong (errors introduced during import or transformation, tampering, outdated data, a misleading aggregate), how likely is it, and what is the impact of a wrong decision?
The impact of enterprise data warehousing
A data warehouse integrates cleaned data from many systems over many years. For visualisation this enables:
- Analysis and use of historical data trends and patterns: long-term trends, seasonality and year-on-year comparisons.
- Correlation with current data: placing today's figures against history and against other datasets (sales against marketing spend), so unusual results stand out.
- Data refinement/optimisation: ETL removes duplicates and standardises codes and units; pre-aggregated summaries (OLAP) make dashboards fast and consistent across the organisation (one version of the truth).
How big data affects visualisation design
- Scope of visible information: a screen can only show so much. Designers aggregate (averages, totals), filter, sample, and use dense forms (heat maps, hexbin maps) with interactive drill-down so users can move from overview to detail.
- Types and depth of insight: big data supports finer-grained insights (individual stores, hours, customer segments), combines varied sources (text, location, sensors) and supports predictive visuals. The risk is overload and false patterns, so designers must focus on the question the audience needs answered.
A hospital group builds a dashboard of emergency department waiting times.
- Ownership: the data belongs to the hospital group and includes patient information, so only aggregated, de-identified figures appear.
- Source: timestamps come from the triage system; staff sometimes enter times late, so the data is compared with automated door-entry logs.
- Validation: waiting times below zero or above 48 hours are flagged for review.
- Warehouse: five years of history show winter peaks, so this week's figures are compared with the same weeks in past years.
- Big data design: hundreds of thousands of visits are summarised as median waits per hospital per hour on a heat map, with drill-down to each hospital.
- Treating data integrity as only security
- It includes accuracy, completeness and consistency, not just protection from tampering.
- Plotting everything
- With big data, aggregation and filtering are design decisions, not shortcuts.
- Ignoring ownership
- Publishing a visual from data you do not have rights to is an integrity and legal problem.
Practice questions
Original practice questions graded from foundation to exam level, each with a full worked solution. Try them before revealing the solution.
foundation3 marksA student builds a chart of local rainfall using data copied from a website. Identify three integrity checks they should make.Show worked solution →
- Source: is the website authoritative (the Bureau of Meteorology) or an unknown blog?
- Ownership and licence: may the data be reused, and how must it be attributed?
- Validation: are there missing days, impossible values (negative rainfall) or unit mix-ups (mm and inches)?
Marking guide: 1 mark per relevant check.
core4 marksExplain how an enterprise data warehouse improves a sales dashboard compared with visualising data straight from the point-of-sale system.Show worked solution →
The point-of-sale system holds recent transactions in its own format. A data warehouse integrates years of cleaned sales data with stock, marketing and customer data, so the dashboard can:
- show historical trends and patterns (seasonal peaks over five years);
- correlate with current data (compare this week's sales with the same week in past years and with current promotions);
- use refined and optimised data (duplicates removed, consistent product codes, pre-aggregated summaries) so charts are accurate and load quickly.
Marking guide: 1 mark for the limitation of the operational system, 1 mark each for historical trends, correlation and refinement.
exam6 marksA transport authority wants to visualise every tap-on and tap-off from its ticketing system (tens of millions per month). Explain how big data affects the design of this visualisation and assess the integrity issues.Show worked solution →
- Scope of visible information
- Tens of millions of trips cannot be plotted individually. The design must aggregate (trips per station per hour), filter (one line or one day) and use suitable forms such as heat maps and flow maps, with drill-down for detail.
- Types and depth of insight
- The volume allows detailed insights: peak loads by station and minute, travel patterns between suburbs, and the effect of timetable changes. Combined with other data (weather, events), it can reveal causes and support predictions of crowding.
- Integrity: ownership and privacy
- The authority owns the data but it describes individuals' movements, so visuals must be aggregated to prevent anyone being tracked.
- Integrity: source and validation
- Missing tap-offs, faulty readers and duplicate taps must be detected and handled (validation rules, estimates clearly labelled).
- Integrity: risk
- Errors in aggregation or out-of-date data could lead to poor timetable decisions, and a breach of the raw data would be serious, so access must be controlled.
- Assessment
- With careful aggregation, validation and privacy protection, the visualisation can deliver deep insight; without them, its conclusions and the public's trust are at risk.
Marking guide: 2 marks for big data design effects (scope and depth), 3 marks for integrity issues (ownership, source/validation, risk), 1 mark for an overall assessment.