Data Science
13 dot points across 3 inquiry questions, in syllabus order. Each dot point has a focused answer, with exam-style questions and worked answers where available.
Collecting, storing and analysing data
How a blockchain works (blocks, hashes, a distributed ledger and consensus), why that makes records tamper-evident, and how it is applied to online voting, online identities, tracking items of value and recordkeeping, with the limitations you need for evaluate questions.
Big data defined by volume, variety and velocity; how a data warehouse integrates data for analysis (ETL, historical, subject-oriented); the benefits and risks of data mining; and the impact of data scale on storage, streaming, machine learning, human behaviour and digital footprints.
Quantitative data is numeric and measurable; qualitative data describes qualities in words or categories. How each is stored as data types (integer, real, string, Boolean, date), and how nominal, ordinal, interval and ratio levels decide which statistics are meaningful.
How enterprises sample and collect data (manual or computerised, active or passive), how to judge primary and secondary data for relevance, accuracy, validity and reliability, and how errors, uncertainty, raw versus processed data and bias limit what data can tell you.
What informatics is and how it turns data into understanding, how to choose between graphs, infographics, dashboards, reports, network diagrams and maps, the difference between structured and unstructured datasets, and how likes, emoticons and memes work (and fail) as alternative feedback data.
How four everyday software features affect the privacy and security of data: autofill, public versus private connections, checkboxes (including pre-ticked consent) and terms of agreement, with the risks, the protections and how to evaluate a design.
How to evaluate local storage, cloud storage, portable storage media and data warehouses against criteria such as capacity, cost, speed, accessibility, security, reliability, scalability and legal location, with a worked recommendation for a small business.
Processing and presenting data
How statistical models (averages, regression, correlation, forecasting) and machine learning (supervised, unsupervised, training and testing) analyse big data and make predictions in enterprises, with the four types of analytics, examples, and the limits of prediction.
How to summarise, analyse and present data in a spreadsheet: functions and statistics, charts, what-if modelling (goal seek, data tables, scenarios), filtering, grouping and sorting, linking sheets, conditional formatting, forms and reports, and building a dashboard with pivot tables and slicers.
How a flat-file database works and why it causes redundancy, how to design a relational database with computational thinking (entities, primary and foreign keys, a data dictionary and user views), and how to sort and search with SQL, forms and reports.
Data quality
The ethical use of data in social and enterprise research (consent, de-identification, purpose), and the issues NESA lists: bias, accuracy, metadata, copyright and acknowledging sources, intellectual property including ICIP, permissions, privacy and cultural responsibility, and security.
How data that is selected (curated) and communicated to people changes what they do: data literacy, the effect of timeframes, signals such as notifications, streaks and ratings, the problem of data swamps, and how enterprises and governments educate users.
The Australian and NSW laws that govern collecting and handling data (Privacy Act 1988 and the Australian Privacy Principles, the Notifiable Data Breaches scheme, NSW privacy and health records laws, copyright and spam laws), the authorities that enforce them (OAIC, IPC NSW), and Indigenous data sovereignty principles.
