Inquiry Question 2: Do non-infectious diseases cause more deaths than infectious diseases?
Collect and represent data from secondary sources to evaluate the method used in an example of an epidemiological study, including incidence, prevalence, mortality, and the methods and benefits of epidemiology
A focused answer to the HSC Biology Module 8 dot point on epidemiology. Defines incidence, prevalence and mortality, compares cohort, case-control and cross-sectional study designs, and applies them to the Doll and Hill lung cancer studies.
Reviewed by: AI editorial process; not yet individually human-reviewed
Have a quick question? Jump to the Q&A page
Jump to a section
What this dot point is asking
NESA wants you to define the core epidemiological measures, describe the main study designs, evaluate a real epidemiological study, and explain how epidemiology informs public health.
The answer
Epidemiology is the study of the distribution, causes and control of disease in populations. It uses observational and experimental study designs to identify risk factors, estimate disease burden, and evaluate interventions.
Core measures
- Incidence
- New cases per population per time. Formula: . Reported per 100 000 per year. Tracks how fast a disease is emerging.
- Prevalence
- Existing cases at a point in time. Formula: . Reported as a percentage. Tracks total disease burden.
- Mortality
- Deaths per population per time. Crude mortality counts all deaths; cause-specific mortality counts deaths from a specific disease. Reported per 100 000 per year. Tracks lethality.
- Case fatality rate
- Deaths divided by diagnosed cases. Measures how deadly a disease is once contracted.
- Morbidity
- Total illness in a population, including non-fatal disease burden (often measured as DALYs, disability-adjusted life years).
The single hardest distinction to keep straight is incidence versus prevalence. Picture a bathtub: incidence is the water flowing in from the tap (new cases), recovery and death are the drain, and prevalence is the depth of water sitting in the tub right now (all existing cases). A chronic disease has a slow drain, so even a modest inflow fills the tub.
Study designs
- Cross-sectional study
- Measures prevalence and risk factors in a population at a single point in time. Useful for snapshots but cannot establish temporal sequence.
- Cohort study (prospective)
- Follows a group of healthy people forward in time, recording exposures and waiting for disease to develop. Strong for establishing temporal sequence and calculating incidence and relative risk. Example: the Framingham Heart Study (1948 onwards) identified cholesterol, smoking and hypertension as cardiovascular risk factors.
- Case-control study (retrospective)
- Compares people with the disease (cases) to matched people without (controls), looking backward at exposures. Efficient for rare diseases. Vulnerable to recall and selection bias.
- Randomised controlled trial (RCT)
- Participants are randomly assigned to intervention or control groups. The gold standard for testing whether an intervention causes an outcome. Used for treatment trials, less often for risk factor studies (cannot ethically assign people to smoke).
- Ecological study
- Compares disease rates across populations (e.g. fluoride in water versus dental caries). Cannot make individual-level claims (ecological fallacy).
The designs fall into two families. Descriptive studies just describe the pattern of disease (who, where, when) - cross-sectional snapshots and ecological comparisons. Analytical studies test a hypothesis about a cause by comparing groups - cohort and case-control. The key difference between the two analytical designs is the direction of enquiry: a cohort starts from exposure and looks forward to disease; a case-control starts from disease and looks back to exposure.
Worked example: Doll and Hill and lung cancer
In 1950, Richard Doll and Austin Bradford Hill published a case-control study of 1298 patients in London hospitals. Cases were lung cancer patients; controls were matched patients without lung cancer. Smoking history was recorded by interview.
Result. Smokers had a much higher rate of lung cancer than non-smokers, with a dose-response gradient: more cigarettes per day, higher cancer risk.
Follow-up. The British Doctors Study (1951 onwards) followed 40 000 male doctors prospectively for over 50 years. It confirmed:
- Lung cancer mortality 25 times higher in heavy smokers than non-smokers.
- Half of long-term smokers die from a smoking-related disease.
- Quitting at any age reduces risk.
Bradford Hill criteria. Hill later proposed nine criteria for inferring causation from observation: strength of association, consistency, specificity, temporality, biological gradient, plausibility, coherence, experiment and analogy. Smoking and lung cancer satisfied all nine.
Impact. The studies led to public health warnings, advertising restrictions, taxation, plain-packaging laws (in Australia from 2012), and a roughly two-thirds reduction in adult smoking rates in developed countries.
One of the most cited epidemiological graphs is the lagged parallel between cigarette consumption and lung-cancer deaths in 20th-century men: cigarette use rose from the 1900s, and lung-cancer mortality climbed too - but delayed by roughly two to three decades, because the disease takes that long to develop after exposure. That time lag, and the matching shapes of the two curves, are exactly the temporality and biological-gradient evidence the Bradford Hill criteria look for.
Benefits of epidemiology
- Identifies causes. Smoking and lung cancer, asbestos and mesothelioma, HPV and cervical cancer.
- Targets prevention. Identifies high-risk groups for screening (e.g. women over 50 for breast cancer).
- Evaluates interventions. Did the cervical cancer vaccine reduce incidence? (Yes, by over 50 percent in vaccinated cohorts.)
- Tracks emerging disease. Surveillance systems detect new outbreaks early (COVID-19, HIV).
- Allocates resources. Prevalence data informs hospital capacity, drug stockpiles and staffing.
Limitations of epidemiology
- Cannot prove causation in observational studies. Only RCTs can do that directly; observational studies use Bradford Hill criteria.
- Confounding. Hidden variables may explain associations.
- Bias. Selection bias, recall bias, reporting bias.
- Generalisability. A study in one population may not apply elsewhere.
Examples in context
Example 1. NSW BreastScreen and population-level mortality reduction. BreastScreen NSW invites women aged 50 to 74 for free biennial mammographic screening, with participation reaching about 55 percent of eligible women. Epidemiological evaluation by the Cancer Institute NSW compares age-adjusted breast cancer mortality between regular screeners and non-screeners, controlling for confounders such as socioeconomic status. Modelled data show a roughly 25 percent reduction in breast cancer mortality among regularly screened women. This is a population-level intervention evaluation using cohort study design, and it illustrates how prevalence (women living with breast cancer), incidence (new diagnoses) and mortality (deaths) move differently when screening is introduced: incidence rises initially due to detection, then mortality falls.
Example 2. Mater Mothers' Hospital and the COVID-19 maternal cohort study. During 2020-2022, the Mater Mothers' Hospital in Brisbane led an Australian prospective cohort study following 8500 pregnant women, tracking SARS-CoV-2 infection, vaccination and obstetric outcomes. Results published in The Lancet Regional Health showed vaccinated women had no increase in adverse pregnancy outcomes, while unvaccinated infected women had a 4.5-fold higher risk of preterm birth. The cohort design (following exposed and unexposed groups forward in time) was essential to establish temporal causation, which a cross-sectional snapshot could not have shown. The findings directly informed Royal Australian and New Zealand College of Obstetricians and Gynaecologists guidelines for COVID-19 vaccination in pregnancy.
Practice and recall drills for this dot point are in the question bank above (graded practice questions with full marking criteria, plus short fluency drills).
Exam-style practice questions
Practice questions written in the style of NESA exam questions on this dot point, with worked answer explainers. The year tag is the paper they imitate, not the source.
2025 HSC7 marks[A population lives across regions A, B and C; A and B are linked by a road bridge while C is isolated. A graph shows the risk (%) of developing an environmental disease according to age at exposure (10, 20, 30 years) over up to 55 years after exposure.] Design an epidemiological study that could be used to produce the results shown in the graph. Justify the features of your design.Show worked answer →
Band-marked: top marks design a study producing most features of the graph AND justify features with reference to the stimulus.
- Sample/cohort: survey individuals from age groups 10, 20 and 30 (matching the graph's exposure ages) across areas A, B and C, followed over ~55 years (the graph's time span).
- Use of a control: A and B mix freely via the bridge, but C is isolated so it can serve as a control group — justify this with the map.
- Validity features: include equal numbers of males and females; large sample size; and collect data on confounding factors (diet, exercise, exposure to disease-causing agents, lifestyle, general health) so risk can be attributed to age at exposure rather than other variables.
- Analysis: correlate the data to find common features/activities, linking exposure age to disease risk. Examiners reward designs that explicitly use the stimulus and show sound grasp of reliability and validity.
Source: NESA 2025 HSC Biology examination and marking guidelines.
2023 HSC7 marks[Air pollution has been linked to non-infectious neurological disorders. 500 people from each of three major cities, males and females aged 20–50, were monitored for 12 months; results gave the % of each sample with symptoms.] Evaluate the method used in this epidemiological study in determining a link between air pollution and the symptoms.Show worked answer →
An evaluate question — top marks need a thorough grasp of what makes an epidemiological study valid, a comprehensive analysis of THIS study, and an informed judgement.
Judgement: The study is not valid and does not establish a statistically significant cause–effect link.
Weaknesses to analyse (each earns credit):
- Confounding risk factors not controlled — age, sex, ethnic group and especially occupation (which could cause similar symptoms) are not accounted for; different ethnic groups are not indicated.
- Exposure not measured — no indication of locality within each city or proximity to industry; subjects should have varying, measured levels of exposure so greater exposure can be matched to greater symptom incidence.
- Time too short — 12 months may be insufficient for symptoms to develop.
- No measure of symptom severity, and the sample/geographic and socioeconomic range may be too narrow for reliable trends.
Conclude with a clear judgement that the design flaws mean any apparent link is unreliable.
Source: NESA 2023 HSC Biology examination and marking guidelines.
2022 HSC4 marks[A historical study followed non-smoking married women aged 40+ across 29 Japanese health districts for 14 years; lung-cancer mortality was assessed by husbands' smoking. Non-smoker women with non-smoker husbands: 8.7 per 100 000; with smoker husbands: 15.5; women who smoke: 32.8.] Evaluate the method used in this epidemiological study.Show worked answer →
Top marks (4) need a thorough understanding of the methodology plus a suitable judgement.
- Strengths
- The study used large numbers of women matched into three exposure categories and followed them for a long period (14 years). These features make the sample size and duration adequate for a valid study.
- Limitations
- The categories assume each woman spends similar time with a smoker and that each smoker smokes a similar amount; in reality exposure time and smoke volume vary greatly, which could compromise the data. (Longer follow-up would give even more definitive data.)
- Judgement
- Because the large cohort should average out the variation in individual exposure, the method is overall valid despite these limitations.
Source: NESA 2022 HSC Biology examination and marking guidelines.
2020 HSC4 marks[An 11-year study of 58 406 young adults related drinking-water arsenic exposure to mortality; survival graphs are shown for three exposure bands (<90, 90–223, >223 µg/L) in males and females.] The hypothesis was that exposure to arsenic in drinking water increases mortality in young adults. Discuss the data presented in the graphs in relation to this hypothesis.Show worked answer →
Marks come from making points for and/or against the hypothesis and relating each to the data.
Support for the hypothesis:
- In both sexes, increasing arsenic dose led to decreased survival, suggesting arsenic causes the decline; the dose–response was clearest in males.
- Survival declined progressively over the 11 years, consistent with cumulative exposure reducing survival.
Reference the control/qualifications:
- The <90 µg/L group had the highest survival and acts as a near-control; survival there was high even though this is above the WHO limit.
- In females, all doses >90 µg/L gave a similar decrease, hinting at other interacting factors (e.g. nutrition, genes).
- Note the magnitude caveat: despite the large sample, survival only dropped by ~0.1% or less, so the effect, while present, is small.
Source: NESA 2020 HSC Biology examination and marking guidelines.
Practice questions
Original practice questions graded from foundation to exam level, each with a full worked solution. Try them before revealing the solution.
foundation3 marksDefine incidence, prevalence and mortality, and state how each is calculated for a population.Show worked solution →
- 1 mark - incidence
- Incidence is the number of new cases of a disease arising in a population at risk over a set time. It is calculated as new cases ÷ population at risk, per unit time (commonly per 100 000 per year).
- 1 mark - prevalence
- Prevalence is the total number of existing cases (new plus pre-existing) in a population at a point in time. It is calculated as existing cases ÷ total population, usually given as a percentage or per 1000.
- 1 mark - mortality
- Mortality is the number of deaths in a population over a set time. It is calculated as deaths ÷ population, per unit time (commonly per 100 000 per year); cause-specific mortality counts deaths from one named disease.
The mark for each hinges on the distinction: incidence = NEW cases (a rate), prevalence = ALL existing cases (a proportion), mortality = deaths. Naming the measure without the calculation, or blurring incidence with prevalence, caps the response.
foundation2 marksExplain why a chronic disease such as type 2 diabetes can have a high prevalence but a relatively modest incidence, while an acute fatal disease can show the opposite pattern.Show worked solution →
1 mark - the chronic case. Because incidence counts only new cases while prevalence counts all existing cases, a long-lasting disease accumulates cases: people diagnosed years ago are still alive and still counted, so prevalence climbs well above the yearly incidence.
1 mark - the acute fatal case. An acute, rapidly fatal disease produces many new cases (high incidence) but sufferers either recover or die quickly, so few exist at any one moment and prevalence stays low.
The discriminator is duration: prevalence ≈ incidence × average disease duration. An answer that does not link the gap to how long cases persist does not earn the second mark.
core3 marksIn a NSW town of 50 000 people, 200 new cases of type 2 diabetes were diagnosed during 2025, and 3200 people were living with diabetes at the end of the year. Calculate (a) the annual incidence rate per 100 000, and (b) the point prevalence per 1000. Show your working.Show worked solution →
- 1 mark - incidence working
- Incidence rate = new cases ÷ population at risk = = 0.004 per person per year.
- 1 mark - incidence expressed per 100 000
- 0.004 × 100 000 = 400 new cases per 100 000 per year.
- 1 mark - prevalence
- Point prevalence = existing cases ÷ total population = = 0.064 = 64 per 1000 (6.4%).
Full marks require correct working AND correct units (per 100 000 per year for incidence, per 1000 for prevalence). A bare number with no denominator or time frame is penalised. Note the contrast the data makes plain: only 200 new cases this year, but 3200 people living with the disease, because diabetes is chronic.
core4 marksCompare cohort and case-control study designs. (a) Describe each design. (b) State one advantage and one limitation of each. (c) Justify which is more suitable for studying a rare disease such as mesothelioma.Show worked solution →
- 1 mark - describe each design
- A cohort study selects a group free of the disease, records their exposures, and follows them forward in time to see who develops the disease. A case-control study starts with people who already have the disease (cases) and matched people without it (controls), then looks backward at past exposures.
- 1 mark - advantages/limitations
- Cohort: can calculate incidence and relative risk and establishes temporal sequence, but is slow and expensive and poor for rare diseases. Case-control: fast, cheap and efficient for rare diseases, but is vulnerable to recall and selection bias and cannot directly give incidence.
- 1-2 marks - justified judgement
- For a rare disease such as mesothelioma, a case-control design is more suitable: a cohort would have to follow an enormous number of people for decades to capture enough cases, whereas a case-control study starts with existing cases, so it gathers the needed numbers quickly and cheaply.
The judgement mark requires an explicit reason tied to rarity (cases are pre-assembled), not just "case-control is faster".
core5 marksA researcher reports that across 20 countries, nations with higher average chocolate consumption also have more Nobel laureates per capita, and concludes that eating chocolate improves cognitive performance. Analyse this claim using epidemiological reasoning.Show worked solution →
- 1 mark - identify the design
- This is an ecological study: it compares population-level averages (chocolate per capita, laureates per capita) across countries, not individuals.
- 1 mark - correlation is not causation
- A statistical association between two population variables does not establish that one causes the other; the data are observational.
- 1 mark - confounding
- A confounding variable plausibly drives both: national wealth / education spending raises both chocolate affordability and research output. Wealth, not chocolate, may explain the link.
- 1 mark - the ecological fallacy
- Population-level correlation cannot be used to infer individual-level behaviour - we have no data showing that the individuals who ate chocolate are the ones who won prizes.
- 1 mark - what would be needed
- To test causation you would need an individual-level design (cohort or, ideally, a randomised controlled trial) that controls for confounders such as wealth and education.
Band 6 answers name the ecological fallacy AND a specific confounder, and propose a controlled individual-level study. Merely saying "correlation is not causation" caps at 2 marks.
exam6 marksDescribe the features of a well-designed epidemiological study, and explain how each feature improves the reliability or validity of the conclusions drawn.Show worked solution →
Award up to 6 marks for naming design features AND linking each to reliability (reproducibility) or validity (measuring what was intended).
- Large sample size (1 mark)
- A large sample averages out random individual variation and increases statistical power, improving reliability so the result is unlikely to be a chance finding.
- Control or comparison group (1 mark)
- An unexposed (or lower-exposure) comparison group lets the effect of the exposure be isolated; without it there is no baseline, so any apparent effect lacks validity.
- Controlling confounders (1-2 marks)
- Matching or measuring confounders (age, sex, diet, occupation, socioeconomic status) ensures the outcome is attributed to the exposure under study rather than a hidden variable - central to validity.
- Measured, graded exposure (1 mark)
- Recording how much exposure each subject had allows a dose-response relationship to be tested; a clear gradient strengthens the case for causation.
- Adequate duration and follow-up (1 mark)
- The study must run long enough for the disease to develop, otherwise true effects are missed - a validity issue (e.g. a 12-month study of a slow cancer).
Full marks need at least four features, each tied to WHY it improves reliability or validity, not merely listed.
exam7 marksEvaluate the use of epidemiology in establishing the link between cigarette smoking and lung cancer, with reference to the study designs used and the criteria for inferring causation.Show worked solution →
"Evaluate" demands a judgement weighing strengths and limitations, supported by the methodology and the causal reasoning.
- The evidence and designs (2-3 marks)
- Doll and Hill's 1950 case-control study compared lung-cancer patients with matched controls and found smokers far over-represented, with a dose-response gradient (more cigarettes, higher risk). The British Doctors Study then followed ~40 000 doctors prospectively (cohort design) for decades, confirming lung-cancer mortality was many times higher in heavy smokers and that quitting reduced risk. Using both designs together is a strength: the case-control study generated the hypothesis efficiently; the cohort study established temporal sequence and measured incidence.
- Inferring causation (2 marks)
- Because these are observational studies they cannot, alone, prove causation - it would be unethical to randomly assign people to smoke. The Bradford Hill criteria (strength, consistency, temporality, biological gradient, plausibility, coherence, experiment, analogy, specificity) were applied: smoking and lung cancer satisfied them, with a strong consistent dose-response across many studies and a plausible carcinogenic mechanism.
- Limitations weighed (1-2 marks)
- Observational data carry confounding and bias risks (e.g. recall bias in the case-control interviews); no single study is decisive.
- Judgement (1 mark)
- Despite the inability of any one observational study to prove causation, the convergence of case-control and cohort evidence, the dose-response gradient, and satisfaction of the Bradford Hill criteria make the causal link between smoking and lung cancer firmly established - a landmark demonstration of epidemiology's power. An answer lacking an explicit judgement, or omitting the causal criteria, caps below full marks.
