Login

Please fill in your details to login.





009. bias in data: who is missing from the dataset (ks5)

Critically analyse datasets for historical biases that lead to flawed and discriminatory AI algorithms.

We often trust data to give us the objective truth. But what if the data itself is prejudiced? If a machine learning algorithm is trained on incomplete information, its decisions will be inherently flawed. In this advanced module, we will critically interrogate datasets. You will learn to ask "who is missing?", analysing how historical and structural biases in data collection can lead to discriminatory technology.

Bias in the Machine: The Ghost in the Data


We tend to think of computers as impartial and logical, making decisions based on pure data, free from human emotion and prejudice. However, the rise of Artificial Intelligence has revealed a dangerous flaw in this assumption. AI systems learn to make decisions by analysing vast quantities of training data - and if that data is biased, the AI will inevitably learn, automate, and even amplify those biases. This is known as algorithmic bias.

Where Does Bias Come From?


Bias isn't programmed into a machine deliberately; it creeps in through the data we feed it. There are several common types:

Historical Bias: This occurs when the data reflects past societal prejudices. For example, if an AI is trained on historical hiring data from a time when a certain profession was male-dominated, it might learn to unfairly favour male candidates for that role, even if gender is not an explicit factor. It learns a correlation that is a product of history, not of skill.

Selection Bias (or Sampling Bias): This happens when the data used to train the model is not representative of the real world. Imagine creating a facial recognition system trained predominantly on images of one ethnic group. The system would likely have a much higher error rate when trying to identify people from other, underrepresented groups. The sample data does not reflect the diversity of the population.

Measurement Bias: This arises when the data is collected or measured in a flawed way. For example, if a company uses 'arrests' as a proxy for 'crime rate' to deploy police patrols, it might create a feedback loop. Areas with a higher police presence will naturally have more arrests, leading the algorithm to send even more police to that area, regardless of the actual underlying crime rate. The metric itself is skewed.

The Real-World Consequences


Algorithmic bias is not a theoretical problem. It has significant real-world consequences. Biased AI has been found in systems that decide on loan applications, score job applicants, and even assist in medical diagnoses. When these systems make flawed judgements, they can reinforce systemic inequality and deny people opportunities based on skewed historical data rather than their own merit.

As future technologists, it is our responsibility not just to build powerful systems, but to build fair ones. The first step is learning to critically interrogate every dataset and ask the crucial question: whose story is this data not telling?

Plugged Task: The Algorithmic Audit


image
The Scenario
You are a Junior Data Analyst at a financial technology company called 'FinTechForward'. The company has developed an AI model to automate loan approvals. However, early testing suggests the model may be showing unfair bias against applicants from certain areas. Your manager has given you a small sample of the training data and asked you to conduct an audit and write a brief report identifying potential sources of bias.

The Persona
As The Analyst, your job is to be sceptical and meticulous. You must look beyond the surface of the numbers, question how the data was collected, and clearly communicate the risks of using a potentially flawed dataset.

1
Get Organised

Open a new word processing document and create a simple report structure with the following headings: Introduction, Analysis of Dataset, Identified Biases and Risks, and Recommendations.

2
Research a Real-World Precedent

To understand the severity of the issue, you need to know how this has happened before. Research a real-world case of algorithmic bias. A famous example is Amazon's experimental AI recruiting tool. Use the link below to get started.

In the Introduction of your report, briefly summarise the real-world case study you researched. Explain what the AI was for and why it was found to be biased.

3
Analyse the Sample Data

Your manager has sent you a link to a spreadsheet containing a sample of 100 fictional loan applicants. Your task is to look for patterns that could lead to bias.

Open the spreadsheet. Spend some time sorting and filtering the data. Do you notice any patterns? Look at things like:
The number of applicants from each postcode. Is it evenly distributed? This could be a sign of selection bias.
The average income for different genders or the 'Job Title' provided. Does it reflect historical stereotypes? This could be a sign of historical bias.
How is 'Credit Score' represented? Is it a simple number, or is it a category like 'Good'/'Bad'? Could this measurement be flawed? This might be measurement bias.

image
You are a data science tutor. Explain the difference between historical bias, selection bias, and measurement bias. Use simple analogies. Keep your answer under 150 words. NO intro, NO outro, NO deviation from the topic, NO follow-up questions


4
Document your Findings

In the Analysis of Dataset and Identified Biases and Risks sections of your report, document the patterns you found.

For each potential bias, name the type (e.g., "Potential Selection Bias").
Describe the evidence you found in the data (e.g., "Only 5% of applicants are from the 'M41' postcode, suggesting this area is underrepresented.").
Explain the risk this poses (e.g., "The AI may learn to unfairly penalise applicants from 'M41' simply because it has not seen enough successful examples from that area.").

5
Make Recommendations

In the final section of your report, propose two concrete actions the company could take to improve the dataset before training the final AI model. Think about how you could address the biases you have found.

Outcome
Your final report is professionally formatted with clear headings.
You have correctly summarised a real-world case of algorithmic bias.
You have identified at least two different types of potential bias in the sample dataset.
For each bias, you have explained the evidence and the potential negative consequence.
Your report includes two sensible recommendations for the company to improve its data.

Unplugged Task: Data Equity Infographic


On a blank piece of A4 paper, design a clear and visually engaging infographic for the managers at 'FinTechForward' who are not data experts.

Your infographic must:
Have a powerful, clear title (e.g., "Building Fairer AI").
Include a simple, non-technical definition of 'Algorithmic Bias'.
Use simple doodles or diagrams to explain one type of bias (e.g., for selection bias, you could draw a basket of mostly red apples with only one green one).
List three "golden rules" the company should follow to ensure their data collection is fair and representative in the future.
Last modified: June 24th, 2026
The Computing Café works best in landscape mode.
Rotate your device.
Dismiss Warning