YOURPACEAI WORKSHOP · Beginner+
Data & machine learning
Build a coding foundation, explore datasets, and learn how models make predictions.
75–120 minutes · Five modules, a portfolio workbook and automatic assessment.
A spreadsheet is enough for the workshop. Optional Python/Pandas extension requires basic coding. The supplied CSV is fictional.
Open interactive workshopRead all lessons freely here. The interactive workshop lets you save your workbook, take the automatic assessment and earn a non-accredited certificate.
By the end, you can
- Clean a small dataset with a documented rule.
- Calculate a reproducible result and choose a chart.
- Explain limitations and avoid leakage in a prediction design.
Practice materials
Download practice packRead the fictional source material
FICTIONAL SALES CSV
id,day,category,sales_gbp
1,Mon,Books,40
2,Tue,Books,60
2,Tue,Books,60
3,Wed,Books,
4,Thu,Books,80
5,Fri,Books,100
Question: What is the average sales value across unique records with a recorded amount?
Unit: GBP. Blank means missing, not zero. Repeated id 2 is an exact duplicate. No population sampling claim is made.
From a dataset to a useful answer
MODULE 1 OF 5
Start with a question
Choose a question before choosing a model. A spreadsheet and a chart may answer it. Use a public dataset with clear permission and definitions; record what each column means.
Worked exampleInput
Question: average recorded sales for unique records; CSV has five unique IDs.
Reviewed result
Use id to identify the exact duplicate, and sales_gbp for the amount. Keep the raw CSV unchanged.
Define the question and units first so the cleaning rule is linked to the analysis.
Quick check: Should a missing amount automatically mean zero sales?
No. Missing and zero are different states.
Your activityPick a small public dataset. Write one question and identify the columns you need.
MODULE 2 OF 5
Check the data
Look for missing values, duplicates, inconsistent units and unrepresentative samples. AI can suggest cleaning code, but inspect and test it. Keep the original dataset and document changes.
Worked exampleInput
Six raw rows; id 2 appears twice; id 3 has a blank amount.
Reviewed result
Drop the repeated exact record. Exclude id 3 only from this mean calculation and document it. Four recorded values remain: 40, 60, 80, 100.
This is a stated rule for this exercise. Real duplicates or missing values require context rather than automatic deletion.
Quick check: What is the mean under this rule?
(40 + 60 + 80 + 100) / 4 = £70.
Your activityCount missing values and duplicates. Explain one cleaning decision and how it changes your answer.
MODULE 3 OF 5
Avoid misleading results
A chart shows a relationship, not necessarily a cause. For predictive models, keep test data separate from training and avoid using information unavailable at prediction time. Compare against a simple baseline.
Worked exampleInput
A model predicts purchases using a field recorded after purchase.
Reviewed result
Remove post-purchase information from prediction inputs; keep a test set separate and compare with a simple baseline.
Information unavailable at prediction time creates leakage and misleading evaluation.
Quick check: Does a rising chart prove what caused the increase?
No. These values alone do not establish causation.
Your activityMake one chart and write a finding plus two limitations. If predicting, define the test split before training.
MODULE 4 OF 5
Make your result reproducible
In a spreadsheet, keep Raw and Clean tabs. Document duplicate removal and the missing-value rule, then calculate count, sum and mean. In Python, inspect duplicated IDs and missing values before transforming the data.
Worked exampleReproduction checklist: 6 raw rows → 5 unique IDs → 4 recorded amounts → total £280 → mean £70.
Your activityCreate a bar chart of the four recorded day amounts. Label GBP and note the excluded missing record.
MODULE 5 OF 5
Write a finding with limits
Separate your arithmetic from broader claims. This is a tiny fictional set, not evidence about a business or population. A useful report includes what you did and what you cannot conclude.
Worked exampleFinding: the four unique recorded amounts average £70. Limitations: one value is missing and this invented week cannot establish a trend or cause.
Your activityWrite a finding, the cleaning log and two limitations. Add a separate plan for how you would evaluate a predictive model.
A reproducible data analysis
Clean the fictional CSV, calculate the mean, design a chart and explain what the data cannot show.
Workbook sections
- Question, column meanings and cleaning rules
- Clean values, formula or code, count and mean
- Chart description and evidence-backed finding
- Missing-data limits and a leakage-free evaluation plan
Review rubric
- The raw data and stated cleaning rules are retained.
- The exact duplicate is handled and the missing amount is not silently replaced with zero.
- The mean is reproducible and chart units are clear.
- Findings do not claim causation or population generality.
Automatic assessment and certificate
Five knowledge questions and three applied scenario checks are marked immediately. Pass with at least 4/5 knowledge answers and all 3/3 applied checks correct. Feedback and retries are available. Certificates also require five completed activities and a four-section workbook. The portfolio is recorded, not independently graded; the assessment is open-book, unproctored and non-accredited.
Take the workshop assessmentContinue with external study
Go further with a portfolio project
Use a public dataset to answer one question. Make a chart, explain your findings, and record the limitations of the data.
External course access and fees are set by their providers.