Instructor Notes
General Teaching Approach
This course follows the Carpentries live-coding pedagogy: the instructor types code live while participants follow along. Avoid slides for code; always demonstrate in RStudio.
Key principles: - Start from what they know: Every R operation is introduced alongside its SPSS equivalent. Use SPSS terminology first, then introduce the R term. - Wow first, skills second: Episode 1 is pure motivation. Show impressive things before asking anyone to type. - Sticky notes: Use colored sticky notes (or digital equivalents) for real-time feedback. Green = I’m following. Red = I need help. - Helpers: Aim for 1 helper per 5-8 participants to assist with individual issues without stopping the class.
Session Structure
Session 1 (5-6 hours with breaks)
| Episode | Time | Notes |
|---|---|---|
| 01 - The Case for Switching | 35 min | Instructor demo only, no participant coding. Demo is the UA SIDS
reference-list pull from island-research-reference-data
(see Episode 1 instructor block). Second half of the demo loads the
squad data and counts where the Curacao men’s squad plays; that half
needs no network and is the one to protect. Also shows the squad-report
capstone HTML as the Friday target, rendered at
episodes/files/blue-wave-squad-report.html and sourced at
episodes/files/blue-wave-squad-report-template.Rmd. |
| Break | 15 min | |
| 02 - Your First R Session | 65 min | First hands-on. Go slow. Many will struggle with typos. The “Before
you import” subsection is a deliberate whole-room synchronized moment:
project the download links on the screen, wait for green stickies in the
Files pane before typing read_csv(). |
| Break | 15 min | |
| 03 - Data Manipulation | 60 min | The pipe operator is the key “aha” moment |
| Break | 15 min | |
| 04 - Visualization | 35 min | End on a high, everyone leaves with a beautiful chart. Two datasets in this episode: squad data for the categorical charts, FIFA rankings for the continuous ones. Do not suppress the missing-data warning on the first scatterplot; it is taught. |
| Wrap-up + homework brief | 10 min | Project the homework page on the screen. Walk through the four-step assignment out loud. Tell participants the page URL is bookmarked under “For Learners → Homework brief” on the course site so they can open it on any device overnight. Emphasise: 30–60 minutes is enough, do not attempt R Markdown yet (that is Day 2). Bring the script to the Day 2 recap; there is no open lab in the four-hour format. |
Session 2 (5-6 hours with breaks)
| Episode | Time | Notes |
|---|---|---|
| Review, homework and troubleshooting | 20 min | Address questions from between-session practice. Close it with one question and do not answer it: “The national team named a new squad on 11 September. How long would it take you to redo Wednesday’s analysis?” Episode 6 answers it with one changed line. |
| 05 - Statistical Analysis | 95 min, split either side of the coffee break | Core for survey researchers. The normality-testing section (histogram, Q-Q plot, Shapiro-Wilk, robustness note) maps directly onto the SPSS Explore output most participants will recognise. Take your time. |
| Break | 15 min | |
| 06 - Reproducible Reporting | 50 min | R Markdown is often the biggest “wow” for SPSS users. Ends with the
squad-report capstone that Episode 1’s opening teased. Participants pull
blue-wave-squad-report-template.Rmd and
blue-wave-report.css from the GitHub raw URL via the
download.file() block in the episode; walk through the
template’s structure live once both files are in their working
directory. Finish by changing params$team from
CUW-M to ARU-W and re-knitting, so they see
one file produce a different report. |
| Break | 15 min | |
| 07 - Where to Go from Here | 40 min | End with practical next steps. The new UA datasets subsection
(CAS_election_data and island-research-reference-data) is a chance to
live-demo read.csv() straight from a raw GitHub URL, and
most SPSS users have never seen data load over HTTPS without a manual
download. |
Per-episode scene transitions
The course carries a series of cartoons, not eight
unrelated images. One Curacaoan analyst runs through all of them,
starting as a fan with a wall map on the index page and finishing at the
tiller of a boat heading out of the harbour. The renders and full scene
briefs are in scene-briefs.md.
This is a deliberate change from the Aruba course, where the images were decoration and instructors were told to ignore them. Here you may point at them. Use the callback in Episode 6 and again in Episode 7, where the payoff sits.
Each episode opens with its scene image and a one-line quip caption. The captions carry most of it. The transition lines below are optional single beats for instructors who want one on arriving at a new episode.
| Ep | Caption on page | Optional transition line |
|---|---|---|
| Index | Twenty-six players. Ten countries. One spreadsheet. | (Shown on the landing page and on the opening slide. No spoken line; let people read it while the room settles.) |
| 1 | One gate charges you every season. The other one only asks you to learn the way in. | (Episode 1 opens with the workshop’s full opening sequence; no separate transition needed.) |
| 2 | The dominoes can wait. The console is blinking. | “Laptop open, console blinking, iguana unimpressed. Time to type something.” |
| 3 | The stew takes twenty minutes. The chopping takes an hour. | “Three jars on the counter today: filter, select, mutate. Everything else in dplyr is a variation on those three.” |
| 4 | SPSS hands you a chart. ggplot2 hands you a grammar. | “Those are the Handelskade houses and they are also a bar chart. By the end of this episode you will be writing the sentence that draws them.” |
| 5 | Sometimes the review says nothing happened. You report that too. | “The tests you know from SPSS are all here. What is new is that one of today’s two comes back null, and we are going to report it anyway.” |
| 6 | The squad changed overnight. Again. Good thing the report rebuilds itself. | “This is where R Markdown earns the price of admission. New squad in, finished document out, one button.” |
| 7 | You have the basics. The map runs well past the harbour mouth. | “She started this course pinning photos to a wall map. She is leaving the harbour with a chart. That is roughly where you are too.” |
Pick one beat, deliver it, move into the page’s first heading. Do not stack a second sentence on top.
What the four-hour format costs
The course runs 09:00 to 13:00, which is 480 minutes of room time for 380 minutes of episodes. The overhead is welcome, recap, two breaks a day and the two wrap-ups. Three things were cut to make it fit, and instructors should know which:
The open lab on Day 2 is gone. Homework review moved into the 20-minute Day 2 recap, so protect that slot; it is now the only place a stuck participant gets unstuck.
Exercise time in Episodes 2, 3 and 4 took most of the reduction, roughly 25 minutes each. Episode 5 was left slightly long on purpose because it is the hardest and it carries the null result.
The Day 1 questions-and-consolidation block is gone. Fold consolidation into the end of each episode rather than saving it up.
The AI and package-trust material, added 22 September 2026
Two additions, both deliberate and both cuttable under time pressure.
Episode 2 gains a callout after the
install.packages() versus library() box: what
CRAN’s review actually guarantees, the July 2026 Hugging Face intrusion
as the contrast (a malicious dataset abusing code execution paths in the
upload processing pipeline, so reading data is running code), and what
to do where an employer blocks installations, which is an internal
mirror plus renv pinning. Budget four minutes spoken. Two
Central Bank staff are in the room and their institution treats package
installation as a security exposure, so deliver this as a fair position
with a professional answer rather than as an obstacle. If Episode 2 is
running behind, say the last paragraph only and leave the rest to be
read.
Episode 7 gains “Letting AI write the R” before the learning resources: that generating R is now the faster route for routine work, that the job becomes judging the output, the four failure modes that read well, and the rule about never pasting supervisory or personal data into a public model. Episode 7 is self-guided, so this costs nothing from the timetable. It is worth naming out loud in the wrap-up even if nobody reads the section.
The index says both are covered, which matters for institutions deciding whether to send staff.
What went wrong on Day 1, Curacao, 23 September 2026
Recorded the same evening, from delivery. Every item here cost time in the room.
Getting the data into R was the whole problem.
Participants had the files in a folder and no clear route from there to
a loaded data frame. The instructions said to create a project but never
said, in one line, how to point R at the folder afterwards. Rendell set
the working directory to the data folder rather than to its
parent, which breaks every "data/..." path in the course by
one level, and the paths had to be edited live. Episode 2 now has a step
4 that prints getwd() and list.files("data")
before anything else, and setup.md says to pick the parent folder in as
many words.
Nobody knew whether to upload through the menu or load by command. Walk the room through one route and name it as the route. Do not offer both.
The downloaded files collided. All three arrived under the same name and overwrote each other, so everyone had to rename before anything would load. setup.md now says to check names and extensions in the Downloads folder first.
The Excel import silently returned half the data.
read_excel() takes the first sheet, and this workbook
splits 48 rows of Curacao and 46 of Aruba. Episode 2 now counts the
rows, names the trap, and stacks the sheets with
bind_rows(). Teach it as the lesson it is: an Excel import
that returns half your cases looks exactly like one that worked.
The fix that removes all of this at once is reading
from the web. Episode 5 now opens with a base URL and three
read_csv(paste0(base, ...)) lines, so Day 2 starts with
data in memory whatever state anyone’s folders are in. Put that block on
the screen first and let people catch up while you talk.
There was no second pair of hands. One instructor cannot debug twenty installations and keep a timetable. Where no helper is funded, recruit two or three participants who got through setup quickly and ask them openly to help their neighbours. It costs nothing and it works.
Common Issues
- Installation problems: The registration email asks participants to reply with any install error before the day, so most should arrive fixed. Have a USB drive with R and RStudio installers as backup.
- Typos: SPSS users are not used to typing commands. Expect many syntax errors. Normalize this: “error messages are how R talks to you.”
- Parentheses and quotes: The most common beginner errors. Show how RStudio auto-completes these.
-
Loading packages: Participants will forget
library(). Remind them at the start of each episode.
Local Data Notes
The course uses Dutch Caribbean datasets to keep examples relevant.
Everything learners compute on is generated by
scripts/00_build_teaching_data.R from verified sources and
committed to episodes/data/. Nothing is scraped live in the
room.
-
blue_wave_squad.csv, player-level squad lists for
the four ABC island national teams, 94 players. Primary teaching
dataset, Episodes 2 to 4 and 6. Also shipped as
.xlsx(two sheets) and.savfor the import demonstrations. Curaçao’s men are the World Cup squad. Scraped from pinned Wikipedia revisions byscripts/00_build_teaching_data.R; variable definitions and limitations are inepisodes/data/blue_wave_squad_codebook.md. - blue_wave_squad_2026-09.csv, the September 2026 call-up, same columns, 91 players. Episode 6 only, where the capstone report is re-knit on it and the Episode 3 region code is shown misfiling Bosnia and Gibraltar.
- fifa_rankings.csv, FIFA rank and points against population and diaspora for 211 national associations. Continuous variables for Episodes 4 and 5.
- diaspora_change.csv, diaspora stock in 1990, 2010, and 2024. Supplies the paired t-test in Episode 5.
-
island-research-reference-data, a country reference
list with SIDS, SNIJ, and World Bank classifications, pulled live from
GitHub in Episode 1 and Episode 7. Offline fallback committed at
episodes/data/countries_backup.csv. - CAS_election_data, Aruba, Curacao, Sint Maarten election results 1985-2025. Used in Episode 7.
- World Bank indicators via
WDI(Episode 7) and CBS Netherlands viacbsodataR(Episode 7).
Regenerate the derived files with
Rscript scripts/00_build_teaching_data.R from the
repository root. Test all live downloads before the course; URLs and
APIs change.
The gaps in the squad data are the point
Two of 94 players cannot be placed at a club, and the squads are only as current as Wikipedia. Episodes 3, 5, and 6 each stop to name what is being excluded and what it costs. Do not tidy this away or apologise for it. Most participants have never been shown what to do with a gap other than delete it, and this is the most transferable thing in the two days.
Two is a small number, and that is deliberate rather than unfortunate. The habit of reporting exclusions is easiest to build when the exclusion changes nothing.
The island comparison does not work, and Episode 5 uses that
Testing whether Curaçao and Aruba differ on players-based-abroad returns a p-value around 0.64. Testing men against women returns 0.006. Episode 5 runs both in that order on purpose: the first thing you try fails, the second works, and the write-up has to admit both. Do not skip to the one that works.
Note on the elections example
The Episode 7 election-data example summarises fragmentation across all parties in each Curacao election rather than singling out any one party. In a room that may contain civil servants and ministry staff, filtering on a single party name reads as partisan even when it is not meant to. If extending the example live, default to all-parties views. If a participant asks why, this is a small editorial choice that protects the course and the network’s neutrality, and it is worth naming briefly.
Note on the Papiamentu example
The Episode 7 text-analysis sample is written in Curacao orthography, not the Aruban etymological standard used in the master course. The accompanying callout uses the difference to make a point about stopword lists being analytical choices. If someone in the room raises the orthography question, that is a good outcome, not a derailment.
Train-the-Trainer
This course is designed for replication. If you are adapting it for another island or institution: 1. Replace datasets with locally relevant equivalents 2. Adjust the SPSS operations covered based on your pre-course survey results 3. Keep the “wow first, skills second” structure 4. All materials are CC-BY 4.0, so please attribute the DCDC Network
The Case for Switching
Live demo script
This is the complete script to run live. Practice this before
the workshop. Make sure tidyverse is installed.
Everything the demo needs comes with it.
The demo has two halves. The first shows R reaching out to a live research dataset the network owns. The second lands the local hook on the squad data the rest of the course uses. Run the first if the Wi-Fi is cooperating, and the second regardless, because the second is the one that makes the room lean forward.
Part A, step 1: Frame the source before you type
Before the first keystroke, name what the room is about to see. The
CSV about to load is in a GitHub repository maintained at the University
of Aruba, island-research-reference-data, part of the DCDC
Network’s shared infrastructure for island research. The network owns
and curates it, and keeps it up. That framing matters, because the point
here goes past R being able to read a URL: the data layer underneath
belongs to us.
Part A, step 2: Pull the SIDS reference list
Open a new R script in RStudio and type or paste the following. Run it line by line so participants can watch each step.
R
# Load packages (install tidyverse once, before the workshop)
# install.packages("tidyverse")
library(tidyverse)
# Pull the UA island-research reference list straight from GitHub
countries <- read_csv(
"https://raw.githubusercontent.com/University-of-Aruba/island-research-reference-data/main/countries/countries_reference_xlsform.csv"
)
# Quick look at what we got
head(countries)
Pause. Point out: “No browser. No download dialog. No save-as. The file is now a live object in my session, with over a dozen columns per country.”
Part A, step 3: Filter to SIDS and chart by region
R
countries |>
filter(is_sids == 1) |>
count(wb_region) |>
ggplot(aes(x = reorder(wb_region, n), y = n)) +
geom_col(fill = "#2e8894") +
coord_flip() +
labs(
title = "Small island developing states by World Bank region",
x = NULL,
y = "Number of SIDS"
) +
theme_minimal(base_size = 14)
Pause again. Key talking points:
- “This chart is ready for a report as it stands. Title, axis labels, colour, proportions, all set in code.”
- “If the UA team adds a country to the reference list tomorrow, I re-run this script and the chart updates. No re-click, no re-export.”
- “Every editorial choice, what counts as a SIDS and which region goes where, is traceable, because the definitions sit in the source CSV you just pulled.”
Part B: the local hook
This is the moment the course earns the room’s attention. Load the squad data and answer the question the island already argues about.
R
squad <- read_csv("data/blue_wave_squad.csv")
squad |>
filter(team_code == "CUW-M") |>
count(club_country, sort = TRUE)
Read the result out loud. Then draw it:
R
squad |>
mutate(
island = if_else(str_starts(team_code, "CUW"), "Curaçao", "Aruba"),
home_code = if_else(str_starts(team_code, "CUW"), "CUW", "ABW"),
region = case_when(
club_country == "X" ~ "Unknown",
club_country == home_code ~ "Home island",
club_country == "NLD" ~ "Netherlands",
club_country == "USA" ~ "North America",
club_country %in% c("GBR", "GRC", "TUR", "DEU", "BEL", "CHE", "XKX") ~ "Rest of Europe",
.default = "Rest of world"
)
) |>
count(island, region) |>
ggplot(aes(x = island, y = n, fill = region)) +
geom_col(position = "fill") +
scale_y_continuous(labels = scales::percent) +
labs(
title = "Where the players actually play",
subtitle = "All four ABC island national squads, 2026",
x = NULL, y = NULL, fill = NULL
) +
theme_minimal(base_size = 14)
Talking point: “Two islands, 80 kilometres apart, with populations and colonial histories that track each other closely. Both send most of their players abroad. But look where. Curaçao’s are scattered across ten countries. Aruba’s are overwhelmingly in one. Nobody in this room needed a statistician to care about that question. What you needed was a way to ask it, and you just watched it take twelve lines.”
Do not resolve the question. Let the room argue. Say: “By Wednesday you will be able to run this yourself, and change it to ask what you actually want to know.”
Step 4: Show the contrast with SPSS
Ask the audience: “How would you have done this in SPSS?”
Walk through it slowly. Make it sting a little. This is the moment the cost of the current workflow lands.
- Go looking for the squad lists. Wikipedia? FotMob? The federation’s Facebook page? Pick one and hope it is current.
- Copy 94 player names, clubs, and positions into Excel by hand. An hour if you are quick and do not lose concentration.
- Discover that the club’s country is only shown as a little flag you cannot copy, that two players have no club listed at all, and that Jong Holland is a Curaçao club despite the name. Clean by hand.
- Import into SPSS. Recode club country into a grouping variable, because the raw text is unusable as it stands.
- Analyze > Descriptive Statistics > Frequencies. Copy the output table.
- Graphs > Chart Builder, drag variables, format the chart, copy, paste into Word.
Then say: “That is half a morning, on a good day. In R it was twelve lines and ten seconds, and the next person who asks gets the same answer from the same file.”
Backup plan
If the Wi-Fi is unreliable, the reference list is saved locally at
episodes/data/countries_backup.csv. Swap the
read_csv() call for the local path:
R
countries <- read_csv("data/countries_backup.csv")
Then proceed with the
filter() |> count() |> ggplot() pipeline as normal.
Part B is local already and needs no network at all, which is why it is
the half you never skip.
Your First R Session
Pacing notes
- Spend time on the RStudio pane orientation. Have participants identify each pane on their own screen before moving on.
- The
<-assignment operator trips people up. Give them a few minutes to practice creating objects with different names and values. - When loading tidyverse, the startup messages can be alarming to beginners. Reassure them that the “Attaching packages” and “Conflicts” messages are normal and expected.
- Do the “Before you import” subsection as a
whole-room moment, not as reading. Project the download page, walk
everyone through the browser download, the New Project dialog, and
creating the
datasubfolder. Wait for green sticky notes in the Files pane before typingread_csv(). This is the most common point of failure in the course and it is worth five deliberate minutes up front to avoid twenty scattered minutes of troubleshooting later. - The
.savimport is worth a beat of theatre. Most of the room has years of.savfiles they assume are trapped. Let them see one open in three lines. - Encoding warning is not pedantry. If a participant opens the CSV in Excel and saves it, the ç in Curaçao usually breaks and they will hit an error later that looks unrelated. Say it once, firmly.
- If a participant cannot download (blocked network, locked laptop), have a USB stick or shared-drive copy of the three files ready as fallback.
- The two-way table at the end is a deliberate cliffhanger. Do not explain it. Let someone in the room say it out loud.
Data Manipulation
Teaching tips
- Write the pipe
|>on the whiteboard and say “and then” out loud every time you use it. This mental model sticks. - Build pipelines live, one step at a time. Run after each added line so participants can see how the output changes.
- The most common beginner mistake is putting
|>at the start of a line instead of at the end of the previous line. Emphasise that the pipe goes at the end of the line, so R knows the expression continues. - Compare nested function calls to piped code side by side. The readability advantage sells itself.
- The
mean()of a logical is the single highest-leverage idea in this episode. SPSS users are trained to build dummy variables by hand. Show them the shortcut, then show them the SPSS way they would otherwise have used, and let the contrast do the work. - The
filter(club_country %in% c("CUW", "ABW"))result usually produces an audible reaction in a Curaçao room. Let it. Then say “we are not explaining that today, we are learning how to ask it.” - If someone asks why two players have
club_country == "X", the honest answer is that the source is a Wikipedia squad table, those two rows list no club, and the build script recorded that rather than guessing. Point them toblue_wave_squad_codebook.mdin the data folder. This is a good moment for a word about documenting your own uncertainty. - Someone may well ask how current the squads are. They are as current as Wikipedia, which is to say: maintained by volunteers, lagging real call-ups by weeks or months, and not necessarily consistent across the four pages. Say so. It is a better answer than pretending, and it sets up Episode 5.
Visualization with ggplot2
Instructor note
The faceting example is a good place to pause and let learners experiment. Encourage them to try:
-
facet_wrap(~ island, ncol = 1)to control the layout - swapping the facet formula to
facet_grid(island ~ gender)and asking which comparison each version makes easy - removing
coord_flip()to see why the labels needed it
The position = "fill" versus default stacking contrast
is the most transferable idea in the episode. Most people in the room
have at some point presented a counts chart where the group sizes
differed, and drawn a conclusion from bar height that the proportions
did not support. Say that out loud.
Do not skip past the missing-data warning on the first scatterplot. Learners who have been taught to make warnings go away need to hear, from you, that this particular one is information.
Statistical Analysis in R
Instructor note
This is a good time to reinforce the reproducibility advantage. In SPSS, if a reviewer asks you to re-run an analysis on a different subset, you click through the dialogs again. In R you change one line and re-run the script.
Emphasise that broom::tidy() produces a data frame, so
learners can use every dplyr verb from Episode 3 on their statistical
results: filter to significant terms, arrange by p-value, join to
labels.
Three moments are worth slowing down for, and none of them is about syntax.
The non-significant diaspora coefficient is the most important. Most of the room has been taught, implicitly, that a good table has stars in it. Say plainly that you are leaving a dead predictor in the model on purpose, and why. If anyone asks whether they should drop it, that is the discussion you want.
The pair of chi-square tests is the second. Run the island one, let the room see the p-value, and let the disappointment sit for a moment before you run the gender one. The sequence is the lesson: the first thing you tried did not work, you tried a second thing, and the write-up has to admit both. Almost everyone in the room has at some point reported only the second.
The third is the filter(club_country != "X") line. Ask
what it did before you tell them. Someone will spot that it removed two
players. Two is small enough that nobody would object, which is exactly
why it is a good example: the habit has to be built when the stakes are
low.
Reproducible Reporting
Common knitting problems
The most common issue is that knitting fails because the
.Rmd file does not load packages or data that earlier
chunks depend on. Remind participants that knitting starts from a
blank environment. Every package and dataset must be
loaded within the .Rmd file itself, even if it is already
loaded in their current R session.
Another common issue is file paths. If participants write
read_csv("data/blue_wave_squad.csv"), the working directory
during knitting is the folder where the .Rmd file is saved.
Make sure the data file is in the right relative location.
Third, and specific to this dataset: encoding. If a participant’s
Windows locale mangles the ç in Curaçao during knitting, have them save
the .Rmd as UTF-8 via File > Save with
Encoding. It is worth mentioning pre-emptively rather than
debugging it live in four separate laptops.
Timing the re-run
This section costs about eight minutes, most of it the download and
the knit. If the room is behind, run it from the front instead:
pre-rendered copies of both reports are in episodes/files/,
and the region callout needs nothing but the chunk above. Keep the match
itself out of it. The finding holds whatever the score on Thursday
night.
Where to Go from Here
Before delivery: two things to replace
The two survey links above are the Aruba pilot’s forms. Either create Curaçao cohort forms and swap the URLs, or accept mixed responses across editions and add a cohort question to the existing forms. Decide before promotion goes out, not on the morning.
The Papiamentu sample text in the text-analysis section is written in Curaçao orthography, which differs from the Aruban standard used in the master course. That was deliberate for this room. If you would rather teach it as a comparison, the Aruban version of the same passage is in the master course repository and running both through the same stopword list makes the point in the callout better than the callout does.