Instructor Notes

General Teaching Approach


This course follows the Carpentries live-coding pedagogy: the instructor types code live while participants follow along. Avoid slides for code; always demonstrate in RStudio.

Key principles: - Start from what they know: Every R operation is introduced alongside its SPSS equivalent. Use SPSS terminology first, then introduce the R term. - Wow first, skills second: Episode 1 is pure motivation. Show impressive things before asking anyone to type. - Sticky notes: Use colored sticky notes (or digital equivalents) for real-time feedback. Green = I’m following. Red = I need help. - Helpers: Aim for 1 helper per 5-8 participants to assist with individual issues without stopping the class.

Session Structure


Session 1 (5-6 hours with breaks)

Episode Time Notes
01 - The Case for Switching 35 min Instructor demo only, no participant coding. Demo is the UA SIDS reference-list pull from island-research-reference-data (see Episode 1 instructor block). Second half of the demo loads the squad data and counts where the Curacao men’s squad plays; that half needs no network and is the one to protect. Also shows the squad-report capstone HTML as the Friday target, rendered at episodes/files/blue-wave-squad-report.html and sourced at episodes/files/blue-wave-squad-report-template.Rmd.
Break 15 min
02 - Your First R Session 65 min First hands-on. Go slow. Many will struggle with typos. The “Before you import” subsection is a deliberate whole-room synchronized moment: project the download links on the screen, wait for green stickies in the Files pane before typing read_csv().
Break 15 min
03 - Data Manipulation 60 min The pipe operator is the key “aha” moment
Break 15 min
04 - Visualization 35 min End on a high, everyone leaves with a beautiful chart. Two datasets in this episode: squad data for the categorical charts, FIFA rankings for the continuous ones. Do not suppress the missing-data warning on the first scatterplot; it is taught.
Wrap-up + homework brief 10 min Project the homework page on the screen. Walk through the four-step assignment out loud. Tell participants the page URL is bookmarked under “For Learners → Homework brief” on the course site so they can open it on any device overnight. Emphasise: 30–60 minutes is enough, do not attempt R Markdown yet (that is Day 2). Bring the script to the Day 2 recap; there is no open lab in the four-hour format.

Session 2 (5-6 hours with breaks)

Episode Time Notes
Review, homework and troubleshooting 20 min Address questions from between-session practice. Close it with one question and do not answer it: “The national team named a new squad on 11 September. How long would it take you to redo Wednesday’s analysis?” Episode 6 answers it with one changed line.
05 - Statistical Analysis 95 min, split either side of the coffee break Core for survey researchers. The normality-testing section (histogram, Q-Q plot, Shapiro-Wilk, robustness note) maps directly onto the SPSS Explore output most participants will recognise. Take your time.
Break 15 min
06 - Reproducible Reporting 50 min R Markdown is often the biggest “wow” for SPSS users. Ends with the squad-report capstone that Episode 1’s opening teased. Participants pull blue-wave-squad-report-template.Rmd and blue-wave-report.css from the GitHub raw URL via the download.file() block in the episode; walk through the template’s structure live once both files are in their working directory. Finish by changing params$team from CUW-M to ARU-W and re-knitting, so they see one file produce a different report.
Break 15 min
07 - Where to Go from Here 40 min End with practical next steps. The new UA datasets subsection (CAS_election_data and island-research-reference-data) is a chance to live-demo read.csv() straight from a raw GitHub URL, and most SPSS users have never seen data load over HTTPS without a manual download.

Per-episode scene transitions


The course carries a series of cartoons, not eight unrelated images. One Curacaoan analyst runs through all of them, starting as a fan with a wall map on the index page and finishing at the tiller of a boat heading out of the harbour. The renders and full scene briefs are in scene-briefs.md.

This is a deliberate change from the Aruba course, where the images were decoration and instructors were told to ignore them. Here you may point at them. Use the callback in Episode 6 and again in Episode 7, where the payoff sits.

Each episode opens with its scene image and a one-line quip caption. The captions carry most of it. The transition lines below are optional single beats for instructors who want one on arriving at a new episode.

Ep Caption on page Optional transition line
Index Twenty-six players. Ten countries. One spreadsheet. (Shown on the landing page and on the opening slide. No spoken line; let people read it while the room settles.)
1 One gate charges you every season. The other one only asks you to learn the way in. (Episode 1 opens with the workshop’s full opening sequence; no separate transition needed.)
2 The dominoes can wait. The console is blinking. “Laptop open, console blinking, iguana unimpressed. Time to type something.”
3 The stew takes twenty minutes. The chopping takes an hour. “Three jars on the counter today: filter, select, mutate. Everything else in dplyr is a variation on those three.”
4 SPSS hands you a chart. ggplot2 hands you a grammar. “Those are the Handelskade houses and they are also a bar chart. By the end of this episode you will be writing the sentence that draws them.”
5 Sometimes the review says nothing happened. You report that too. “The tests you know from SPSS are all here. What is new is that one of today’s two comes back null, and we are going to report it anyway.”
6 The squad changed overnight. Again. Good thing the report rebuilds itself. “This is where R Markdown earns the price of admission. New squad in, finished document out, one button.”
7 You have the basics. The map runs well past the harbour mouth. “She started this course pinning photos to a wall map. She is leaving the harbour with a chart. That is roughly where you are too.”

Pick one beat, deliver it, move into the page’s first heading. Do not stack a second sentence on top.

What the four-hour format costs


The course runs 09:00 to 13:00, which is 480 minutes of room time for 380 minutes of episodes. The overhead is welcome, recap, two breaks a day and the two wrap-ups. Three things were cut to make it fit, and instructors should know which:

The open lab on Day 2 is gone. Homework review moved into the 20-minute Day 2 recap, so protect that slot; it is now the only place a stuck participant gets unstuck.

Exercise time in Episodes 2, 3 and 4 took most of the reduction, roughly 25 minutes each. Episode 5 was left slightly long on purpose because it is the hardest and it carries the null result.

The Day 1 questions-and-consolidation block is gone. Fold consolidation into the end of each episode rather than saving it up.

The AI and package-trust material, added 22 September 2026


Two additions, both deliberate and both cuttable under time pressure.

Episode 2 gains a callout after the install.packages() versus library() box: what CRAN’s review actually guarantees, the July 2026 Hugging Face intrusion as the contrast (a malicious dataset abusing code execution paths in the upload processing pipeline, so reading data is running code), and what to do where an employer blocks installations, which is an internal mirror plus renv pinning. Budget four minutes spoken. Two Central Bank staff are in the room and their institution treats package installation as a security exposure, so deliver this as a fair position with a professional answer rather than as an obstacle. If Episode 2 is running behind, say the last paragraph only and leave the rest to be read.

Episode 7 gains “Letting AI write the R” before the learning resources: that generating R is now the faster route for routine work, that the job becomes judging the output, the four failure modes that read well, and the rule about never pasting supervisory or personal data into a public model. Episode 7 is self-guided, so this costs nothing from the timetable. It is worth naming out loud in the wrap-up even if nobody reads the section.

The index says both are covered, which matters for institutions deciding whether to send staff.

What went wrong on Day 1, Curacao, 23 September 2026


Recorded the same evening, from delivery. Every item here cost time in the room.

Getting the data into R was the whole problem. Participants had the files in a folder and no clear route from there to a loaded data frame. The instructions said to create a project but never said, in one line, how to point R at the folder afterwards. Rendell set the working directory to the data folder rather than to its parent, which breaks every "data/..." path in the course by one level, and the paths had to be edited live. Episode 2 now has a step 4 that prints getwd() and list.files("data") before anything else, and setup.md says to pick the parent folder in as many words.

Nobody knew whether to upload through the menu or load by command. Walk the room through one route and name it as the route. Do not offer both.

The downloaded files collided. All three arrived under the same name and overwrote each other, so everyone had to rename before anything would load. setup.md now says to check names and extensions in the Downloads folder first.

The Excel import silently returned half the data. read_excel() takes the first sheet, and this workbook splits 48 rows of Curacao and 46 of Aruba. Episode 2 now counts the rows, names the trap, and stacks the sheets with bind_rows(). Teach it as the lesson it is: an Excel import that returns half your cases looks exactly like one that worked.

The fix that removes all of this at once is reading from the web. Episode 5 now opens with a base URL and three read_csv(paste0(base, ...)) lines, so Day 2 starts with data in memory whatever state anyone’s folders are in. Put that block on the screen first and let people catch up while you talk.

There was no second pair of hands. One instructor cannot debug twenty installations and keep a timetable. Where no helper is funded, recruit two or three participants who got through setup quickly and ask them openly to help their neighbours. It costs nothing and it works.

Common Issues


  • Installation problems: The registration email asks participants to reply with any install error before the day, so most should arrive fixed. Have a USB drive with R and RStudio installers as backup.
  • Typos: SPSS users are not used to typing commands. Expect many syntax errors. Normalize this: “error messages are how R talks to you.”
  • Parentheses and quotes: The most common beginner errors. Show how RStudio auto-completes these.
  • Loading packages: Participants will forget library(). Remind them at the start of each episode.

Local Data Notes


The course uses Dutch Caribbean datasets to keep examples relevant. Everything learners compute on is generated by scripts/00_build_teaching_data.R from verified sources and committed to episodes/data/. Nothing is scraped live in the room.

  • blue_wave_squad.csv, player-level squad lists for the four ABC island national teams, 94 players. Primary teaching dataset, Episodes 2 to 4 and 6. Also shipped as .xlsx (two sheets) and .sav for the import demonstrations. Curaçao’s men are the World Cup squad. Scraped from pinned Wikipedia revisions by scripts/00_build_teaching_data.R; variable definitions and limitations are in episodes/data/blue_wave_squad_codebook.md.
  • blue_wave_squad_2026-09.csv, the September 2026 call-up, same columns, 91 players. Episode 6 only, where the capstone report is re-knit on it and the Episode 3 region code is shown misfiling Bosnia and Gibraltar.
  • fifa_rankings.csv, FIFA rank and points against population and diaspora for 211 national associations. Continuous variables for Episodes 4 and 5.
  • diaspora_change.csv, diaspora stock in 1990, 2010, and 2024. Supplies the paired t-test in Episode 5.
  • island-research-reference-data, a country reference list with SIDS, SNIJ, and World Bank classifications, pulled live from GitHub in Episode 1 and Episode 7. Offline fallback committed at episodes/data/countries_backup.csv.
  • CAS_election_data, Aruba, Curacao, Sint Maarten election results 1985-2025. Used in Episode 7.
  • World Bank indicators via WDI (Episode 7) and CBS Netherlands via cbsodataR (Episode 7).

Regenerate the derived files with Rscript scripts/00_build_teaching_data.R from the repository root. Test all live downloads before the course; URLs and APIs change.

The gaps in the squad data are the point

Two of 94 players cannot be placed at a club, and the squads are only as current as Wikipedia. Episodes 3, 5, and 6 each stop to name what is being excluded and what it costs. Do not tidy this away or apologise for it. Most participants have never been shown what to do with a gap other than delete it, and this is the most transferable thing in the two days.

Two is a small number, and that is deliberate rather than unfortunate. The habit of reporting exclusions is easiest to build when the exclusion changes nothing.

The island comparison does not work, and Episode 5 uses that

Testing whether Curaçao and Aruba differ on players-based-abroad returns a p-value around 0.64. Testing men against women returns 0.006. Episode 5 runs both in that order on purpose: the first thing you try fails, the second works, and the write-up has to admit both. Do not skip to the one that works.

Note on the elections example

The Episode 7 election-data example summarises fragmentation across all parties in each Curacao election rather than singling out any one party. In a room that may contain civil servants and ministry staff, filtering on a single party name reads as partisan even when it is not meant to. If extending the example live, default to all-parties views. If a participant asks why, this is a small editorial choice that protects the course and the network’s neutrality, and it is worth naming briefly.

Note on the Papiamentu example

The Episode 7 text-analysis sample is written in Curacao orthography, not the Aruban etymological standard used in the master course. The accompanying callout uses the difference to make a point about stopword lists being analytical choices. If someone in the room raises the orthography question, that is a good outcome, not a derailment.

Train-the-Trainer


This course is designed for replication. If you are adapting it for another island or institution: 1. Replace datasets with locally relevant equivalents 2. Adjust the SPSS operations covered based on your pre-course survey results 3. Keep the “wow first, skills second” structure 4. All materials are CC-BY 4.0, so please attribute the DCDC Network

The Case for Switching


Live demo script

This is the complete script to run live. Practice this before the workshop. Make sure tidyverse is installed. Everything the demo needs comes with it.

The demo has two halves. The first shows R reaching out to a live research dataset the network owns. The second lands the local hook on the squad data the rest of the course uses. Run the first if the Wi-Fi is cooperating, and the second regardless, because the second is the one that makes the room lean forward.

Part A, step 1: Frame the source before you type

Before the first keystroke, name what the room is about to see. The CSV about to load is in a GitHub repository maintained at the University of Aruba, island-research-reference-data, part of the DCDC Network’s shared infrastructure for island research. The network owns and curates it, and keeps it up. That framing matters, because the point here goes past R being able to read a URL: the data layer underneath belongs to us.

Part A, step 2: Pull the SIDS reference list

Open a new R script in RStudio and type or paste the following. Run it line by line so participants can watch each step.

R

# Load packages (install tidyverse once, before the workshop)
# install.packages("tidyverse")
library(tidyverse)

# Pull the UA island-research reference list straight from GitHub
countries <- read_csv(
  "https://raw.githubusercontent.com/University-of-Aruba/island-research-reference-data/main/countries/countries_reference_xlsform.csv"
)

# Quick look at what we got
head(countries)

Pause. Point out: “No browser. No download dialog. No save-as. The file is now a live object in my session, with over a dozen columns per country.”

Part A, step 3: Filter to SIDS and chart by region

R

countries |>
  filter(is_sids == 1) |>
  count(wb_region) |>
  ggplot(aes(x = reorder(wb_region, n), y = n)) +
  geom_col(fill = "#2e8894") +
  coord_flip() +
  labs(
    title = "Small island developing states by World Bank region",
    x = NULL,
    y = "Number of SIDS"
  ) +
  theme_minimal(base_size = 14)

Pause again. Key talking points:

  • “This chart is ready for a report as it stands. Title, axis labels, colour, proportions, all set in code.”
  • “If the UA team adds a country to the reference list tomorrow, I re-run this script and the chart updates. No re-click, no re-export.”
  • “Every editorial choice, what counts as a SIDS and which region goes where, is traceable, because the definitions sit in the source CSV you just pulled.”

Part B: the local hook

This is the moment the course earns the room’s attention. Load the squad data and answer the question the island already argues about.

R

squad <- read_csv("data/blue_wave_squad.csv")

squad |>
  filter(team_code == "CUW-M") |>
  count(club_country, sort = TRUE)

Read the result out loud. Then draw it:

R

squad |>
  mutate(
    island = if_else(str_starts(team_code, "CUW"), "Curaçao", "Aruba"),
    home_code = if_else(str_starts(team_code, "CUW"), "CUW", "ABW"),
    region = case_when(
      club_country == "X"       ~ "Unknown",
      club_country == home_code ~ "Home island",
      club_country == "NLD"     ~ "Netherlands",
      club_country == "USA"     ~ "North America",
      club_country %in% c("GBR", "GRC", "TUR", "DEU", "BEL", "CHE", "XKX") ~ "Rest of Europe",
      .default                  = "Rest of world"
    )
  ) |>
  count(island, region) |>
  ggplot(aes(x = island, y = n, fill = region)) +
  geom_col(position = "fill") +
  scale_y_continuous(labels = scales::percent) +
  labs(
    title = "Where the players actually play",
    subtitle = "All four ABC island national squads, 2026",
    x = NULL, y = NULL, fill = NULL
  ) +
  theme_minimal(base_size = 14)

Talking point: “Two islands, 80 kilometres apart, with populations and colonial histories that track each other closely. Both send most of their players abroad. But look where. Curaçao’s are scattered across ten countries. Aruba’s are overwhelmingly in one. Nobody in this room needed a statistician to care about that question. What you needed was a way to ask it, and you just watched it take twelve lines.”

Do not resolve the question. Let the room argue. Say: “By Wednesday you will be able to run this yourself, and change it to ask what you actually want to know.”

Step 4: Show the contrast with SPSS

Ask the audience: “How would you have done this in SPSS?”

Walk through it slowly. Make it sting a little. This is the moment the cost of the current workflow lands.

  1. Go looking for the squad lists. Wikipedia? FotMob? The federation’s Facebook page? Pick one and hope it is current.
  2. Copy 94 player names, clubs, and positions into Excel by hand. An hour if you are quick and do not lose concentration.
  3. Discover that the club’s country is only shown as a little flag you cannot copy, that two players have no club listed at all, and that Jong Holland is a Curaçao club despite the name. Clean by hand.
  4. Import into SPSS. Recode club country into a grouping variable, because the raw text is unusable as it stands.
  5. Analyze > Descriptive Statistics > Frequencies. Copy the output table.
  6. Graphs > Chart Builder, drag variables, format the chart, copy, paste into Word.

Then say: “That is half a morning, on a good day. In R it was twelve lines and ten seconds, and the next person who asks gets the same answer from the same file.”

Backup plan

If the Wi-Fi is unreliable, the reference list is saved locally at episodes/data/countries_backup.csv. Swap the read_csv() call for the local path:

R

countries <- read_csv("data/countries_backup.csv")

Then proceed with the filter() |> count() |> ggplot() pipeline as normal. Part B is local already and needs no network at all, which is why it is the half you never skip.



Your First R Session


Pacing notes

  • Spend time on the RStudio pane orientation. Have participants identify each pane on their own screen before moving on.
  • The <- assignment operator trips people up. Give them a few minutes to practice creating objects with different names and values.
  • When loading tidyverse, the startup messages can be alarming to beginners. Reassure them that the “Attaching packages” and “Conflicts” messages are normal and expected.
  • Do the “Before you import” subsection as a whole-room moment, not as reading. Project the download page, walk everyone through the browser download, the New Project dialog, and creating the data subfolder. Wait for green sticky notes in the Files pane before typing read_csv(). This is the most common point of failure in the course and it is worth five deliberate minutes up front to avoid twenty scattered minutes of troubleshooting later.
  • The .sav import is worth a beat of theatre. Most of the room has years of .sav files they assume are trapped. Let them see one open in three lines.
  • Encoding warning is not pedantry. If a participant opens the CSV in Excel and saves it, the ç in Curaçao usually breaks and they will hit an error later that looks unrelated. Say it once, firmly.
  • If a participant cannot download (blocked network, locked laptop), have a USB stick or shared-drive copy of the three files ready as fallback.
  • The two-way table at the end is a deliberate cliffhanger. Do not explain it. Let someone in the room say it out loud.


Data Manipulation


Teaching tips

  • Write the pipe |> on the whiteboard and say “and then” out loud every time you use it. This mental model sticks.
  • Build pipelines live, one step at a time. Run after each added line so participants can see how the output changes.
  • The most common beginner mistake is putting |> at the start of a line instead of at the end of the previous line. Emphasise that the pipe goes at the end of the line, so R knows the expression continues.
  • Compare nested function calls to piped code side by side. The readability advantage sells itself.
  • The mean() of a logical is the single highest-leverage idea in this episode. SPSS users are trained to build dummy variables by hand. Show them the shortcut, then show them the SPSS way they would otherwise have used, and let the contrast do the work.
  • The filter(club_country %in% c("CUW", "ABW")) result usually produces an audible reaction in a Curaçao room. Let it. Then say “we are not explaining that today, we are learning how to ask it.”
  • If someone asks why two players have club_country == "X", the honest answer is that the source is a Wikipedia squad table, those two rows list no club, and the build script recorded that rather than guessing. Point them to blue_wave_squad_codebook.md in the data folder. This is a good moment for a word about documenting your own uncertainty.
  • Someone may well ask how current the squads are. They are as current as Wikipedia, which is to say: maintained by volunteers, lagging real call-ups by weeks or months, and not necessarily consistent across the four pages. Say so. It is a better answer than pretending, and it sets up Episode 5.


Visualization with ggplot2


Instructor note

The faceting example is a good place to pause and let learners experiment. Encourage them to try:

  • facet_wrap(~ island, ncol = 1) to control the layout
  • swapping the facet formula to facet_grid(island ~ gender) and asking which comparison each version makes easy
  • removing coord_flip() to see why the labels needed it

The position = "fill" versus default stacking contrast is the most transferable idea in the episode. Most people in the room have at some point presented a counts chart where the group sizes differed, and drawn a conclusion from bar height that the proportions did not support. Say that out loud.

Do not skip past the missing-data warning on the first scatterplot. Learners who have been taught to make warnings go away need to hear, from you, that this particular one is information.



Statistical Analysis in R


Instructor note

This is a good time to reinforce the reproducibility advantage. In SPSS, if a reviewer asks you to re-run an analysis on a different subset, you click through the dialogs again. In R you change one line and re-run the script.

Emphasise that broom::tidy() produces a data frame, so learners can use every dplyr verb from Episode 3 on their statistical results: filter to significant terms, arrange by p-value, join to labels.

Three moments are worth slowing down for, and none of them is about syntax.

The non-significant diaspora coefficient is the most important. Most of the room has been taught, implicitly, that a good table has stars in it. Say plainly that you are leaving a dead predictor in the model on purpose, and why. If anyone asks whether they should drop it, that is the discussion you want.

The pair of chi-square tests is the second. Run the island one, let the room see the p-value, and let the disappointment sit for a moment before you run the gender one. The sequence is the lesson: the first thing you tried did not work, you tried a second thing, and the write-up has to admit both. Almost everyone in the room has at some point reported only the second.

The third is the filter(club_country != "X") line. Ask what it did before you tell them. Someone will spot that it removed two players. Two is small enough that nobody would object, which is exactly why it is a good example: the habit has to be built when the stakes are low.



Reproducible Reporting


Common knitting problems

The most common issue is that knitting fails because the .Rmd file does not load packages or data that earlier chunks depend on. Remind participants that knitting starts from a blank environment. Every package and dataset must be loaded within the .Rmd file itself, even if it is already loaded in their current R session.

Another common issue is file paths. If participants write read_csv("data/blue_wave_squad.csv"), the working directory during knitting is the folder where the .Rmd file is saved. Make sure the data file is in the right relative location.

Third, and specific to this dataset: encoding. If a participant’s Windows locale mangles the ç in Curaçao during knitting, have them save the .Rmd as UTF-8 via File > Save with Encoding. It is worth mentioning pre-emptively rather than debugging it live in four separate laptops.



Timing the re-run

This section costs about eight minutes, most of it the download and the knit. If the room is behind, run it from the front instead: pre-rendered copies of both reports are in episodes/files/, and the region callout needs nothing but the chunk above. Keep the match itself out of it. The finding holds whatever the score on Thursday night.



Where to Go from Here


Before delivery: two things to replace

The two survey links above are the Aruba pilot’s forms. Either create Curaçao cohort forms and swap the URLs, or accept mixed responses across editions and add a cohort question to the existing forms. Decide before promotion goes out, not on the morning.

The Papiamentu sample text in the text-analysis section is written in Curaçao orthography, which differs from the Aruban standard used in the master course. That was deliberate for this room. If you would rather teach it as a comparison, the Aruban version of the same passage is in the master course repository and running both through the same stopword list makes the point in the callout better than the callout does.