Agent Sigma's Field Notebook
Case Files from Statsville โ a story-based guide to Collecting Data
Welcome, new recruit. The Statsville Data Bureau collects information every day โ surveys, experiments, opinion polls โ but not every dataset can be trusted. Your job is to follow Agent Sigma through four real cases, learn how good data is collected, and spot the tricks that make data go wrong. Read each case, study the field notes, work the examples, then crack the practice questions to close the case.
The Sampling Mystery
How do you pick a group of people that actually represents the whole town? SRS, stratified, cluster & systematic sampling.
The Bias Trap
The numbers don't add up. Learn to spot undercoverage, nonresponse, response, and voluntary response bias.
The Experiment Lab
Randomization, control groups, blocking, and blinding โ how to design an experiment nobody can poke holes in.
Cause or Coincidence?
Breakfast and better grades โ connected, or just a coincidence? Observational studies vs. experiments and confounding variables.
The Sampling Mystery
"Every sample tells a story โ but only if you pick it the right way."
Mayor Whitfield wants to know: do Statsville residents support building a new park? She sends three assistants out to find 100 opinions each. Agent Sigma is called in when the three results come back wildly different โ someone picked their sample the wrong way, and it's your job to find out how a sample should be picked.
A good sample is a small group that honestly represents the larger population. The method you use to choose that group decides whether you can trust the result.
SRSSimple Random Sample
Every individual โ and every possible group of individuals โ has an equal chance of being chosen. Like pulling names blindly from a hat.
STRATStratified Sampling
Split the population into similar subgroups (strata), then take a separate SRS from each stratum. Guarantees every subgroup is represented.
CLUSCluster Sampling
Split the population into clusters that each look like a mini-population, then randomly choose a few whole clusters and use everyone inside them.
SYSSystematic Sampling
Pick a random starting point, then choose every k-th individual after that (e.g., every 10th person).
โConvenience Sampling
Choosing whoever is easiest to reach. Fast, but almost always biased โ not random at all.
โVoluntary Response
Individuals choose to participate themselves (call-in polls, online surveys). Overrepresents strong opinions.
Assistant #1 sorted Statsville into 5 neighborhoods, then randomly chose 2 entire neighborhoods and surveyed every single house inside those two.
Whole groups (neighborhoods) were randomly selected, and everyone inside the chosen groups was surveyed โ that's the signature of cluster sampling. Notice the other 3 neighborhoods got zero representation, which is fine as long as each neighborhood is a realistic mini-version of the whole town.
Assistant #2 asked every 10th customer leaving the grocery store, starting from a randomly chosen 4th customer.
A random starting point (the 4th customer) followed by a fixed interval (every 10th after that) is the classic pattern of systematic sampling.
Assistant #3 divided students by grade level (9th, 10th, 11th, 12th) and randomly selected 20 students from each grade.
The population was split into subgroups (grades) first, and a separate random sample was taken from within each subgroup โ that's stratified sampling, not cluster, since not every student in a grade was chosen.
A local radio show asks listeners to call in and vote for their favorite park design.
Listeners chose whether to participate. People with the strongest feelings are the most likely to call, so quieter opinions get left out โ this is not a trustworthy sample.
The Bias Trap
"The sample was random. So why do the numbers feel... off?"
Agent Sigma reviews the mayor's park survey and something is wrong: the results claim 90% support, but grumbling around town suggests otherwise. The sampling method wasn't the problem this time โ bias crept in somewhere else. Bias is any systematic tendency for a sample to misrepresent the population, and it can sneak in even when the method looked fine on paper.
UCUndercoverage Bias
Some part of the population is left out of the sampling frame entirely, so it has zero chance of being chosen (e.g., surveying only landline owners).
NRNonresponse Bias
People chosen for the sample don't respond โ and the people who do respond tend to differ from those who don't.
RBResponse Bias
Respondents answer inaccurately because of how a question is worded, who's asking, or social pressure to give a "good" answer.
VRVoluntary Response Bias
People with strong opinions are far more likely to volunteer, skewing results toward the extremes.
Agent Sigma mailed 1,000 surveys to random Statsville addresses. Only 60 came back.
The sample was chosen randomly โ the problem is that only 6% replied. People who feel strongly enough to mail a survey back may not represent the other 94%, so the results likely skew toward extreme opinions.
A survey question reads: "Given the rising crime problems in Statsville, do you support more police patrols?"
The question is loaded โ mentioning "rising crime problems" nudges people toward answering "yes," regardless of their real opinion.
A city survey is conducted only by calling home landlines, completely missing every household that uses only cell phones.
Cell-phone-only households were never even in the sampling frame โ they had no chance of being selected at all, which likely skews the sample toward older residents.
A TV news show asks viewers to text in "yes" or "no" on whether the park should be built.
Only viewers who feel strongly enough bother to text in, so the result overrepresents passionate opinions on both extremes.
The Experiment Lab
"Anyone can run a test. Only a careful design can be trusted."
A Statsville scientist claims her new "Focus Juice" boosts test scores. Agent Sigma is brought in to make sure the experiment is designed fairly โ because a badly designed experiment can make an ordinary juice look like a miracle drink, and a great treatment look useless.
VARExplanatory & Response Variable
The explanatory variable (treatment) is what's deliberately changed โ like the drink given. The response variable is what's measured afterward โ like the test score.
CTRLControl Group
A comparison group that does not get the treatment (often a placebo), giving a baseline to measure the real effect against.
RANDRandomization
Subjects are randomly assigned to treatment groups so the groups start out similar in every way โ this cancels out lurking variables like existing ability.
REPReplication
Each treatment is applied to many subjects, not just one or two, so the results reflect a real effect rather than random chance.
BLKBlocking
Group subjects into blocks that share a trait likely to affect the outcome (like sex or age), then randomize treatments within each block. Reduces variability from that known trait.
BLINDBlinding
Single-blind: subjects don't know their group. Double-blind: neither subjects nor the people evaluating results know โ this prevents the placebo effect and biased grading.
The scientist gives Focus Juice to the 10 fastest-finishing students and plain water to the 10 slowest-finishing students, then compares test scores.
The groups were not randomly assigned โ they were already different in ability before the juice was ever poured. Any score difference could just be the original skill gap, not the juice. This is a confounded design.
Instead, 40 students are randomly assigned by coin flip to Focus Juice or a juice-flavored placebo. Neither the students nor the teacher grading the tests knows who got which drink until after grading is finished.
This design uses randomization (fair starting groups), a control group (the placebo), and double-blinding (no one's expectations can sneak in). This is the gold standard for testing a real effect.
Since boys and girls might react to the juice differently, students are first split by sex. Within each group, students are then randomly assigned to juice or placebo.
Splitting by a known trait (sex) before randomizing is blocking โ it removes that source of variability so the real treatment effect is easier to see.
All 40 students are given Focus Juice, and their scores are compared to their own scores on a test from last month.
There is no comparison group getting a different (or no) treatment at the same time โ other things could have changed since last month (easier test, more studying), so the juice can't be given credit.
Cause or Coincidence?
"Two things happening together doesn't mean one caused the other."
The Statsville Gazette runs a headline: "Kids Who Eat Breakfast Get Better Grades!" Parents start forcing pancakes on their kids the very next morning. But Agent Sigma isn't convinced yet โ did breakfast really cause the better grades, or is something else going on?
OBSObservational Study
Researchers watch and record what's already happening, without imposing any treatment. Can reveal an association between variables โ but not cause and effect.
EXPExperiment
Researchers deliberately impose a treatment and control other variables. A well-designed (randomized, controlled) experiment can support a causal conclusion.
CONFConfounding (Lurking) Variable
An outside variable connected to both the explanatory and response variable, making it impossible to tell which one is really responsible for the effect.
KEYAssociation โ Causation
Two variables moving together could be a coincidence, reverse causation, or explained by a confounder โ randomized experiments are the strongest tool for proving true cause and effect.
Researchers observed 500 students and found that students who eat breakfast tend to have higher grades. No breakfast was assigned by the researchers.
Because breakfast wasn't randomly assigned, we can't rule out a confounding variable โ for example, families who prioritize breakfast might also prioritize homework and sleep, and that could be the real reason for the higher grades.
A school randomly assigns half its students to a free breakfast program for a semester and compares their grades to a control group that receives no breakfast program.
Because breakfast was randomly assigned and there's a control group, differences in grades can more confidently be credited to the breakfast program itself.
A study finds that towns with more ice cream shops also tend to have more drowning incidents.
Ice cream shops don't cause drownings. Hot weather (a lurking variable) increases both ice cream sales and swimming, which increases drowning risk.
Case Closed: Statsville Summary
Everything Agent Sigma learned, filed in one place.
Case 1 โ Sampling Methods
- SRS: every individual/group equally likely
- Stratified: subgroups first, then SRS within each
- Cluster: pick whole random groups, use everyone inside
- Systematic: random start, then every k-th person
- Convenience / Voluntary response: not random โ biased
Case 2 โ Types of Bias
- Undercoverage: a group is left out of the frame entirely
- Nonresponse: chosen subjects don't answer
- Response bias: question wording / interviewer causes inaccurate answers
- Voluntary response: strong opinions self-select in
Case 3 โ Experimental Design
- Randomization: spreads out lurking variables evenly
- Control group: baseline for comparison (often placebo)
- Replication: treat many units, not just one
- Blocking: group by a known trait, then randomize within
- Double-blind: neither subject nor evaluator knows the group
Case 4 โ Observation vs. Experiment
- Observational study: just watch โ shows association only
- Experiment: impose treatment โ can support causation if randomized
- Confounding variable: hidden factor tied to both variables
- Golden rule: association does not equal causation
| Term | Quick Definition |
|---|---|
| Population | The entire group you want to learn about. |
| Sample | The smaller group actually studied. |
| Sampling frame | The actual list used to select the sample. |
| Bias | A systematic tendency for a sample to misrepresent the population. |
| Treatment | The specific condition applied to subjects in an experiment. |
| Placebo | A fake treatment used so the control group doesn't know it's untreated. |
| Lurking variable | An unmeasured variable that could explain an observed relationship. |
| Causation | One variable directly produces a change in another โ best shown by randomized experiments. |
๐ Statsville Data Bureau โ Certificate of Completion
Cases closed: 0 / 4
"A well-collected sample is worth a thousand guesses." โ Agent Sigma