Define the Two Terms Like a Scientist (Not a Philosopher)

  • Correlation is a statistical relationship: two variables move together — positively (both rise/fall together), negatively (one rises as the other falls), or not at all. It is measured by a coefficient (ranging from −1 to +1) such as Pearson's r, and it describes what the data do. Correlation is a mathematical fact about a dataset, nothing more.
  • Causation is a mechanistic claim: changes in one variable produce changes in the other through a defined pathway. Causation always implies correlation (a cause must show up as association in data), but correlation does not imply causation — because there are at least five other ways the association could have arisen (below).

The crux: a correlation tells you the data are connected. It does not tell you why they are connected. Science assignments reward the student who names the possible "whys" and works to rule them out, not the student who asserts the first plausible story.

The Five Non-Causal Explanations (Memorize This List)

Whenever you observe a correlation in an assignment, your first job is to audition these five alternatives before concluding anything:

1. The confounder (the big one)

A confounding variable is a third factor associated with both variables that explains their apparent link. The classic: ice cream sales and drowning both rise in summer — but the confounder is temperature/season, which drives both. Confounders are everywhere in student datasets: the "study hours correlate with grades" finding is confounded by prior ability; "screen time correlates with anxiety" is confounded by sleep. Study.com on confounding variables Your assignment's hidden confounder is the single most common place students lose analysis points.

2. Reverse causation

The arrow may point the other way. "Depression correlates with social isolation" — but does isolation cause depression, or does depression cause withdrawal? Without temporal or experimental evidence, the data cannot tell you the direction. Many student papers assert a direction the dataset simply does not support.

3. Chance (sampling error)

A correlation in one sample may be a fluke. Small samples produce unstable correlations; every published r should carry a p-value precisely so readers know whether the association could plausibly be chance. In your assignments: "our sample was only 20 participants" should automatically lower your confidence in any causal-sounding claim.

4. Selection bias

The sample may not represent the population, and the correlation may be an artifact of who was included. Studying the "relationship between exercise and GPA" using only varsity athletes bakes selection into the data, since athletes are pre-screened on both variables.

5. The measurement problem

The "correlation" may live in the instruments, not the world. If your survey measures both variables with the same tool (self-report bias), or if two variables are mathematically derived from each other, the association is partly manufactured. Example: "temperature in Celsius correlates with temperature in Fahrenheit" — a perfect correlation that is pure definitional overlap, not science.

The Methods That Do Support Causal Claims

Professors do not expect you never to use the word "cause" — they expect you to know what evidence licenses it:

  • Randomized experiments: random assignment to treatment/control conditions breaks the confounder problem by design. This is the gold standard ("the group assigned to the new study intervention improved significantly compared to control").
  • Longitudinal designs with temporal ordering: the cause is measured before the effect, ruling out reverse causation (a two-year panel showing isolation precedes depression is stronger than a one-shot survey).
  • Natural experiments and quasi-experiments: policy changes, cutoffs, or exogenous shocks approximate random assignment (e.g., comparing outcomes before/after a law change).
  • Statistical control (with caveats): regression models that adjust for measured confounders reduce the risk — but they can never adjust for the confounders you did not measure. In student writing, say "controlling for age and income, the association persisted," then still hedge: "although unmeasured confounders cannot be ruled out."

The writing formula: match the strength of your language to the strength of the design. A randomized experiment licenses "caused" or "produced"; a cross-sectional survey licenses "was associated with" or "was related to" — never "proved."

How to Write It Correctly (The Phrase Bank)

Amateur phrasing to drop: "proves," "causes," "led to," "is due to," "confirms," "shows that X causes Y." Professional phrasing to adopt:

  • "was positively correlated with (r = .42, p < .05)"
  • "was associated with / was related to"
  • "is consistent with a causal interpretation, although confounding cannot be ruled out"
  • "the temporal ordering suggests that…, but reverse causation remains possible"
  • "random assignment allows a stronger inference: the intervention caused…"
  • "after adjusting for [confounders], the association remained significant"

Notice the pattern: each sentence either states the statistical fact (correlation) or explicitly earns the causal word by naming the design (experiment, temporal order, statistical control) that supports it. That syntax — data first, causal word only when the design earns it — is the entire skill in one habit.

The Discussion-Section Move That Wins Points

In lab reports and analysis papers, a paragraph that performs the causal audit visibly earns disproportionate credit. Structure it as a checklist:

"The observed correlation between sleep and exam performance (r = .38) could, in principle, reflect several non-causal processes. First, a confounder — trait conscientiousness — plausibly drives both better sleep and better preparation; our survey did not measure it. Second, reverse causation is possible: students who perform poorly may sleep worse from anxiety. Third, with n = 34, sampling error cannot be excluded. The temporal ordering (sleep logged for two weeks before the exam) partially addresses reverse causation, but only a randomized sleep intervention would license a causal claim. We therefore conclude that sleep and performance are associated — and treat the causal pathway as a hypothesis for future work."

That paragraph, written in plain English, is better than a page of fancy statistics — because it demonstrates exactly the critical thinking the assignment exists to test.

Spotting Correlation–Causation Errors in Your Own Drafts (The Audit)

Before submitting, run these three passes:

  1. The verb scan: find every "causes," "proved," "led to," "is due to." For each, ask: what design supports this word? If the answer is a survey or observational data without temporal order, downgrade the verb to "associated with."
  2. The confounder hunt: for every correlation you report, generate three candidate confounders by asking "what else correlates with both variables?" If you cannot name at least one, you have not tried hard enough — and the grader will name three.
  3. The direction check: could the reverse arrow explain the data? If yes, say so — explicitly marking reverse causation as possible is itself a sign of mastery.

Also: beware the ecological fallacy (drawing individual conclusions from group data) and the correlation-of-trends trap (two variables rising over time almost always correlate — e.g., my height and Google's stock price both increased from 2000–2020 — but that tells you nothing). Time-trend correlations are a favorite hidden trap in assignment datasets; check before you claim.

Conclusion

Correlation is a fact about data; causation is a claim about mechanisms — and the distance between them is filled with confounders, reverse arrows, chance, selection, and measurement artifacts. The scientist's job, in a lab report, a discussion post, or a full paper, is to audition those alternatives out loud and only then say what the data support. Match your verbs to your design: "associated with" for observational data, "caused" only for experiments and designs that earn it. Run the verb scan, the confounder hunt, and the direction check on every draft. Your professors are not grading whether you found the truth — they are grading whether you know how much your data entitle you to claim. That calibration, demonstrated in one clean paragraph, is what separates the chemistry of the assignment from the chemistry of the grade.

Your next step: Take the dataset from your current assignment and write the "causal audit paragraph" right now — correlation coefficient, three candidate non-causal explanations, your design's actual strengths and limits, and the correct concluding verb. Show it to a classmate and ask them to find a confounder you missed. There is always at least one — and finding it before your professor does is the whole point of the exercise.