Correlation means two variables change together; causation means a change in one produces a change in the other. To tell them apart, do not rely on the size of a correlation or a small p-value. Check whether the proposed cause came first, whether plausible rival explanations were addressed, and whether the study design can support a causal comparison. Randomized experiments usually provide stronger causal evidence than observational snapshots, but every design still needs a bias check.
This guide gives students a repeatable five-question audit for judging causal claims in journal articles, literature reviews, and news reports.
A correlation describes the direction and strength of a relationship. The Australian Bureau of Statistics explains that the correlation coefficient, r, ranges from −1.0 to +1.0: its sign gives direction, while its magnitude describes the strength of a linear relationship. An r near zero can still miss a nonlinear relationship, and even an r near either extreme does not identify a cause.
A causal claim is stronger. It says that, for a defined population and outcome, changing an exposure would change what happens compared with a relevant alternative. That comparison is often called a counterfactual: what would have happened to the same people at the same time under a different exposure. Because both outcomes cannot usually be observed for the same person, researchers use a comparison group and a design intended to make that group credible.
Fast rule: ask what comparison would reveal the effect of changing X. Then ask whether the study actually approximated that comparison.
The Centers for Disease Control and Prevention warns that an observed association may reflect a causal connection, but may instead result from chance, selection bias, information bias, confounding, or other errors in design and analysis. These alternatives are not technical footnotes. They are the main reasons a result can be statistically real in the sample yet fail to represent the causal effect claimed.
A confounder is a pre-exposure factor related to both the exposure and the outcome. In the classic example, ice-cream sales and sunscreen sales rise together because hot weather increases both. Buying ice cream does not make people buy sunscreen; temperature creates the association. The Cochrane Handbook defines confounding in non-randomized intervention studies as common causes of intervention choice and outcome, making the observed association differ from the causal effect.
Adjustment can reduce confounding only for variables that were identified, measured well, and modeled appropriately. Cochrane distinguishes residual confounding, where control is incomplete, from unmeasured confounding, where an important common cause is absent from the analysis. Therefore, “adjusted for age and income” is useful information, not a guarantee that every alternative explanation has disappeared.
Suppose a survey finds that students who sleep less report more stress. One story is that short sleep raises stress; another is that stress disrupts sleep. A single cross-sectional measurement cannot reliably establish which came first. Longitudinal timing helps, but time order alone does not remove confounding or measurement error.
Selection bias appears when entry into or retention in a study depends on factors connected to both exposure and outcome. Information bias appears when variables are measured differently or inaccurately across groups. Chance remains relevant too, but statistical significance addresses sampling variation under a model; it does not rule out systematic bias, confounding, reverse causation, or a poorly chosen comparison. The CDC Field Epidemiology Manual explicitly separates the role of chance from these other explanations.
Study design changes which rival explanations remain plausible and what assumptions a causal conclusion requires. Read the methods before accepting the verbs in an abstract or headline.
Cochrane explains that unbiased randomized trials can estimate causal effects, while non-randomized studies face greater risks from confounding, selection, and measurement. Randomization strengthens a comparison, but poor execution or missing outcomes can still weaken it.
Use these questions in order. They work for journal articles, news reports, essays, and class discussions. Write one sentence of evidence under each question before deciding how strong the claim is.
Austin Bradford Hill proposed nine causal viewpoints in 1965, including temporality, consistency, biological gradient, plausibility, and experiment. A modern peer-reviewed review links that framework with directed acyclic graphs, sufficient-component cause models, and GRADE. Use the viewpoints to build a reasoned case, not as boxes that mechanically prove causation.
Consider a fictional study of 200 students: every additional five weekly study hours is associated with an exam score four percentage points higher. These hypothetical figures are for interpretation practice, not a real finding.
The defensible first sentence is: “In this sample, students who reported more weekly study time tended to have higher exam scores.” This describes an association. It does not yet show that assigning any student five extra study hours would raise that student’s score by four points.
If the study measured time and score once, used self-reports, and adjusted only for age, the causal claim remains weak. A stronger design might measure baseline achievement, record study behavior prospectively, compare similar students, and explain missing data. A randomized study of a supported study-time intervention could strengthen causal inference, although it would test the full intervention—not “hours” in isolation—and would still require attention to adherence and outcome measurement.
A careful final sentence is: “More study time was associated with higher scores, but the observational design cannot exclude differences in prior achievement, motivation, course context, or measurement error.” That wording is informative without pretending the evidence answers a question it was not designed to settle.
Match your verb to the design and the authors’ actual analysis. This protects accuracy and makes your academic writing more credible.
Do not upgrade “associated” to “caused” in your literature review. Also separate practical importance from statistical evidence: a precise effect may be too small to matter, while an imprecise estimate may leave several important effects compatible with the data.
Copy these prompts into your notes and answer them in one or two lines per paper. In Snitchnotes, you can turn the completed prompts into flashcards or practice questions so that evaluating evidence becomes an active skill rather than a definition you memorize.
No. A strong correlation shows that two variables move together closely under the chosen measure, not why they move together. Confounding, reverse causation, selection, measurement error, or a shared trend can produce a strong association. Causal interpretation requires a credible comparison and a design that addresses plausible alternatives.
Sometimes. Natural experiments, interrupted time series, instrumental variables, regression discontinuity designs, and carefully analyzed longitudinal data can support causal inference when their assumptions are credible. The key is not the label of the method but whether the design approximates the relevant counterfactual comparison and makes rival explanations unlikely.
No. Regression adjustment can reduce confounding from appropriately measured variables, but it cannot automatically fix unmeasured confounding, poor measurement, selection bias, or reverse causation. Adjusting for a variable affected by the exposure can even introduce bias. Ask why each covariate was chosen and how the causal structure was justified.
Association is the broader idea that the distribution of one variable differs with another. Correlation usually refers to a numerical measure of how variables vary together, often the Pearson coefficient for a linear relationship. Both describe patterns in data. Neither, by itself, demonstrates that changing one variable would change the other.
A cause must occur before its effect, so correct time order rules out some reverse-causation stories. However, an earlier variable can still be only a marker for a third factor that produces the later outcome. Temporality clears one hurdle; it does not eliminate confounding, bias, chance, or an incorrect causal mechanism.
To distinguish correlation from causation in research, treat an association as the beginning of the analysis. Define the causal claim, establish time order, identify common causes, inspect how the comparison was created, and test whether the result withstands plausible alternatives. Strong causal conclusions come from design plus reasoning, not from an impressive coefficient or a single statistical threshold.
Use the five-question audit on the next paper you read, then keep the answers beside your summary. If you use Snitchnotes, convert the audit into practice questions and rehearse the judgment process. The goal is not to dismiss observational evidence; it is to say exactly what the evidence can—and cannot—support.
Notes, quiz, podcasts, flashcards et chat — en un seul upload.
Essaie ta première note gratuitement