Reading Research Critically: A Practical Checklist for Evaluating Whether a Study Can Be Trusted
A headline announces that a new study has identified a food additive linked to anxiety, or that a single gene predicts academic performance, or that a common medication doubles the risk of a rare condition. The story is shared widely. Readers absorb the finding. And then, months or years later, a quieter story appears: the original study could not be replicated. The sample was too small. The statistical method was poorly suited to the question. The funding came from a source with a direct financial interest in the outcome.
Scientific literacy is not the same as scientific training. You do not need a graduate degree to read research critically. What you need is a structured set of questions—a framework for examining what a study actually claims, how it was conducted, and whether the confidence placed in its conclusions is justified by the evidence presented.
The checklist that follows is designed for exactly that purpose. It is organized around six areas of evaluation, each of which addresses a distinct and common source of unreliable findings.
1. Sample Size and Population
The first question to ask of any study is deceptively simple: how many people—or organisms, or data points—were involved, and who were they?
Small samples are among the most reliable predictors of non-replicable findings. A study of twelve participants that reports a statistically significant effect is far less convincing than one reporting the same effect in twelve hundred participants. This is not merely intuitive; it reflects a mathematical reality about statistical power. Small samples are prone to producing dramatic-looking results by chance, particularly when researchers test multiple outcomes and report only the ones that reach significance—a practice known as p-hacking or outcome switching.
Red flags to watch for:
- Fewer than fifty participants in a study claiming broad generalizability
- A sample drawn entirely from one university's undergraduate population (the so-called WEIRD problem: Western, Educated, Industrialized, Rich, and Democratic)
- No description of how participants were recruited or selected
- Claims of population-wide significance from a sample that is clearly unrepresentative
2. Study Design and Methodology
Not all study designs are equally capable of establishing causation. The hierarchy of evidence in medicine and public health runs roughly from expert opinion at the bottom, through case reports, observational studies, and cohort studies, up to randomized controlled trials (RCTs) and systematic reviews at the top.
An observational study can identify correlations—people who eat a particular food are more likely to develop a particular condition—but it cannot, by itself, establish that the food causes the condition. Confounding variables, which are unmeasured factors that influence both the exposure and the outcome, are the persistent enemy of observational research.
Red flags to watch for:
- A correlation presented as a causal finding without a plausible mechanistic explanation
- No control group, or a control group that differs from the treatment group in important ways
- Self-reported data used as the primary measure for something that could be directly observed or measured
- A cross-sectional study (a snapshot in time) used to draw conclusions about change over time
3. Statistical Reporting
Statistics are the language in which research findings are expressed, and they can be used with varying degrees of transparency and care. Several specific reporting practices should raise immediate questions.
P-values alone are insufficient. A p-value below 0.05 is conventionally treated as the threshold for statistical significance, but this threshold is widely misunderstood and frequently misused. A p-value tells you the probability of observing results at least as extreme as those found, assuming the null hypothesis is true. It does not tell you the probability that the hypothesis is correct, the size of the effect, or whether the finding is clinically meaningful.
Effect sizes matter more than p-values. A study can achieve statistical significance while demonstrating an effect so small as to be practically irrelevant. Look for reported effect sizes—Cohen's d, odds ratios, relative risk—and consider whether the magnitude of the effect is meaningful in context.
Relative risk versus absolute risk. A treatment that reduces the risk of an event from 2% to 1% has reduced relative risk by 50%—but the absolute risk reduction is only 1 percentage point. Headlines almost always report relative risk; the absolute numbers tell a more honest story.
Red flags to watch for:
- No confidence intervals reported alongside point estimates
- A p-value just below 0.05 in a study that tested many variables
- No pre-registration of hypotheses (pre-registration means the researchers publicly committed to their hypotheses and analysis plan before collecting data, reducing the opportunity for after-the-fact adjustment)
- Language like "trending toward significance," which typically means the result did not reach significance
4. Funding Sources and Conflicts of Interest
The relationship between funding source and research outcome is one of the most consistently documented phenomena in the sociology of science. Studies funded by pharmaceutical companies are significantly more likely to report favorable results for the sponsor's product than independently funded studies of the same interventions. The same pattern has been documented in food science, environmental research, and behavioral economics.
This does not mean that industry-funded research is necessarily wrong—it means it warrants additional scrutiny. A well-designed, transparently reported industry-funded study may be more reliable than a poorly designed independent one. Funding is a flag, not a disqualifier.
Red flags to watch for:
- No disclosure of funding source
- Funding from an organization with a direct financial interest in the outcome, combined with no independent replication
- Authors with undisclosed financial relationships to the sponsor
- Research conducted entirely within a company's own laboratories with no external validation
5. Peer Review Status and Journal Quality
Peer review is the mechanism by which the scientific community vets research before publication. It is imperfect—reviewers miss errors, conflicts of interest affect recommendations, and the process is slow—but its presence is still a meaningful signal of basic quality control.
Not all peer-reviewed journals are equal. A study published in Nature or The New England Journal of Medicine has passed a more rigorous review process than one published in a recently launched journal with an unfamiliar editorial board. The Directory of Open Access Journals (DOAJ) and databases like PubMed provide a reasonable filter for journal legitimacy.
Red flags to watch for:
- Publication in a journal not indexed in PubMed, DOAJ, or other reputable indices
- A journal whose editorial board is composed primarily of authors who publish in that same journal
- A preprint (which has not been peer-reviewed) reported in the media as though it were a final, certified finding
- Rapid publication timelines that suggest cursory review
6. Independent Replication
A single study, however well-designed, is a single data point. The gold standard in science is replication: the finding holds up when independent researchers, using independent samples, reproduce the result.
High-profile findings that have not been replicated—or that have been directly contradicted by replication attempts—should be held provisionally. The history of science is populated with celebrated discoveries that did not survive independent scrutiny.
A practical final question: Has this finding been replicated by researchers with no connection to the original team, in a different population, using a different methodology? If the answer is no, treat the finding as preliminary—interesting, perhaps, but not yet established.
Critical reading is not cynicism. It is the appropriate response to a scientific literature that is vast, uneven in quality, and frequently filtered through media incentives that reward drama over accuracy. The checklist above will not make you a scientist. But it will make you a more discerning consumer of scientific claims—which, in an era of information abundance, is a skill of genuine and lasting value.