Eight short lessons that help education partners read and use the statistics that appear in research briefs, program evaluations, and What Works Clearinghouse reviews. Each lesson pairs a plain-language explanation with a small example you can try yourself.
Click any lesson to expand it. Read the short explanation, then move the slider or flip the toggle in the example to see how the statistic behaves. There is no fixed order, start with whatever you meet most often in your work. Use the filter below to jump to a theme.
Partners make better decisions when they can judge evidence on their own rather than relying on a summary. These lessons give district and agency staff a shared vocabulary for reading effect sizes, uncertainty, and study design, which builds partners' lasting capacity to use research. The goal is confident, independent readers of evidence, not one-time answers.
An effect size puts a result on a common ruler so you can compare studies that used different tests. A common one is the standardized mean difference, often written as d. It says how far apart two groups are in standard-deviation units. A d of 0.20 is small, 0.50 is medium, and 0.80 is large, though in education even 0.10 to 0.20 can matter at scale. Move the slider to see what a given effect size looks like and what it means in plain terms.
A p-value answers one narrow question. If the program truly had no effect, how often would we see a result this large just from random chance. A small p-value (by convention below 0.05) means chance is an unlikely explanation, so the result is called statistically significant. A p-value does not tell you the effect is large or important, only that it is probably not zero. Move the slider and watch the label change at the 0.05 line.
Every estimate carries uncertainty. A confidence interval shows the range of values that are consistent with the data, usually at 95 percent confidence. A wide interval means a lot of uncertainty, often from a small sample. A narrow interval means a more precise estimate. A useful habit is to check whether the interval crosses zero. If it does, the study cannot rule out no effect. Increase the sample size and watch the interval tighten.
The way a study forms its comparison group shapes how much you can trust a causal claim. In a randomized controlled trial (RCT), chance decides who gets the program, so the two groups start out similar on average. In a quasi-experimental design, groups form some other way, for example by district choice or timing, and the researcher adjusts for differences statistically. Quasi-experiments are valuable and often more feasible, yet they lean harder on assumptions. Use the toggle to compare.
A forest plot stacks several studies on one chart so you can see the whole body of evidence at a glance. Each row is one study. The dot is its estimate and the horizontal line is its confidence interval. The vertical line marks no effect. Studies whose line crosses it did not detect an effect. The diamond at the bottom is the combined, or pooled, estimate across all studies. Hover or tap a row to read it in words.
These two ideas are easy to mix up. Statistical significance asks whether an effect is probably not zero. Practical significance asks whether the effect is large enough to matter for students, staff time, or budget. With a very large sample, even a tiny effect can be statistically significant while making little real difference. Set the two dials and see which of the four cases you land in.
A regression describes how an outcome changes as one input changes, holding other things fixed. The coefficient is the slope. It says how much the outcome moves for a one-unit increase in the input. In the example, each extra hour of tutoring per week is linked to a change in a test score. Move the coefficient slider to see the fitted line tilt and read the plain-language sentence it produces.
Federal education law (ESSA) sorts evidence into tiers so partners can tell how strong the backing for a program is. The tiers reflect study design and whether the program itself was tested. Higher tiers come from stronger designs. The fourth level, often called "demonstrates a rationale," means a program has a logic model and ongoing study, but has not yet shown effects. Select a tier to see what it requires.
| Tier | Design behind it | What it tells a partner |
|---|