Methods explainedAugust 2026
How we choose which studies to replicate
When you join a Meta-Student “replication” type project you are handed a specific published study to repeat. It wasn't picked because a supervisor happened to like it. It was picked by a system we built called RRIS — the Replication Research Identification System — which searched 7,575 studies and ranked every one of them by how much the field would learn if somebody actually checked it. Here is how that ranking works.
The problem: there are far too many studies to check.Replication means running a study again, independently, to see whether the original result holds up. It is how science checks itself. The difficulty is that sport and exercise science publishes thousands of papers a year and nobody could ever re-run them all. Replicate at random and you mostly end up re-testing things nobody was relying on in the first place. So the real question isn't “is replication good?” — it is “which studies, out of thousands, are worth the effort?”
The idea: influence × uncertainty. A study is worth replicating when two things are true at the same time.
• People are relying on it. If a paper is widely cited — referenced by other researchers in their own work — then its conclusion has spread. Coaches may have changed what they do because of it. If it is wrong, the error has travelled with it.
• The evidence underneath is thin. The most common reason to doubt a result is a small sample size — the number of participants, written N. A study with 8 participants gives a far less stable estimate than one with 200. Small studies aren't bad science; they are simply uncertain, and that uncertainty rarely survives into the way the finding gets talked about later.
A study scoring high on both — lots of people relying on it, not much data underneath it — is the most efficient thing you can possibly replicate. That is what RRIS goes looking for.
The score. Every study gets a Replication Value:
RV = (citations per year) × 1 / ln(N + 1)
In words: how fast the study is being cited, divided by how much data it rests on. Two deliberate choices in that formula change which studies rise to the top:
Citations per year, not total citations. A 2018 paper has had eight years to collect citations; a 2024 paper has had two. Ranking on the raw total would simply reward being old. Dividing by the years since publication measures how quickly a paper is being taken up, which is what we actually care about.
The logarithm of sample size, not 1 ÷ N. This one is less obvious. If you divide by Ndirectly, a study with 5 participants scores twice as “uncertain” as one with 10 and four times one with 20 — while above about N = 30 the score barely moves at all, so every large study ends up looking identical. Using ln(N + 1) reflects how uncertainty actually behaves: the jump from 5 to 10 participants matters enormously, the jump from 100 to 105 matters very little, but a study of 200 is still meaningfully more solid than one of 40.
A worked example. The current top-ranked study is a 2020 paper on the physiological effects of training in face masks: 245 citations since 2020, which is about 35 citations per year, from 16 participants. That gives RV = 35.0 ÷ ln(17) = 35.0 ÷ 2.83 = 12.4. For comparison, a paper collecting 5 citations a year from 40 participants scores 5 ÷ ln(41) = 1.3. The mask study is roughly nine times the replication priority — a widely-used finding resting on sixteen people.
Where the citation counts come from. Different databases index different sets of citing papers, so they disagree with each other. Every shortlisted study is therefore checked against three independent sources — OpenAlex, Crossref and Semantic Scholar — and the highest count is used, since each database catches citations the others miss. Google Scholar usually reports a higher number again, but it offers no public interface for automated checking, so it cannot be part of a reproducible ranking, and it also often includes non-peer-reviewed citations.
The second filter: could a student actually do it? A high score means a study is worth replicating. It does not mean you can replicate it in a semester with the equipment in your department. So every candidate then passes a feasibility screen, and most are removed. To survive, a study needs:
• 40 participants or fewer — achievable across a student cohort
• Any training block within four weeks — has to fit inside a teaching term
• No venous blood draws or muscle biopsies— these need clinical staff and ethics approval you won't have (finger-prick sampling is fine)
• No MRI, DXA or CT scanning — equipment most departments cannot book
• Whole-body, applied outcomes — not molecular or mechanistic laboratory work
• Healthy young or trained participants — not clinical patients, children or older adults, which are entirely different ethics propositions
• A design the platform can analyse — it has to map onto one of the Meta-Student study types
What this means for your project. You are not doing a practice exercise. You have been handed a finding that the field is actively citing, whose original evidence base is a handful of participants. Your data then joins the same pooled analysis as every other site running that study. Twelve students in twelve labs, each testing fifteen participants, produce something none of them could produce alone: an estimate with enough participants behind it to actually power the question. That is the point of the whole platform. The RRIS method simply makes sure the question is one worth settling when we use replications as projects.
Methods explainedAugust 2026
How we decide whether a study replicates
When a group of student labs re-runs an earlier study, the obvious question is: did it replicate? That turns out to be a much harder question than it sounds, and answering it with a single yes or no throws away most of what your data can tell us. Here is how Meta-Student does it instead.
Your data is never mixed in with the original study. The original result stays separate, as a benchmark to compare against. That matters because published effects tend to be exaggerated — striking findings are the ones that get published, so the first estimate of an effect is usually too big. Our recent review of replications in sports and exercise science found that effects shrank by around 75% on average when studies were repeated. If we simply averaged the original in with your data, that exaggeration would contaminate the answer.
We also don't rely on “was it significant again?”Whether a replication scrapes under p < .05 tells you very little about whether the effect is the same size, and it depends heavily on how many participants happened to be recruited. So instead we report several things, each answering a different question:
1. What is the combined answer? The pooled effect from every contributing lab, together with a range showing how much labs differ from one another. Because it draws on many labs at once, this estimate is far more precise than any single study could manage — including the original.
2. Was the original study even big enough to see what it claimed?A small study can only reliably detect very large effects — a bit like using a telescope that can just about show the moon, then reporting that you saw Pluto. This check (called “small telescopes”) asks whether the true effect is smaller than the original study could realistically have detected. Usefully, it needs nothing more than the original study's sample size, so it still works when the original paper reported very little.
3. Is there genuinely nothing here?Failing to find an effect is not the same as showing there isn't one. An equivalence test does the second job properly: it asks whether the effect is small enough to be practically meaningless. Passing it is a real finding i.e- “we looked carefully, and there is nothing here worth caring about” . This is one of the most valuable results a replication can produce, and it is very hard to achieve without pooling many labs.
4. Does it agree with the original? Where the original reported enough detail, we check whether the new estimate matches it, and whether it lands inside the range the original study would have predicted.
The verdict is decided in advance.Which of these counts as the main criterion is chosen and locked before a single participant is recruited, so nobody can move the goalposts once the results are in. And if the original paper didn't report enough detail for one of the checks, we say so plainly, the fact that a published study can't be properly checked is itself worth knowing.
The upshot for you: your dataset isn't just another data point. It feeds into a pre-registered, multi-criterion evaluation of whether a published finding holds up — the kind of evaluation most individual studies are far too small to attempt. The full technical detail lives in the project whitesheet.
Methods explainedJune 2026
What actually happens to your data
You collect data from a dozen or so participants, upload a spreadsheet, and some time later a result appears. This post explains everything that happens in between, no statistics background assumed.
Most meta-analyses are built from the numbers printed in published papers: a mean, a standard deviation, a sample size. We do something better. Because you upload participant-level data, the analysis sees every individual row from every lab. This is called individual participant datameta-analysis, and it is considered the gold standard: it lets us check the data properly, treat every lab's data the same way, and model the results far more accurately than summary numbers allow.
2. All the data goes into one model — it isn't just averaged. The analysis uses a mixed model, which keeps track of which participant came from which lab. That matters because there are two quite different sources of variation going on at once: people differ from each other within a lab, and labs differ from one another. Think of comparing exam results across schools — you'd want to know both how much pupils vary inside a school and how much the schools themselves differ. Ignoring the second one makes results look more certain than they really are.
3. Everyone gets put on the same scale.To combine them, each study's result is converted into a standardised effect size— how big the difference is compared with the natural spread of the measurements. It answers “how big is this effect, in units of how much people naturally vary?”, which is comparable across labs and equipment. For studies looking at relationships between two variables rather than differences between groups, we use a rank-based measure (Kendall's τ) that is robust to odd distributions and outliers.
4. We measure how much labs disagree — and report it honestly. If every lab finds much the same thing, the pooled result is a solid summary. If labs disagree substantially, that is a genuine scientific finding. So alongside the headline effect we report a prediction interval: the range a brand-new lab running the same protocol could reasonably expect to find. It is usually wider than the confidence interval, because the confidence interval describes the average effect while the prediction interval describes what happens in a single new setting. If a result is going to be useful in the real world, the prediction interval is often the number that matters most.
5. The analysis re-runs every time someone new submits. A normal meta-analysis is a snapshot: someone gathers the literature, runs the numbers once, publishes. Meta-Student is a livinganalysis. Each accepted dataset triggers a full re-run, so the result grows more precise as contributions arrive, and you can watch it evolve on the study's results page.
6. Each study knows when to stop. Before a study opens, we set the smallest effect size of interest — the smallest result that would actually matter in practice. Everything else follows from that. After at least four datasets have arrived, every re-run checks whether the study has reached a conclusion, and there are three ways it can finish:
• An effect is there. The estimate is precise enough and clearly different from zero.
• There is no meaningful effect.The estimate is precise enough, and the whole plausible range of the true effect is smaller than the smallest effect worth caring about. This is a positive result in its own right, not a failure. Note that this is not the same as claiming “non significant p”, which does not test if there is no meaningful effect.
• We won't get there. Even if data kept arriving, the study realistically could not reach either of the above. Rather than let it drift on indefinitely, it closes and says so. Either we cannot get the number of studies and samplre sizes we need, or the effect we are trying to measure is just too small.
To avoid being fooled by an early lucky run, nothing can trigger before four datasets, and the same conclusion has to appear on two consecutive re-runs before the study actually closes.
7. Some questions simply can't be answered at our estimated effect size — and we say so before you start.This is the least glamorous and possibly most important feature. A two-group comparison (one group does the intervention, another doesn't) is statistically expensive: with roughly ten student studies of 10–15 participants each, it can only reliably detect fairly large effects, around 0.7 standard deviations or more. Designs where the same people are measured twice — pre-post and crossover — are far more efficient, and can chase effects roughly half that size with exactly the same number of participants, because each person acts as their own comparison. So when a study is created, the platform checks the design, the target effect size and the expected number of contributing labs, and refuses combinations that could never reach a conclusion. Far better to discover that in ten seconds at the planning stage than after eighteen months of data collection. Thats why we may ask for much larger individual sample sizes for some studies such as RCT.
8. Every result comes with its workings.You'll see a forest plot (each lab's result as its own line, with the pooled estimate underneath), a funnel plot and formal tests for publication bias, checks for unusual or overly influential datasets, and model diagnostics. Nothing is a black box — the full statistical specification is public in the project whitesheet.
9. Where we're uncertain, we say so.If there are too few studies to estimate how much labs differ, we report “not estimable” rather than printing a reassuring zero. And because studies that stop early tend to give slightly optimistic effect estimates, results from closed studies carry that warning on the page. Being straightforward about the limits of a result is part of the point of this project.
The short version: your spreadsheet joins everyone else's in a single properly specified model, that model re-runs the moment new data lands, and the study closes only when it has genuinely earned a conclusion — one way or the other.
Platform updateMarch 2026
Meta-Student launches in prototype
We are inching closer to a release - stay tuned. The study pipeline (study creation, student registration, data submission and meta-analysis) are all stable and under review. We are now planning our intial studies, and preparing to engage in a large outreach exercise to raise awareness. YOU CAN HELP! Please locate the social media buttons at the bottom of the page and like, share, comment, or otherwise share the project with any interested parties.