Lore

Evaluating Evidence & Supplement Quality

Из Read: Creatine, comprehensive guide

This chapter sets out the yardstick the rest of the book is measured against: how to weight a nutrition or supplement claim by the quality of the evidence behind it rather than by how compelling a single study sounds, and how to check that a product physically contains what its label promises. It covers the evidence hierarchy and Layne Norton's five-point checklist, the gap between mechanism and measured outcome, study designs that can and cannot detect the effect they're testing, the ways results get over-read (subgroups, equivalence framing, duplicate-counted meta-analyses), the speaker's own incentives and certainty language as signals, and the verification tools — COAs, GMP, stability studies, off-shelf testing — that stand between a buyer and documented supplement fraud.

The hierarchy is a starting weight, not a verdict

The framework this chapter runs on comes from Dr. Layne Norton, and it starts with an ordering: meta-analyses and systematic reviews sit above randomized controlled trials, which sit above cohort and epidemiological studies, which sit above animal studies, which sit above case studies. The point of the ordering is not to rank studies for sport. It's to decide how much a given piece of evidence is allowed to move your beliefs before you've seen anything else — a case report is permission to be curious, not permission to conclude.

But the level of a study is only the first filter. Evaluating Nutrition & Fitness Evidence specifies a five-point checklist for what "strongly supported" actually means: a plausible mechanism, agreeing animal data, a dose-response relationship, supporting human RCTs, and supporting epidemiology — together, not any one alone. Most contested nutrition claims fail this not because the evidence is bad but because it's partial: a pathway with no outcome trial, or an epidemiological signal with no dose-response.

Dose-response deserves its own emphasis because it's the cheapest fraud detector available. A genuine causal or toxic effect should scale with exposure. When a low-dose group shows a bigger effect than a high-dose group, that pattern points at confounding rather than causation — the reasoning used in this material to deflate a headline aspartame-cancer finding. Absent dose-response, a striking conclusion deserves scrutiny of how it was reached: which subgroup was selected, how the outcome was reframed. "If you torture the data enough it will confess."

Above any single study sits the consensus of data: conclusions that replicate across many labs, decades, funding sources, and countries. Creatine's roughly 92% expert consensus for muscle building is the example given, and it is a different kind of claim than "a 2026 trial found." Reading forest plots is the granular version of the same instinct — where the dots cluster, how wide the confidence intervals are, how heavily each study is weighted by size, rather than whichever study got the headline.

The most institutionalized form of this discipline in the chapter's material is the NIH-funded Interventions Testing Program, which tests candidate longevity molecules in non-inbred mice, replicated in triplicate across three independent institutions. Only two interventions have ever extended lifespan across all four model-organism categories — yeast, worms, flies, mammals — namely caloric restriction and rapamycin. ITP-confirmed molecules also include acarbose, SGLT2 inhibitors, and 17-alpha estradiol in males only; ITP-failed molecules include nicotinamide riboside, metformin, fisetin, and resveratrol, each of which had a perfectly plausible mechanistic story going in. That list is the cleanest available demonstration that a good mechanism predicts almost nothing, and Longevity Interventions carries what it means for those specific molecules.

Mechanism is a hypothesis; the outcome is the claim

Every genuine outcome has some mechanism behind it. The inverse does not hold: a biochemical pathway existing tells you a thing is possible, not that it happens at a meaningful size in a living person. Norton's line for the failure mode is blunt — "there's nothing more dangerous than somebody who's read a biochemistry book" — and the operational rule it produces is equally blunt: you cannot claim something does X if you don't actually measure X. Before accepting a "does X" claim, look for measured outcome data, not pathway data.

The cleanest case in Evaluating Nutrition & Fitness Evidence is the origin of the "creatine causes hair loss" claim. The founding study ran three weeks and measured DHT — a mechanistic precursor to hair loss, not hair loss. It never measured hair. A later 12-week double-blind RCT measured the actual outcome, hair growth and loss, directly: no group difference, and no difference in DHT, testosterone, or free testosterone either. The reversal is worth noticing, but the deeper point is that the mechanism-only study should never have supported the outcome claim in the first place, in either direction. What the claim itself now looks like is Creatine Fundamentals' business; the methodology is what belongs here.

The same rule has a quieter, more common application: when mechanism and clinic disagree on numbers, the clinic wins. Mechanistic studies on omega-3 fats and muscle attribute the effect to reduced inflammation, an explanation that should be weighed against the broader and decidedly mixed inflammation literature rather than accepted at face value. And the doses used in those mechanistic studies — 2 to 5 grams — should not override clinical dose-response data showing effects from around 1.4 grams when someone asks what to actually take. Mechanistic-study parameters are chosen to make an effect visible in an assay, not to be a recommendation.

What a study design can and cannot detect

A study's design determines the set of answers it is physically capable of returning. Two studies can both be RCTs and differ enormously in whether a null result means anything.

Start with what people report about their own eating. The "good pupil phenomenon" — subjects in free-living nutrition studies over-report healthy behavior and under-report cheating — is, in Alan Aragon's framing, a structural reason nutrition science stays permanently contested. Metabolic-ward studies control for it but are artificial; free-living studies are natural but unreliable. There is no design that is both. This is why controlled feeding trials, where all food is provided and calories and protein are equated, beat self-reported diet studies, and why one-to-one nutrient swap designs — a seed oil substituted gram-for-gram for saturated fat — beat designs that let the intervention smuggle in extra calories. Mendelian randomization is the observational world's attempt at the same trick, using genetic variants as proxies for lifelong exposure to approximate a randomized experiment.

The confound has a signature result attached. Ad libitum ketogenic diets beat controls, but the reason is that they equate to higher protein intake and produce a spontaneous, unmeasured caloric deficit of 400 to 900 calories a day — not a distinct carb-restriction mechanism. When protein and total calories are actually equated, well-controlled trials show no significant difference in fat loss regardless of macronutrient split: "protein and calories are the great equalizer." Diet Strategy & Body Composition is where that conclusion gets used; here it's an illustration of an uncontrolled variable doing the work the headline credits to something else.

Windows that are too short for the thing being measured

Researchers use a washin or washout period to isolate a true physiological effect from a confound — separating real muscle-mass change from water retention, for instance. If that window is shorter than the time the confound takes to resolve, a null result can be manufactured by measuring too early rather than by the intervention failing. A 12-week, 5 g/day creatine study used a 7-day washin for exactly this purpose and then found no muscle-mass advantage over training alone. The "Don't Die" team's critique is that true muscle saturation at 5 g/day takes three to four weeks, so a 7-day window couldn't have separated real gain from residual water weight, making the null a plausible methodology artifact.

Norton reads the same study through a different lens, the baseline-reset fallacy. People say caffeine "stopped working," but more likely their baseline arousal shifted upward and stayed there, so they're comparing against a new elevated floor rather than noticing the effect is still present. Applied to the creatine study: a lean-mass advantage appeared during the one-week wash-in and then showed no further divergence across 12 weeks of training. Read naively that looks like creatine ceasing to matter; read as a baseline reset, the one-time water-driven gain established a new floor and later measurement simply no longer detects a relative difference even though the underlying intracellular-water effect persists. He also flags a population confound — untrained subjects respond so strongly to any resistance training that a smaller ongoing supplement effect can vanish inside the noise. The general heuristic: before accepting any "the effect washed out" study, check whether the design can even distinguish an already-realized, front-loaded gain from an ongoing incremental one.

Dose chosen for sensitivity

Dose selection is a design choice with evidentiary consequences. Norton's creatine HCl versus monohydrate trial tested at a low maintenance dose of about 2.1 g/day, deliberately: HCl's marketing claim runs higher solubility → higher bioavailability → smaller effective dose needed, so any real advantage should show up most clearly at a low dose, while a high dose could produce a ceiling effect that masks a genuine difference. Under that framing, a null at the sensitivity-maximizing dose is far more probative of "no real difference" than the identical null would be at a high dose. Ask of any supplement trial whether the dose was picked to give the claimed effect its best chance, or merely for convenience.

How results get over-read

Even a well-designed study yields a number that can be stretched past what it supports. Evaluating Nutrition & Fitness Evidence collects four specific stretches.

Statistical significance is not clinical significance. A result can be unlikely due to chance and still too small to matter. Be especially suspicious when significance is reached only through a subgroup or subanalysis of an otherwise null primary outcome — the example is a fatty-liver-disease trial of NR plus pterostilbene ('Basis') with no primary-outcome effect, where only a subanalysis restricted to participants below 27% hepatic fat turned significant. Peter Attia's rule of thumb covers the pattern: "if you have to resort to really interesting statistical machinations to see something, there probably isn't something very interesting there."

"No difference detected" is not "causation disproven." A 2026 meta-analysis of more than 13,000 participants found creatine and placebo produced statistically indistinguishable adverse-event-reporting rates, around 13% each. The source video treats this as proof creatine "does not cause" side effects, collapsing an equivalence finding into a causal acquittal, where the careful statement is insufficient evidence of a difference. The tell to watch for is absolute, individual-variability-free framing — "case closed," "no side effects" — layered on pooled-cohort statistics, particularly when the same source discloses ownership of a monohydrate-only supplement company that benefits from the conclusion.

A meta-analysis can be worse than the trials inside it. In a critique of a meta-analysis cited in Dr. Darren Candow's creatine-and-cognition claims, relayed on Rhonda Patrick's podcast, the analyst Physionic found the same trials counted repeatedly — Alvis and McMorris entered the pooled analysis as separate entries up to seven times, because each trial reported multiple measurements of the same outcome. That hugely inflates the apparent sample size and likely manufactured statistical significance in a memory outcome that would disappear once duplicates were removed. A useful corroborating tell: the same analysis showed no effect on processing speed, suggesting the memory "hit" was a duplication artifact rather than a real effect. Note also the discipline Physionic modeled — rejecting the meta-analysis as poor evidence while still endorsing the underlying claim that creatine likely helps cognition in stressed or aging populations, on the strength of other corrected analyses and RCTs. Criticizing a source is not the same act as rejecting a claim.

Verify the citation, not the paraphrase. When a credentialed source cites a specific study for a specific number, the number can still be wrong even when the source is generally careful. Dr. Rhonda Patrick's claim that 10 g of creatine raises brain creatine traced to a Tübingen, Germany study that in fact used 20 g/day. The related trap is raw effect size diverging from significance in small trials: a dose arm can post the largest numeric change — 10 g creatine on phosphocreatine in adolescent girls — without that difference being statistically significant in a small, short, or differently-powered study. Patrick flagged that the underlying evidence came from small studies, which is good practice worth crediting even when a specific number turns out to be off.

The speaker is data too

Two properties of a communicator carry information independent of their argument: how certain they sound, and what they sell.

On certainty: credible experts hedge — "could," "uncertain," "the data suggest." Overconfident communicators reach for absolutes: "always," "never," "best," "worst." A speaker's own certainty language is itself a signal about how well the evidence supports the claim, and it's freely observable before you've read anything. The companion discipline is one Norton states about himself: change your conclusion to fit unexpected data, not the reverse — "I care more about getting the right answer than being right."

On incentives, Evaluating Nutrition & Fitness Evidence is deliberately even-handed and correspondingly unsatisfying. A disclosed commercial stake doesn't by itself invalidate an argument — a supplement formulator who sells the cheaper, already-established option is not thereby wrong about it — but it is worth weighing against how much certainty the framing projects. In the monohydrate case, the persuasive work is being done by the evidence base (one new RCT plus decades of prior monohydrate research), not by the source's authority, and that's the distinction to check. Norton discloses that his line sells creatine monohydrate while defending monohydrate as "king" among creatine forms.

The mirror-image move deserves equal scrutiny. A physician who broadly recommends creatine to patients but avoids it himself states explicitly that he takes no sponsors and sells no related product, framing this as what lets him evaluate the data cleanly. That's a common rhetorical preemption of bias objections, and the chapter's guidance is to name it as a pattern rather than accept it at face value. Independence is a claim like any other. The same applies in Supplement Fraud & Quality Verification: Nootropics Depot's lab tests of turkesterone products are useful data, and Nootropics Depot also sells competing supplements.

A third variant is the gap between what a communicator recommends and what they personally do. Mike Israetel of RP Strength argues that food covers roughly 95% of what's needed to build muscle and get lean, that only a handful of supplement categories are actually effective — creatine, protein powder, vitamins and minerals, maybe stimulants — and that much of the rest of the market is either repackaged basics or has no real main effect. He also discloses using prescription weight-loss medications and growth hormone: "I'm a walking pharmacy and you don't want to end up like me." His personal efficacy bar is a decent field heuristic in its own right — "if you can't tell at all when you're very keenly paying attention, probably not doing shit" — which he uses to write off creatine ethylester and non-stimulant fat burners like berberine, capsaicin, and green tea extract as having real but mechanistically "teenytiny" effects that stay below noticeable even stacked. It is worth flagging that this heuristic sits uneasily with everything in the preceding sections: subjective noticeability is exactly the kind of signal blinded trials exist to replace.

Whether the tub contains what the label says

All of the above assumes the product in your hand matches the substance in the study. That assumption is not safe, and the reason starts with regulation. The lay intuition that "natural supplements" are categorically different from and safer than medicine doesn't hold: supplements aren't regulated by the FDA with the rigor applied to pharmaceuticals, and the "generally regarded as safe" category is not an especially rigorous process. Counterintuitively, this material judges the supplement space to have on the order of 10 to 100 times more nefarious quality-control and marketing behavior than the pharmaceutical space, despite the looser regulatory optics implying the opposite risk ordering. The practical implication is to apply the scrutiny you'd apply to a prescription drug, not less.

Supplement Fraud & Quality Verification supplies the vivid end of the problem, via the testing lab Light Labs: a best-selling Amazon creatine gummy that tested at 0% of its claimed creatine, and a retail protein powder that tested at roughly 3 g of protein against a ~21 g label claim, the gap made up with carbohydrate filler. Misses of that magnitude are treated as deliberate fraud rather than manufacturing error, on the grounds that manufacturers are responsible for testing inbound raw materials. The ecdysteroid category shows the same pattern layered on top of an efficacy problem: turkesterone has zero human studies of its own, the closest evidence being a 2019 Eisenmenger study on a related ecdysteroid — whose pills, on testing, contained 6% of the labeled ecdysterone. Nootropics Depot's 2022 tests found Greg Doucette's "Turk Builder" at 0.15% of its labeled dose (0.7 mg against 5,100 mg claimed), Gorilla Mind's product under 1%, and some products with none detectable at all.

Four tools exist to check a specific product, in ascending order of rigor and descending order of availability:

Risk is not evenly distributed across formats. Plain powders — a scoop of creatine in a bag — are the lowest-risk matrix; gummies, drinks, proprietary blends, and multivitamins are higher. Multivitamins are "hard mode" because many vitamins are chemically incompatible with each other and with the product matrix over an 18-month-plus shelf life. Even a correctly formulated product can degrade in the supply chain through heat and light exposure in a warehouse, truck, or shelf — which is what "keep in a cool, dry, dark area" is actually about. Market structure adds a last layer: low minimum order quantities make launching a brand easy and correlate with higher fraud risk in certain formats, while retail channel functions as a rough quality proxy, since retailers like Whole Foods impose their own testing requirements independent of manufacturer claims and Amazon is described as tightening its requirements partly in response to reputational damage from the creatine gummy case. The stated aim of testing-transparency efforts is to let compliant brands win the category over cheaper fraudulent competitors.

Dr. Eric Helms's late adopter principle is the low-effort default that falls out of all this: even the most effective supplements — creatine, caffeine — have only a small impact, so waiting a year or two on an unproven new product costs almost nothing while the quality-control and efficacy pictures resolve. It's worth being honest about what this material doesn't provide: vivid individual cases and a set of tools, but no base rate for how often products fail testing, and no ranking of third-party certifiers beyond "look for the GMP mark and ask for the COA." The failure examples are memorable; they aren't a denominator.

Turning the yardstick on your own case

The chapter ends somewhere slightly at odds with where it began. Alan Aragon proposes an individual-response evidence hierarchy that places your own tracked, personal n=1 response to a protocol above the standard formal hierarchy of RCTs and meta-analyses when the two conflict — peer-reviewed evidence as the default starting point, but your own measured response overriding the general literature recommendation. Set against the opening section's insistence that a single case study is the weakest tier of evidence, this is a real tension in the material rather than a resolved position, and Proactive Health & Self-Advocacy is where systematic self-tracking gets argued for on its own terms.

A cleaner version of the personal-override case is risk calculus rather than efficacy. A physician who broadly recommends creatine to patients avoids it himself because of a family history of polycystic kidney disease and the complete absence of safety data for creatine in that population — not any positive evidence of harm. Population-level trials in people with healthy kidneys, up to five years, at 5 to 30 g/day, show no measured harm to kidney function, but that evidence base simply never tested his subgroup. This is absence of evidence rather than evidence of absence, and it shows how a rational actor with unusual risk factors can land somewhere different from the population analysis without contradicting it. Creatine Fundamentals carries the safety literature itself.

The n=1 license has a hard limit, though, and it's the reason the folklore section exists. A protocol "working" — getting someone lean for a show — does not validate the specific rituals inside it. Fasted cardio burns more fat during the session and less over the following 24 hours, a net wash at best, with the latest systematic review having "pretty much written this off as a myth"; cardio timing should be chosen for schedule and energy, not fat-oxidation folklore. Tilapia's supposed skin-thinning or drying-out property is bodybuilding folklore with nothing behind it — any lean protein does the same job, and the choice only matters for things like fat-source overlap. Whole flax seeds are genuinely nutritious but have no fat-burning mechanism, and the shell isn't digestible unground. Carbs before bed being "instantly stored as fat" is called a complete myth; total daily intake, not timing, determines storage. Rigid food-specific rules tend to function as superstition substituting for actual calorie and macro awareness, and the same critique in Evaluating Nutrition & Fitness Evidence extends to training dogma such as "muscle confusion" — see Resistance Training & Performance Supplements for what does drive the training stimulus.

Last, a framing borrowed from Thomas Sowell that keeps the whole apparatus honest: there are no pure solutions, only trade-offs. When evaluating any protocol, name what is given up, not just what is gained. A chapter of filters like this one is easy to turn into pure negation — every claim has a flaw, so believe nothing. That isn't the intent. The intent is to be able to say how confident to be and why, and then, as the discipline goes, to change the conclusion when the data comes back unexpected rather than defending the position you arrived with.

Открытые вопросы

Концепты

Источники