- Home
- How We Evaluate Claims
How We Evaluate Claims
Every site claims to “follow the science.” Almost none tell you what that means in practice, which makes the phrase worthless. So here is ours, written down, so you can hold us to it.
This page exists for two reasons. First, so you can judge our work by a stated standard rather than by whether you happen to like the conclusion. Second, because we’re asking you to be sceptical of confident claims, and it would be absurd to exempt ourselves from that. That includes being straight with you about how these articles get made, AI and all.
We rank evidence by quality, not by how striking it is
- Systematic reviews and meta-analyses: Many studies pooled and weighed. The closest thing to a verdict.
- Randomized controlled trials: the gold standard for cause and effect in humans.
- Large, long-running observational studies: Useful, but they show association, not causation.
- Small human studies: Suggestive. Easily flukes.
- Animal studies: A hint. Mice are not tiny humans.
- Cell and lab studies: The very beginning of an idea, nothing more.
How We Weight Evidence
One study proves nothing
A single paper is a data point, not a conclusion. We look for replication, and for whether a finding fits the wider body of evidence. When a striking new study contradicts a large existing literature, the prior probability is that the new study is wrong, not that everything before it was.
We report effect sizes in absolute terms
“Doubles your risk” tells you almost nothing without the starting number. If a risk goes from 1 in 10,000 to 2 in 10,000, that is a doubling and also close to irrelevant to your life. We give the absolute numbers wherever we can, because relative risk is the most common way honest data gets used to frighten people.
We follow the money
Funding does not automatically invalidate a study, and pretending otherwise is its own kind of lazy thinking. But industry-funded research reliably skews toward the funder’s interest, so we check who paid, we tell you when it matters, and we weight accordingly. Supplement companies, food industry groups, and pharmaceutical firms alike.
We state our uncertainty rather than hide it
Confidence sells, and hedging doesn’t. We hedge anyway, because the alternative is lying. When the evidence is genuinely unclear, we say “we don’t know” instead of manufacturing a verdict. When something is well established, we say that plainly too, and we don’t invent false balance to seem even-handed.
"Not proven" is not "disproven"
This distinction gets collapsed constantly, including by sceptics. Weak evidence for a claim means nobody has demonstrated it, which is not the same as demonstrating it’s false. We’d rather say “the evidence isn’t there yet” than the more satisfying “this is nonsense.”
We change our minds, visibly
Evidence moves. When it moves against something we’ve published, we update the piece and note what changed and why, rather than quietly editing or leaving it to rot. A visible correction record is the only real proof that a site is following evidence rather than defending a position. Every article carries a last-reviewed date.
How an article gets made?
We use AI in our research and drafting process. We’re telling you that up front, because the alternative is letting you find out and wondering what else we left unsaid.
We also think this is a good use of the technology. Most of the alarm about AI is about what it replaces, and much of it is fair. But used well, as a tool for reading faster, reaching more research, and explaining hard things more clearly, it serves the oldest good cause there is: education. A single person can now work through more of the literature, and turn it into something you can actually use, than was ever possible alone. We think that’s worth doing, and worth doing carefully. The carefulness is the rest of this section.
The order of operations matters enormously, and it’s the thing most AI disclosure statements are vague about. So here is exactly how a Caveat Scientia article is built, in sequence.
A human thesis, written first
Every article begins with a position, written by a person, before any AI touches the work. We identify a claim worth examining, form a view about what the evidence likely shows, and write that thesis down. This is the spine of the piece, and it is ours. This matters because it sets the direction of the inquiry. An article that starts with an AI prompt inherits whatever that model happens to pattern-match. An article that starts with a human thesis has someone accountable for the question being asked.
Authoritative sources, gathered deliberately
We then assemble the evidence base, prioritizing meta-analyses, systematic reviews, and the strongest primary literature we can find on the question, following the hierarchy in Part One. These are the sources the article will actually rest on, and they are selected on the merits, not because they support the opening thesis.
Authoritative sources, gathered deliberately
With the thesis set and the core sources in hand, we use AI tools to go deeper: surfacing additional literature we might have missed, working through dense papers, and synthesizing large volumes of material into something coherent. This is real assistance and we won’t pretend otherwise. It lets one person read further and more carefully than they otherwise could. But it operates inside strict limits:
- AI does not decide what's true. It accelerates the reading. The judgment stays human.
- Every AI-surfaced source is retrieved and read directly. Language models can fabricate citations that look flawless. Nothing enters an article on a model's say-so; if we cite it, we opened it.
- AI never supplies a fact that isn't traceable to a named source. If a claim can't be tracked back to a real study we've read, it does not run.
Every factual claim is verified against the source
Before publication, every factual claim, every statistic, every characterisation of what a study found is checked back against the primary literature by someone with scientific training. That means a master’s-level scientist with a decade in the field reads the paper and confirms that the article says what the evidence actually says. This is the step that makes the rest of it safe. AI can accelerate research; it cannot be trusted to be accurate. So we don’t trust it to be. Verification is the checkpoint every claim must clear, and it is done by a human reading the source.
Human editing, human voice
The final piece is edited by a person for accuracy, tone, and honesty about uncertainty, including the hedges that AI tends to smooth away and that we deliberately put back.
A human decides what to ask and what it means. AI helps read. A scientist checks every fact against the source. Nothing published rests on AI's authority.
What we read
What we're not
Hold us to it
The point of all this is not that we're always right. We won't be. The point is that our reasoning is visible enough that you can catch us when we're wrong, and that we've committed in advance to the standards we'll be judged by. That includes the AI disclosure. We could have said nothing, and almost nobody would have known. We believe AI can genuinely serve education, but that belief only earns your trust if the human checkpoints are real. So we'd rather tell you exactly where the machine helps and exactly where a human has to sign off, and let you judge the work on that basis. If you think we've broken one of these rules in a specific piece, tell us. That email gets read.
Frequently asked questions
Do you use AI to write your articles?
We use AI for research assistance and synthesis, never as the source of truth. Every article starts with a human-written thesis, every source is read directly by a person, and every factual claim is verified against the primary literature before publication.