How Bioma Learn verifies every claim
Last updated: 21 September 2026
Bioma Learn is a nutrition knowledge base built by an editorial system of specialised AI agents, each with one job and a hard boundary. Nothing here is written from memory. Every claim is tied to a primary source, and wherever the source can be stored locally, a machine re-reads the quotation on every build and stops the build if it does not match. This page describes how that works, including the parts that do not work yet.
1. Where our numbers come from
We keep a register of every source we use, with its publisher, version, licence and the date we downloaded it. Today it holds 31 entries. The ones that do most of the work:
| Question | Source |
|---|---|
| What is in a food | USDA FoodData Central (Foundation Foods, SR Legacy), with CoFID and Ciqual as independent second opinions |
| How much a number varies between samples | USDA Foundation sample data: minimum, maximum, median and how many samples were analysed |
| What cooking changes | USDA Table of Nutrient Retention Factors |
| What a real portion is | USDA FNDDS, and NHANES for what people actually eat |
| Whether something works | A local copy of PubMed: 468,455 records, including 305,914 observational studies, plus full texts from PubMed Central where the licence allows |
| Where a recommended intake comes from | Dietary Reference Intake tables, systematic reviews and guidelines from AHRQ, USDA NESR, WHO, CDC and VA, held locally as 15,001 full-text sections |
| Where a legal limit comes from | FDA Daily Values from the Code of Federal Regulations and the Federal Register reasoning behind them |
| How many people are actually short | NHANES blood biomarkers, weighted |
| What is known during breastfeeding | LactMed |
| What people report after taking supplements | FDA CAERS |
| What is actually sold | NIH Dietary Supplement Label Database |
| Which health claims regulators rejected | FDA 21 CFR 101 and the EU health claims register |
Never cited: blogs, commercial sites, aggregators, press releases, popular health portals, science news sites, studies without a DOI or PMID. Wikipedia may help us find a primary source and is never itself the source.
2. Two different questions: how reliable, and how strong
A fact sheet from a government agency is a reliable source carrying weak evidence. A Cochrane review is a reliable source carrying strong evidence. Confusing the two is the cheapest way to mislead a reader without stating a single falsehood, so we track them separately.
How reliable the source is. T1 is peer-reviewed research with a PMID or DOI, and Cochrane. T2 is a government or intergovernmental body: it may set a limit or a position, but a number about an effect needs a T1 card as well. T3 is a professional association or university, allowed only as context. Anything else is not a source.
How strong the evidence is. Levels A to E, shown on the page next to the claim:
| Level | What it means |
|---|---|
| A | Meta-analysis or systematic review of randomised trials |
| B | One or more randomised controlled trials |
| C | Observational studies: cohorts, case-control, cross-sectional |
| D | Mechanistic, animal or laboratory work |
| E | The position of an expert body, without its own analysis |
We do not write that something is proven unless there is level A evidence, and we do not state an effect without saying how large it is, in whom, and over what period.
3. The unit of content is an evidence card
A sentence that makes a claim does not go on a page as prose. It goes in as a card: one statement, the population it applies to, the size and direction of the effect, the study design, the evidence level, the source with its DOI or PMID, and the quotation from that source, word for word. A writer may only use cards that have passed verification. That is a limit built into the tools, not a rule someone has to remember.
We currently hold 816 cards: 271 at level A, 158 at B, 149 at C, 10 at D, 228 at E.
4. What it means to say we have researched something
Getting one fact right is not the same as covering a subject. Every subject we research is taken through nine lines of enquiry: how much is in the food and how much that varies, what the body absorbs, what cooking and storage change, what a real portion looks like, what it does to health, safety limits and interactions with medicines, beliefs the data does not support, who it works differently for, and what science does not yet know. That last one is recorded even when the answer is that nobody has studied it, because an honest gap is a finding and almost nobody publishes it.
A subject does not leave research until it has enough facts behind it: at least twelve cards for a nutrient, four of them at level A or B, covering at least six of the nine lines, and always including health effects and safety. For a food the floor is six cards, two at A or B, covering at least three lines. An individual page is not published if it cites fewer than three cards, and at least one of them has to be stronger than an expert opinion.
Pages built only on measured composition are the exception. A page that reports how much potassium is in a banana does not make a claim that needs a trial behind it. Its obligations are different: every number carries its source, the dataset version and the record identifier, and where a number misleads without context, the context is attached automatically. Iron from plants is the clearest case: 210 pages that quote an iron figure also carry the note that plant iron is absorbed poorly, and the build fails if one of them does not.
5. How a quotation is checked, and what we cannot check yet
When a source can be stored locally, verification is not a promise but a test. The quotation on the card is compared with the text in our copy of the source on every build, after normalising the things that break naive comparison: unicode forms, superscripts, every kind of dash and quotation mark. If the quotation is not there, the build fails. This has already caught a card quoting an abstract that did not contain the sentence.
Not every source can be held locally. Some are behind a paywall, some forbid automated collection, some are web pages that change. A quotation taken from a web page is fetched twice, by two independent requests, and the second one exists only to confirm that the sentence is on the page word for word. If it is not confirmed, there is no card.
The honest number. Of our 816 cards, 12 rest on a web page and were checked by eye. A further 245 were made before we required a card to declare where its quotation came from, so a machine cannot re-check them either. That is 31 per cent of the base that we cannot re-verify automatically, against a ceiling of 25 per cent that we set ourselves. We count this on every build, we publish the number here rather than rounding it away, and closing the gap is current work.
6. Retracted papers
A study can be withdrawn after we have cited it. We check every card with a PMID against two independent lists: the retraction flag in PubMed itself, and the Retraction Watch database, which currently holds 72,453 records. PubMed's own flag is slower and less complete, which is why we do not rely on it alone. A retracted paper stops the build rather than quietly staying on the page.
7. Why our numbers sometimes show a range
A single number for the iron in spinach is a convenient fiction. USDA analysed eight samples of mature spinach and found between 0.63 and 1.62 milligrams per 100 grams. Cultivar, soil, season and storage all move it. Where the source gives us the range and the number of samples, we show them, because a reader who sees three different figures on three different sites deserves to know that the disagreement is real rather than a sign that someone is wrong.
8. Who checks what
The work is divided so that nobody checks their own output. A researcher finds facts and turns them into cards but writes no prose. An editor makes the page readable but cannot add a fact. A nutritionist agent reviews the page for harm and writes the section that translates the science into what to do, using only the cards already on the page. A verifier re-opens the sources and either blocks or passes, and does not edit. A compliance step checks the wording against health claim rules. The build itself enforces what can be enforced mechanically.
Pages are reviewed by this editorial system, not by a named person. We do not invent a reviewer with a medical degree. When a human clinician does review our pages, their name and qualifications will appear on the pages they reviewed.
9. When the evidence changes
New reviews are published, agencies update their reference values, and our own source base grows. When a subject was researched against an older version of our sources, it goes back in the queue rather than keeping a tick next to it. Pages record which cards they rest on, so when a card changes, the pages that used it are identified automatically rather than from memory.
10. How we rank foods
Rankings are built on a realistic portion rather than 100 grams, because nobody eats 100 grams of cinnamon. We exclude raw meat and fish, dry grains and legumes before cooking, large portions of pure fats, and formulated products, since they crowd out the foods a ranking is supposed to help with. Where a food has no household measure in the source data, it has a page but does not appear in rankings.
11. What this site is not
Bioma Learn provides information, not medical advice. It does not diagnose, treat or give personal recommendations. If you have a medical condition, are pregnant, or take medication, talk to a qualified professional before changing your diet.
12. Corrections
Every page shows when it was last verified. To report an error, write to learn@humblebee.ltd. If you tell us a claim is wrong, we re-open the source rather than defend the sentence.