What Can a Study Actually Tell You?

“Clinically proven” is doing a great deal of work on a lot of packaging. This article is not about which studies are good. It is about the four questions that decide whether a study says anything at all about the product in your hand.

The short answer

A study’s conclusion belongs to the exact thing that was tested, applied the exact way it was applied, measured against whatever it was compared to. Change any one of those and the finding does not travel with it. Most citation problems in this industry are not bad studies. They are good studies attached to the wrong product.

Every article in this series links its sources, and I have been asking you to check my work. That is not much use without a way to check. So here is the method I use, in the order I use it.

Question one: what was actually tested?

Not the ingredient in general. The specific material, in a specific state.

In vitro

Cells in a dish, or a reconstructed skin model. Useful for mechanism.

Cannot tell you: what happens on a person. There is no barrier to cross, no immune system, no microbial community already living there.

Ex vivo

Real human skin, usually surgical tissue, kept alive outside the body.

Closer, and genuinely useful for penetration questions. Still not a living person over weeks.

In vivo

Human volunteers, measured over time.

The one that supports a visible claim. Also the slowest and most expensive, which is why it is the rarest.

None of these is fake science. The problem is only when an in vitro finding is written up on a box as though it happened to a face.

Question two: was it the same thing that is in the product?

This is where most of the failures happen, and it is invisible unless you open the paper.

For microorganisms, a study is run on one specific isolate and its findings belong to that isolate. A different strain of the same species is a different organism, and the effect may or may not carry across. Some properties do sit at species level, which I have written about before, but a specific demonstrated effect is not one of them.1

The same applies to state. A study on a heat-killed preparation does not support a claim about live organisms, and the reverse is equally true. They are different materials.

Question three: was it applied the same way?

Route is not a detail. A capsule swallowed daily for twelve weeks and a cream rubbed on a forearm are two entirely different experiments, even with the same organism inside.

This one is worth watching closely in microbiome skincare, because a large share of the literature is oral. Those studies are real and often well conducted. They simply do not tell you what happens when something is applied to skin.

A worked example

There is a well-known set of studies on Streptococcus thermophilus and skin ceramide levels, going back to 1999 and repeated in later work.2,3 Real findings, published in real journals.

Now apply the three questions. The preparations were sonicated, meaning the cells were deliberately broken open, so this is not evidence about live organisms. The strain is not identified as any strain sold in current products. And one of the studies was conducted in patients with a diagnosed skin condition, which is a different population and a different kind of claim entirely.

Three mismatches, and they compound. A product containing live S. thermophilus citing this work would be borrowing credibility that the papers do not extend to it.

I picked that example on purpose

One of my own products contains live Streptococcus thermophilus. So those papers are not a hypothetical for me. They are the most convenient citation available in my own category, and using them would be easy.

I do not cite them, and the three questions are the reason. The published work used sonicated preparations, so it is not evidence about the live organism I sell. It does not identify the strain I use. And the most quoted of those studies was run in patients with a diagnosed condition, which is a population and a claim category I have no business borrowing from.

What I say instead

That the strain is named, that there is a stated count at a stated point in time, and why the organism was selected: its documented safety history, its compatibility with my formulation, and its ability to stay viable through manufacturing, storage and activation.

That is a smaller claim than “clinically shown to increase ceramides”. It is also one I can support, which is the only quality a claim really needs.

If you find those ceramide papers cited on a live probiotic product, including mine, that is worth an email to the company.

Question four: compared to what?

A skincare study without a control is a study of moisturizer.

Almost any cream applied twice a day for eight weeks improves how skin looks and feels. That is the base rate. To show that a particular ingredient did anything, the comparison has to be the identical formula with that ingredient left out. In cosmetics this is called a vehicle control, and it is the difference between testing an ingredient and testing a lotion.

Related, and easy to miss: what was measured? Instrumental measures such as transepidermal water loss or corneometry produce a number from a device. A questionnaire produces an opinion. Both are legitimate, and a study reporting that most participants agreed their skin felt smoother has measured agreement, not smoothness.

How many people were actually in it?

This is the one I would put next to the four questions if I were allowed a fifth.

Cosmetic efficacy studies are small. Around thirty participants is a common size, and thirty per group is widely treated as adequate for a primary endpoint. That is not a scandal. Cosmetics are lower risk than medicines and the claims permitted are correspondingly modest, so the evidence bar is set lower on purpose. It is legal, it is accepted, and it is often perfectly reasonable.

It is also worth understanding what thirty people can and cannot support.

A small study can only find a large effect. Sample size and detectable effect size are linked. Thirty people can reveal something dramatic. Detecting a small change reliably takes far more participants, so a modest effect can be entirely real and entirely invisible at that scale.

An average is not a promise. A significant mean improvement across a group is compatible with a good number of individuals experiencing nothing at all. The published figure is the average of the responders and the non-responders together.

Those thirty are not the public. Participants are screened, often excluded for skin conditions or medications, supervised, and applying the product correctly on a schedule. That is a filtered and unusually compliant group.

Significant does not mean noticeable. A measurable, statistically real change can still be too small for anyone to see in a mirror. There is an active argument inside dermatology about shifting emphasis from statistical significance toward clinical relevance for exactly this reason.4

“87% of women agreed” is a different thing entirely

Percentages like that usually come from a consumer perception study, which asks a panel to rate statements on an agreement scale after using a product.5 It measures satisfaction and preference. It is a survey.

That is a legitimate thing to run and a legitimate thing to report, as long as it is reported as what it is. Eighty-seven percent agreed their skin felt smoother is a fact about agreement. It is not a measurement of smoothness, and there is frequently no control group at all.

Here is the asymmetry that bothers me. Thirty people, once, for twelve weeks. Then that result goes on the box, into every advertisement, and onto every retailer page, for years, in front of millions. The evidence does not scale with the reach of the claim, and nothing requires it to.

So when you see a study cited, the useful instinct is not “is this fake”. It almost never is. It is “how much weight is this being asked to carry”, and the honest answer is usually a great deal more than it was built for.

The smaller things, quickly

✓  How long? Skin turnover takes weeks, which is why cosmetic studies commonly run twelve weeks or more. A two-week result about renewal is measuring something else.

✓  Who knew what? If participants and assessors both knew who got the active product, expectation is part of the result.

✓  Who paid? Industry funding does not make a study wrong. It is a reason to read the methods rather than the abstract.

How a sentence like that gets written

Earlier I said I do not cite the ceramide papers on my S. thermophilus product, because they fail three of the four questions. I want to finish by admitting the other half of that story.

On my other product, a comparable sentence did get through. It described a strain as well researched for effects on the appearance of skin. When I finally read the underlying work properly, applying these same questions, it belonged to other organisms and in some cases a different route of administration entirely. I removed it.

Here is the part worth noticing. That sentence was not written by someone being careless or dishonest. It was written from a general impression of a field rather than from the papers, which is how almost every inaccurate claim in this industry gets made. The four questions are not a test of whether a company is honest. They are a test of whether anyone actually opened the papers, and the answer is often no, including when the person writing is the scientist who owns the company.

The most common failure in this industry is not fabricated evidence. It is real evidence pointed at the wrong product.

✗  The myth

“There is a published study behind it, so the claim is supported.”

✓  The fact

A study supports a claim only if the thing tested, the way it was applied, and the thing it was compared against all match the product making the claim.

The four questions, in order

1.  What was tested? Cells in a dish, tissue, or people.

2.  Was it the same material? Same strain, same state, alive or not.

3.  Same route? Swallowed is not applied.

4.  Compared to what? Without a vehicle control, you are reading a study of moisturizer.

Key takeaways

✓  A finding belongs to the exact material tested, applied the exact way, against whatever it was compared to.

✓  In vitro is for mechanism. Only in vivo supports a claim about what a person will see.

✓  A different strain, or a killed version of the same organism, is a different material.

✓  Oral studies do not support topical claims.

✓  Without a vehicle control, the study measured the base formula.

✓  Most citation problems are good studies attached to the wrong product, not invented data.

✓  Around thirty participants is normal for a cosmetic study. That can only detect a large effect, and an average result hides everyone who saw nothing.

✓  A percentage of people who “agreed” comes from a satisfaction survey, not a measurement.

References

Links go to the peer-reviewed record. You are welcome to check my work.

1. Sanders ME, Benson A, Lebeer S, Merenstein DJ, Klaenhammer TR. Shared mechanisms among probiotic taxa: implications for general probiotic claims. Curr Opin Biotechnol. 2018;49:207–216. PubMed 29128720

2. Di Marzio L, Cinque B, De Simone C, Cifone MG. Effect of the lactic acid bacterium Streptococcus thermophilus on ceramide levels in human keratinocytes in vitro and stratum corneum in vivo. J Invest Dermatol. 1999. PubMed 10417626

3. Di Marzio L, et al. Effect of the lactic acid bacterium Streptococcus thermophilus on stratum corneum ceramide levels and signs and symptoms of atopic dermatitis patients. Exp Dermatol. 2003. PubMed 14705802

4. Pathania YS. Shifting of clinical researches from statistical significance to clinical relevance. J Cosmet Dermatol. 2025. doi:10.1111/jocd.70456

5. Consumer panel size in sensory cosmetic product evaluation: a pilot study from a statistical point of view. Notes that guidance on adequate consumer panel size in the cosmetic sector is rare, and that the minimum depends on the assessment item and the statistical parameters chosen. scirp.org

6. Rezaei F, Rollin JA. A microbial formulation perspective on probiotic skincare: viability, challenges, and current approaches to maintain probiotic viability. Biotechnol Bioeng. 2026. doi:10.1002/bit.70268

Still have a question?

Submit it here, and it may become the next Ask the Scientist article.

Ask your question

About the author

Dr. Farzaneh Rezaei, PhD, Microbial Formulation Scientist

Founder of Fafabiotic, with two decades developing microbial technologies. She writes Ask the Scientist to explain the science behind probiotic skincare without the marketing layer.

More about me

Educational content only. This article is not medical advice and is not intended to diagnose, treat, cure or prevent any disease.

Share