Evidential AI: AI That Shows Its Evidence in Chemistry R&D
Blog11 min readOctober 6, 2026

Evidential AI: AI That Shows Its Evidence in Chemistry R&D

Evidential AI shows what it knows, where it got it, and what it is only assuming. Here is why that matters more than a fluent answer in chemistry R&D.

Back to Resources

Ask a general AI assistant why your batch failed and you will get an answer, usually a confident one. Now ask it which part of that answer came from your data, which part came from a paper, and which part it made up to fill a gap. Most tools can't answer that second question.

In chemistry R&D, the second question matters more than the first. This article is about evidential AI: AI that is built to show its evidence, say when it is unsure, and ask for the information it is missing. We think every AI tool used in a lab should meet this standard.

What you need from an answer

Think about how you judge a claim from a colleague. You look at the conclusion, and then you ask what it rests on. Did they measure it? Did they read it somewhere, and where? Or is it a reasonable guess from experience?

All three can be useful. A good guess from an experienced formulator is worth a lot. But you treat each one differently. You act on the measurement, you check the reference, and you test the guess before you spend a stability slot on it.

An AI answer deserves the same treatment. The problem is that most AI tools mix all three into one paragraph, with no way to tell them apart.


Why generic AI isn't tied to evidence

General-purpose language models are trained to predict the most likely next words. The goal is an answer that reads well and sounds right. Evidence can be part of that answer, but nothing in the design requires it.

There is a second, less obvious issue. A 2025 paper from researchers at OpenAI and Georgia Tech argues that models hallucinate partly because of how they are trained and scored: the usual evaluation methods reward a guess over an honest "I don't know" (Kalai et al., 2025). On a typical benchmark, a guess has some chance of being marked right. Admitting uncertainty scores zero. So models learn to guess.

Chemistry has its own evidence on this. The ChemBench study in Nature Chemistry tested language models on more than 2,700 chemistry questions, and compared them with 19 chemists on a 236-question subset (Mirza et al., 2025). The models did well. The best one outperformed the best human in the study overall. But when the researchers asked the models how confident they were, most gave confidence estimates that did not match how often they were right. On one set of safety questions, GPT-4 gave its lowest confidence (1 on a scale of 1 to 5) to the one question it answered correctly, and a confidence of 4 to the six it got wrong.

So a model can know a lot of chemistry and still be a poor guide to how much of its own answer you should trust. (We covered the accuracy side of this in Why Generic AI Fails at Chemistry.)

People who use these tools for technical work notice the gap. Oran Zucker, who heads a chemical and process engineering division, added a general-purpose assistant to his workflow to save time. It helped, but every technical answer had to be checked, corrected and re-checked before he could use it. He called it supervision work.


Why this hurts more in the lab

In software, a wrong answer is cheap. You run the code, it fails, you fix it, all within minutes. In physical R&D, a wrong answer can cost a cure cycle, an oven study or a scale-up trial. Our CEO, Georgy Maikov, put it this way in his letter Designing for the Slow Loop:

"In a fast loop, a confident guess is cheap, because if it is wrong you find out in a minute. In a slow loop, a confident wrong guess can cost a month."

If you lead R&D and decide which experiments run next, you need to know which parts of an AI's advice are solid enough to act on, and which parts need checking first. As Georgy wrote, "I don't know yet, and here is what would tell us" is part of the job.


What we mean by evidential AI

We use the term for AI whose answers stay tied to evidence and are open about where the evidence runs out. In practice, that means four things:

  1. It starts from evidence. Your own data comes first: protocols, batch records, observations, results. Then published sources. Each claim can be traced back to where it came from.
  2. It shows where each piece of information comes from. Did you report it, was it retrieved from a source, or is the AI inferring it? You should be able to see the difference at a glance.
  3. It states its assumptions and how confident it is. When the evidence doesn't settle the question, it keeps more than one explanation open instead of settling on one too early.
  4. It asks for what is missing. When a key fact isn't there, it asks you for it rather than filling the gap with a guess.

Science has worked this way for a long time. In 1890, the geologist T. C. Chamberlin argued for holding several working hypotheses at once. In 1964, John Platt called the same habit "strong inference." Good scientists already work this way. Evidential AI means the tool works this way too.


What it looks like in a real case

Oran, a REACTOR user, recently wrote up a test he ran on the REACTOR Formulation Troubleshooter. He took two real failures from his master's research on exsolved perovskite catalysts for ammonia synthesis. He already knew how both cases ended, so to keep the test fair he gave the tool only what he had at the time: his protocol, his observations and his data.

Case 1: a gel that wouldn't form. His citrate-nitrate sol-gel synthesis produced a stiff, glassy, tricolored solid instead of a gel. He attached his protocol, and the Troubleshooter worked from it. It came back with four possible causes, ranked by likelihood and labeled low confidence. The real cause was one of the four. Here is what Oran said about the label:

"For me, it shows that the tool doesn't overstate itself, and I feel more comfortable relying on it."

It then asked him three questions, each meant to rule causes in or out. One asked whether his recorded amounts matched the procedure. That sent him back to his notebook, where he found he had never recorded weighing the citric acid. He had left it out. As he put it: "Here, my notebook is what cracked the case, and no AI tool can do that part for you (yet)."

Case 2: a catalyst that wouldn't make ammonia. This one was harder, and Oran ran it twice. The first time he described only the failure. By his account, that run "went completely off track," sending him back to check early synthesis steps. For real work, he wrote, it would have been a waste of time. The second time he added his XRD result, the earlier thesis on the same material, and data from the failed batch showing the material had been made correctly.

"One time I asked REACTOR to think for me. The other time I gave it something to actually reason with."

With that evidence, the tool skipped the checks his data had already ruled out and pointed to the reduction step and nanoparticle formation at the surface from the start. Over several rounds of questions and answers, it asked for a few checks and suggested that the precursor's reducibility had changed, so the reduction step was no longer strong enough. It did not claim certainty. "It kept two possible root causes open and said so," Oran wrote. "The first one turned out to be right."

He also noted how the tool sorted what it knew into three types: You reported (his own data and answers), Retrieved (from a source, such as the thesis he uploaded) and Inferred (the tool's own reasoning).

"'Inferred' is its own reasoning, which is exactly the part to question."

That is the point of evidential AI: it shows you which part to question.


One user, two cases: run the test yourself

Oran's write-up covers one user and two cases, run in hindsight. It shows how the tool behaves. Two cases can't tell you how often it gets the answer right, and we won't claim they do. We have run our own structured tests too, but the most useful test for you is one on your own data.

Oran's method is easy to repeat:

  1. Pick a failure you have already solved. You know the answer, so you can judge the result.
  2. Give the tool only what you had at the time. The protocol, your observations, your data. Leave the answer out.
  3. Check how it reasons, as well as what it concludes. Was the real cause on its list? Did its questions point you toward it? Were the labels honest about what it was inferring?

What you give the tool matters

Oran's two runs of Case 2 show that an evidential tool can only work with the evidence it gets. With a short description, it has to start from the beginning and check everything. With real data, it can rule things out and focus.

You can start with what you already have. A plain description of the formula, what changed between a good batch and a bad one, the conditions, and what you saw is a reasonable first input. In Case 1, Oran attached a single protocol. Your records don't need to be tidy first. Copy in the relevant notes from wherever you keep them. The more you add, the fewer assumptions the tool has to make.

Two habits help most. Keep good records: in Case 1, the deciding clue was in a lab notebook. And question the inferred parts: the labels show you where to push back.


Doesn't that mean more checking?

It's a fair question. Oran Zucker's main complaint about generic AI was all the checking it needed. Asking you to question the inferred parts can sound like the same burden.

The difference is where the checking goes. With a generic answer, every sentence is a possible problem, so you check everything. With labeled information, what you reported is your own data, what was retrieved points to a source you can open, and the inferred items are listed separately. That gives you a short list to review instead of a whole answer to audit. You still check, but you know where to look.

The tool's job is to narrow down what is worth testing and show you why. It also has clear limits: the Formulation Troubleshooter does not run experiments, does not predict property values nobody has measured, and does not replace the bench or your judgment.


How REACTOR is built around this

REACTOR is the AI platform we build at Elementics.AI for chemical and materials R&D. In the Formulation Troubleshooter, each answer comes with a diagnostic card that shows where the investigation stands:

  • Established: the facts you reported.
  • Still assumed: what the tool is inferring, stated as assumptions.
  • To narrow it down: when the cause is still open, a few questions that would rule causes in or out, answerable in one click.
  • Next moves: the checks to run, each tagged as a desk check of your records, bench work or a stability slot, and usually ordered cheapest first.
  • What this cannot tell you: the gaps in the evidence so far.

Across REACTOR, references are numbered and link to the source: a paper, a PubChem record, or a file you uploaded. Reasoning that comes from REACTOR's built-in chemistry knowledge is marked as such, and you can ask for the published precedent behind any suggestion. Ranked causes can carry a confidence label. In Oran's first case, all four were marked low confidence. Each answer is saved in a Research Space, so you keep a record of what was checked and why.

The Formulation Troubleshooter, the tool Oran used, is available on the REACTOR Business plan.


The question to ask any AI tool

When you evaluate an AI tool for your lab, the quality of its answers matters. But ask a second question too: when this tool is wrong, will I be able to tell before I run the experiment?

A tool that shows its evidence, says when it is unsure and asks for what it is missing gives you a fair chance to catch its mistakes. Oran described what he found as "a tool that respected that judgment enough to keep asking for it." In R&D, where every experiment costs time and money, that is the kind of AI worth using.

Test it on a problem you've already solved. Bring one past formulation failure to a working session. Anonymize it if you like. Share your raw notes, keep the answer to yourself, and see how REACTOR sorts the evidence and where it points. Schedule a working session.


Notes and references

A note on the term. "Evidential" also appears in machine learning research, where evidential deep learning refers to a specific method for estimating a model's uncertainty (Sensoy et al., 2018). We use the word in its everyday sense: answers tied to evidence, with the evidence shown.

References

  1. Kalai, A. T., Nachum, O., Vempala, S. S., & Zhang, E. (2025). Why language models hallucinate. arXiv:2509.04664. https://arxiv.org/abs/2509.04664
  2. Mirza, A., et al. (2025). A framework for evaluating the chemical knowledge and reasoning abilities of large language models against the expertise of chemists. Nature Chemistry, 17(7), 1027-1034. https://doi.org/10.1038/s41557-025-01815-x
  3. Sensoy, M., Kaplan, L., & Kandemir, M. (2018). Evidential deep learning to quantify classification uncertainty. Advances in Neural Information Processing Systems 31 (NeurIPS 2018). https://arxiv.org/abs/1806.01768
  4. Chamberlin, T. C. (1890). The method of multiple working hypotheses. Science, ns-15(366), 92-96. https://doi.org/10.1126/science.ns-15.366.92
  5. Platt, J. R. (1964). Strong inference. Science, 146(3642), 347-353. https://doi.org/10.1126/science.146.3642.347

Ready to Transform Your R&D?

See how Reactor's chemistry-native AI can accelerate your research and development workflows.

Schedule a Demo