What your AI doesn't know it doesn't know

What your AI doesn't know it doesn't know

There's a particular failure mode in biopharma AI that doesn't announce itself. The model runs, the analysis completes, the report looks reasonable. And somewhere inside it, a drug was counted twice because it was acquired and recoded. A terminated trial was missed because the stop reason wasn't normalized. A synonym wasn't caught because the ontology wasn't mapped.

The output looks authoritative. The errors are invisible. That's the problem.

LLMs learn language, not facts

To understand why this happens, it helps to understand what a large language model actually is. LLMs are trained to predict the next word in a sequence. They learn from an enormous amount of text, and through that process, they get very good at understanding language, following instructions, summarizing complex information, and reasoning across context.

What they don't do is store facts. Think of an LLM like a well-read but occasionally forgetful colleague. They've absorbed a huge amount of information, but they can't always tell you whether they remember something correctly or are filling in a gap with something that sounds right. The difference between that colleague and a reliable one is that the reliable one knows when to check the source.

That distinction matters enormously when you're asking an AI agent to reason over clinical trial data.

The problem compounds quietly

Here's what a multi-step analysis looks like when the underlying data isn't AI-ready. Take a realistic task: identify all drugs in clinical trials for non-small cell lung cancer, stratify by max phase, and flag trial failures.

Without harmonized, structured data, the errors accumulate at every step.

The agent searches for "non-small cell lung cancer" and misses trials filed under "non-small cell lung carcinoma," a clinically equivalent term. It identifies a drug by one identifier without knowing it was acquired and recoded under a different one, so the drug appears twice in the analysis. It calculates max phase based only on the trials in its current context, missing the fact that some of these drugs are already approved for other indications and being repurposed. It checks for trial failures but misses one that was terminated for safety reasons because the stop reason wasn't in a form the model could reliably interpret. By the time the final report is generated, the context window is straining under the accumulated data, and the model starts missing trials it had seen earlier.

The report gets produced. Parts of it are wrong. The model has no reliable way to flag which parts.

What AI-ready data actually changes

Run the same task with DrugBank's MCP and the failure modes disappear one by one.

A search against DrugBank's disease ontology captures synonyms and child terms automatically, so "non-small cell lung carcinoma" is covered without a separate query. Drug counts are correct because the data tracks acquisitions, recodings, and prior approvals — the model isn't inferring these relationships, they're explicit. Trial stop reasons are curated and normalized, so failure signals are accurate and consistent. And because DrugBank's MCP handles the counting and aggregation directly, the model receives answers rather than raw data, which means the context window stays clean and the final summary is reliable.

The difference isn't that the model got smarter. It's that the data gave the model a complete picture to reason from.

The completeness question

This is the crux of what makes data AI-ready. It's not just accuracy. It's coverage.

An AI agent can only reason over what it has. If a trial is missing from the dataset, the agent doesn't know the trial exists. It can't flag the gap. It can't caveat the conclusion. It produces an analysis based on incomplete information and has no mechanism to tell you that.

A completeness guarantee changes the nature of the task. When the data covers the full picture, the agent's job becomes analysis rather than assembly. That distinction matters for performance, for cost, and for the reliability of the output.

The four properties that define AI-ready data in practice:

Coverage. A completeness guarantee means you're getting the full picture, not a sample of it. Missing data doesn't show up as an error, it shows up as a wrong answer.

Curation. Raw data from regulatory filings and clinical trial registries is inconsistent, incomplete, and often ambiguous. Curated data adds the context, normalization, and expert judgment that turns text into something a model can trust.

Structure and ontologies. Complex biomedical concepts like diseases exist under dozens of synonyms, hierarchical relationships, and coding systems. Structured, ontology-mapped data makes those relationships explicit so models can search and filter accurately without needing to resolve ambiguity themselves.

Model-ready outputs. When an agent can query for a count and receive a number rather than a list of records to parse, it spends less of its context window on assembly and more on reasoning. That's not a minor optimization; it's what separates analyses that hold up under scrutiny from ones that quietly fall apart.

The infrastructure question behind the AI question

The conversation in biopharma right now is largely about models: which one, how to deploy it, what use cases to start with. That's the right conversation. But it's the second conversation. The first one is about the data those models reason from.

A model given incomplete, inconsistent, or unstructured data doesn't fail loudly. It produces a confident output that looks right until someone checks it. In drug discovery and clinical development, the cost of that is not a software bug. It's a wrong conclusion at a decision point that matters.

AI-ready data is what makes the difference between an agent that hallucinates and one you can trust.


DrugBank's intelligence graph covers 156 million structured data points across drugs, targets, diseases, and trials, unified across 20-plus ontologies and accessible via MCP for direct integration into your AI workflows.