Labs — PharmaTools.AI

Labs

Where ideas get tested

Experimental AI projects exploring what's possible at the intersection of machine learning, pharmaceutical data, and healthcare communication. Some of these will become products. Some won't. That's the point.

What we're exploring

Advanced AI Models

Transforming raw pharmaceutical data into actionable intelligence. Detecting patterns invisible to traditional analysis — new connections between compounds, outcomes, and patient responses.

Data-Driven Insights

Bringing together pharmaceutical scientists, AI researchers and healthcare innovators. Contributing to open research initiatives and validation studies.

Community Collaboration

A growing community of pharmaceutical professionals, researchers, and AI enthusiasts. Contributing to experiments and shaping the future of AI in pharma.

Observer Zero

How we measure our own tools

OpenGATE Open source · gates every commit
Open Grounded AI Testing & Evaluation — an open-source verification framework for evidence-grounded AI: hand-labelled gold cases and scorers spanning claim extraction, citation accuracy, verdicts, hallucination, consistency, redaction recall, simplification faithfulness and retrieval fidelity, with a CI regression gate so reliability can't quietly slip. Four implementations across four capability shapes: RefCheckr (evidence QA, where a pre-registered re-run retired the eval's own model-comparison headline and reverted production to sonar), Redacta (redaction — two engine bugs found first run), Patiently AI (simplification — caught dropping safety-critical specifics), and PubCrawl (retrieval — the deterministic layer everything grounds on). Every finding fixed and verified the same day.
0.91claim-extraction f1
100%citation detection
9metric families

Passage hallucination · RefCheckr re-run · five 66-pair runs per arm, pre-registered · lower is better

sonar
0.0%
sonar-pro
0.6%
reasoning-pro
14.6%
Open Source Gold Sets Regression Gate Confusion Matrix
Explore →
Redacta Gauntlet v1 · reasoning + downstream injection measured, CI-gated
An adversarial harness that attacks Redacta at every layer it has — the deterministic engine, the LLM reasoning layer, and the models downstream that consume its output. Recall-first, over-redaction as the cost axis, measured on Claude and Perplexity Sonar. Gaps named, boundaries owned.
91.5%in-scope recall
0downstream leaks
100%injection resistance

Layer-2 identifier capture · deterministic engine alone vs + reasoning layer

engine alone
0%
+ reasoning
100%
Adversarial Cases Reasoning Layer Downstream Consumer Model Comparison Regression Gate
Explore →
LitRAG Open source · benchmarked vs RAGAS & DeepEval
A small, readable RAG pipeline over PubMed abstracts with a citation-faithfulness eval built in: every claim must carry a verbatim quote, checked by a deterministic locator before an LLM judge grades support from the passage alone. Benchmarked against the faithfulness metrics of RAGAS and DeepEval on a 61-claim hand-labelled set with the same Claude judge for all three systems — LitRAG caught all 34 unfaithful claims and was the only system to flag a true claim citing a fabricated quote, the case that separates citation checking from claim checking. Its single error was a conservative false alarm, the safe failure direction for medicine. The corpus is live too: ingest.py pulls fresh abstracts for any PubMed query through the PubCrawl MCP server.
1.000hallucination recall — all 34 unfaithful claims caught
0.984accuracy on the 61-claim gold set
9/9fabricated quotes caught deterministically — no judge call spent

Hallucination recall · same Claude judge, 61 labelled claims · higher is better

LitRAG
1.000
RAGAS
0.912
DeepEval
0.647
RAG Citation Faithfulness Benchmark PubCrawl Open Source
Explore →

Some experiments find their feet. Some don't. Either outcome teaches something useful.

All experiments

Pharma & clinical data

The core of the lab: privacy, literature, biomarkers and adverse-event data.

Redacta Shipped · on the App Store
Pseudonymises patient identifiers — NHS numbers, dates, names — so clinical text can be safely processed by AI, then restores them afterwards. Now a free iPhone app, agent skill, MCP server & self-hosted Kubernetes deployment.
iPhone App OpenClaw Skill MCP Server Privacy PII Redaction Open Source
Explore →
PubCrawl Active
MCP server giving AI assistants verifiable access to PubMed and Europe PMC literature, US/UK drug labelling, and ClinicalTrials.gov — real papers and data, no hallucinated citations.
MCP PubMed Europe PMC Drug labels Trials TypeScript
Explore →
StudyDiff Live · web app & MCP server
Explains why two papers reach opposite conclusions. It extracts each study's design into a structured card, shows which dimensions differ and which are identical — ruled out as the cause — and grounds every value in a verbatim source sentence: click any verified field to see it highlighted in the paper's own text. Fields the source doesn't state come back as "not reported", never guessed, and verification is deterministic (OpenGATE) — no LLM acts as judge. It used to rank the differences and name a likely driver; a benchmark of 15 documented contradictions showed that ranking was no better than always guessing "assay", so it was retired. That result was then confirmed blind on a second, held-out set of 15 contradictions — curated to a protocol fixed before any case was chosen, and measured once. StudyDiff now shows the differences unranked and leaves which one matters to the reader's field.
13.3%top-1 driver on both benchmarks — at or below the always-guess-“assay” baseline
0invented — every value quoted from source
0/25non-assay contradictions the retired ranking identified, across both sets

Reachable ceiling on the development set · before → after the grounding fix · higher is better

before
33%
after
67%
MCP PubMed Grounding Benchmark OpenGATE Open Source
Explore →
BiomarkerFinder Active
Discover key biomarkers in cancer and their role in diagnosis, prognosis, and treatment — powered by Open Targets and AI insights.
NLP Oncology Precision Medicine Open Targets
Explore →
SideEffectViz Active
Interactive visualization and clustering of medication side effects using FDA adverse event data and machine learning.
Python Railway scikit-learn NetworkX
Explore →

Want to collaborate?

Got a research idea, a dataset, or just want to geek out about AI in pharma? I'm always up for a conversation.

Get in Touch