Labs — PharmaTools.AI

Labs

Where ideas get tested

Experimental AI projects exploring what's possible at the intersection of machine learning, pharmaceutical data, and healthcare communication. Some of these will become products. Some won't. That's the point.

What we're exploring

Advanced AI Models

Transforming raw pharmaceutical data into actionable intelligence. Detecting patterns invisible to traditional analysis — new connections between compounds, outcomes, and patient responses.

Data-Driven Insights

Bringing together pharmaceutical scientists, AI researchers and healthcare innovators. Contributing to open research initiatives and validation studies.

Community Collaboration

A growing community of pharmaceutical professionals, researchers, and AI enthusiasts. Contributing to experiments and shaping the future of AI in pharma.

How we measure our own tools

OpenGATE Open source · gates every commit
Open Grounded AI Testing & Evaluation — an open-source verification framework for evidence-grounded AI: hand-labelled gold cases and scorers spanning claim extraction, citation accuracy, verdicts, hallucination, consistency, redaction recall, simplification faithfulness and retrieval fidelity, with a CI regression gate so reliability can't quietly slip. Four implementations across four capability shapes: RefCheckr (evidence QA, where a measured comparison drove the switch to sonar-pro), Redacta (redaction — two engine bugs found first run), Patiently AI (simplification — caught dropping safety-critical specifics), and PubCrawl (retrieval — the deterministic layer everything grounds on). Every finding fixed and verified the same day.
~0.95claim-extraction f1
5.8→2.4%passage hallucination
100%citation detection
9metric families
Open Source Gold Sets Regression Gate Confusion Matrix
Explore →
Redacta Gauntlet v1 · reasoning + downstream injection measured, CI-gated
An adversarial harness that attacks Redacta at every layer it has — the deterministic engine, the LLM reasoning layer, and the models downstream that consume its output. Recall-first, over-redaction as the cost axis, measured on Claude and Perplexity Sonar. Gaps named, boundaries owned.
91.5%in-scope recall
0→100%reasoning lift · layer-2
0downstream leaks
100%injection resistance
Adversarial Cases Reasoning Layer Downstream Consumer Model Comparison Regression Gate
Explore →

Some experiments find their feet. Some don't. Either outcome teaches something useful.

All experiments

Redacta Shipped · on the App Store
Pseudonymises patient identifiers — NHS numbers, dates, names — so clinical text can be safely processed by AI, then restores them afterwards. Now a free iPhone app, agent skill & MCP server.
iPhone App OpenClaw Skill MCP Server Privacy PII Redaction Open Source
Explore →
Time-Machine Chess Live · chess.pharmatools.ai
Chess engines fine-tuned on 150 years of history — face the gambit-happy attackers of 1850, the Classical masters of the 1920s, or the Soviet school. Each era bot is a Maia-2 neural network fine-tuned on period games and validated against the historical record it learned from: the King's Gambit gradient, first-move fashions and draw-rate shifts all reproduce in self-play. The same validation-first method as the pharma tools, applied to something purely for joy.
3playable eras · 1840–1985
2.5Mtraining positions from period games
30×king's gambit era gradient
+4.8ptsera move prediction vs base
PyTorch Maia-2 Self-Play Validation FastAPI Railway Open Source
Play →
SideEffectViz Active
Interactive visualization and clustering of medication side effects using FDA adverse event data and machine learning.
Python Railway scikit-learn NetworkX
Explore →
The RSI Loop Active
A self-improving ergonomic loop, gated by a clinical-compliance auditor. An RSI for RSI — and a pattern for keeping agentic AI inside its lane.
Python MediaPipe AI Safety Agentic AI
Explore →
BiomarkerFinder Active
Discover key biomarkers in cancer and their role in diagnosis, prognosis, and treatment — powered by Open Targets and AI insights.
NLP Oncology Precision Medicine Open Targets
Explore →
PlaceboGPT Active
The world's safest medical AI. A 7,666-parameter model exploring safety through incapability in healthcare AI.
PyTorch LSTM AI Safety Satire
Explore →
Atacama Active
A 7,762-parameter LSTM that predicts Atacama Desert weather with 99.9% accuracy. The world's most overengineered "no."
PyTorch LSTM Satire 30KB
Explore →
PubCrawl Active
MCP server giving AI assistants verifiable access to PubMed and Europe PMC literature, US/UK drug labelling, and ClinicalTrials.gov — real papers and data, no hallucinated citations.
MCP PubMed Europe PMC Drug labels Trials TypeScript
Explore →
Media Monitoring Hub Active
LLM-powered media coverage monitoring, classification, and automated briefing generation.
LLM NLP Monitoring Briefings
Explore →
PathwAIs Active
Personalised eLearning that adapts in real time to a learner's level, available time, and learning style.
LLM Adaptive Learning eLearning Personalisation
Explore →

Want to collaborate?

Got a research idea, a dataset, or just want to geek out about AI in pharma? I'm always up for a conversation.

Get in Touch