Location: London, UK
Employment Type: Full time
Overview
Novogaia is an applied AI drug discovery company. We build machine learning systems that decode the chemistry of natural organisms, starting with fungi, to find the next generation of medicines.
Our models are only as good as the data they learn and are measured against. We work with a large natural-product compound library, generate mass spectrometry data against it through analytical partners, and need someone who can turn that raw output into a dataset our models, and our evaluation pipelines, can actually trust.
We are a small team of AI engineers and chemists building foundation models for molecular structure prediction, and we're looking for someone to own the pipeline that takes raw MS/MS data from the instrument to something clean, annotated, and ready to train or test a model on.
We are seeking a computational data scientist who can define what a trustworthy spectral dataset looks like, and build the schema, QC gates, and annotation process that gets us there. This is a data and cheminformatics role, not a wet-lab role: you'll work closely with analytical chemistry collaborators who run the instruments, but your job is the pipeline, the metadata, and the chemistry judgment that sits on top of it.
The Role
- Develop a deep understanding of Novogaia's compound library, spectral data, and how both feed our models
- Own the metadata schema for MS/MS data: structures, formulae, adducts, precursor masses, collision energies, instrument metadata, retention time, and quality scores
- Build QC and ingest validation for raw and processed spectra arriving from analytical partners, including batch tracking and versioned releases
- This role begins with hands-on schema design, QC, and annotation work, and can grow into ownership of Novogaia's broader data infrastructure and dataset strategy
- Perform chemical annotation: molecular class, natural-product class, structural complexity, and ionisation behaviour, and maintain provenance and licensing tags at the record level
- Work with analytical chemistry collaborators to route ambiguous or high-value spectra for expert review, and structure the resulting data so it can be compared systematically against model predictions
- Support our evaluation team with de-duplication and leakage checks (InChIKey, scaffold, and structural-similarity matching) against public spectral data, and help select representative slices of the dataset for internal and external benchmarking
- In your first year, you'll build the data foundation everything else depends on: a documented schema, a working QC process, and a curated, well-annotated dataset that both our modelling team and our evaluation team can rely on without double-checking your work.
What We Require
- Background in analytical mass spectrometry or cheminformatics (PhD or equivalent industry experience), ideally with exposure to natural products or small-molecule discovery
- Hands-on experience working with LC-MS/MS data and standard formats and tools (e.g. mzML, MSConvert) and spectral databases (e.g. GNPS, MassBank, MoNA)
- Scripting ability in Python or R for building data pipelines