Location: London, UK
Employment Type: Full time
Overview
Novogaia is an applied AI drug discovery company. We build machine learning systems that decode the chemistry of natural organisms, starting with fungi, to find the next generation of medicines.
The same model architectures pushing the frontier of language and reasoning are now being adapted to read the chemical signals nature produces. Getting a model to perform well on a benchmark is easy. Getting it to genuinely reason about chemistry, instead of leaning on database priors, metadata shortcuts, or quirks in how a dataset was built, is the hard part, and the part that matters if the model is going to be trusted with real discovery decisions.
We are a small team of AI engineers and chemists building foundation models for molecular structure prediction. We are seeking a machine learning research engineer who excels at designing rigorous tests for molecular AI systems, and who can turn "does this model actually work" into a concrete, defensible answer. You will own the design of our internal benchmarks and evaluation pipelines; build the adversarial checks that catch shortcut learning and leakage before they reach a customer or a research paper; and work closely with our modelling team to translate evaluation results into research priorities.
The Role
- Develop a deep understanding of Novogaia's models, data, and evaluation needs
- Design benchmark tasks that reflect real discovery problems: structure generation, retrieval, spectrum simulation, and property/class prediction
- Build the datasets and controls that make a benchmark trustworthy: hard negatives, leakage-safe splits, and null baselines that catch a model exploiting shortcuts instead of genuine signal
- This role begins with hands-on benchmark construction and evaluation engineering, and can grow into ownership of Novogaia's broader evaluation strategy and model-development priorities
- Communicate technical work: turn evaluation results into clear findings for the modelling team, and into documentation and data cards that others can trust and reproduce
- Work across the team to scope and lead evaluation work, including:
- Defining what "good" looks like for a given model or task, and choosing the right test for it
- Analyzing model behavior and interpreting results for researchers and non-technical stakeholders alike
- Working with engineers to turn one-off analyses into repeatable, reproducible evaluation pipelines
- Translate lessons from evaluation work into research priorities, data requirements, and modelling direction
In your first year, you'll take the lead on how Novogaia measures model quality, from benchmark design through adversarial testing and reporting. The tests you build will decide how much weight anyone can put on our models' outputs.
What We Require
- Research or applied experience in machine learning, with direct experience building or rigorously evaluating ML benchmarks
- Exceptional technical communication skills, including the ability to explain evaluation findings clearly to both researchers and non-technical stakeholders
- Ability to analyze model behavior and interpret computational results critically
- Strong proficiency in Python, and comfort with reproducible, containerised pipelines