Apply Now

We are looking for a Machine Learning Research Scientist, Evaluations to join our GenAI Research Organisation. You will develop rigorous evaluations and diagnostic methods to reveal where frontier models fail and why. You will collaborate with researchers and engineers to define best practices in evaluation-driven AI development and partner with top foundation model labs to translate failure analysis into technical and strategic input.

Responsibilities:

  • Analyse model behaviour to identify, characterise, and diagnose failure modes in frontier LLMs and Agents.
  • Design and build benchmarks and evaluation methods that measure LLM capabilities in both text and multimodal modalities.
  • Apply post-training expertise to connect observed failures to the data and training interventions that address them.
  • Publish research findings in top-tier AI conferences.

Requirements:

  • Ph.D. or Master's degree in Computer Science, Machine Learning, AI, or a related field.
  • Deep understanding of deep learning, reinforcement learning, and large-scale model fine-tuning.
  • Experience with post-training techniques and LLM evaluation or benchmark development.
  • Excellent written and verbal communication skills.
  • Published research in areas of machine learning at major conferences and/or journals.

Benefits:

  • Comprehensive health, dental and vision coverage
  • Retirement benefits
  • Learning and development stipend
  • Generous PTO
  • Commuter stipend (may be eligible)