Profluent Machine Learning Scientist to drive research into pretraining large-scale deep learning models for biomolecular design. Collaborate across teams to apply and improve pretraining techniques for protein language models.
Responsibilities
We're looking for a motivated and creative Machine Learning (ML) Scientist to drive research into pretraining large-scale deep learning models for biomolecular design. This position offers an opportunity to work at the forefront of generative modeling research across language processing, representation learning, and protein engineering. You should be a self-directed researcher who has the ability to rapidly prototype and evaluate new models and algorithms in the biomolecular domain.
As an early employee, you will proactively shape the direction of our machine learning efforts and collaborate across diverse teams of computational and experimental scientists.
Design and develop state-of-the-art autoregressive, diffusion and representation learning models for protein design
Collaborate across the machine learning and protein design teams to apply, adapt and improve pretraining techniques from other domains to biomolecular deep learning models
Architect, implement, and optimize core infrastructure to support the pretraining of protein language models
Curate relevant datasets and design tasks for rigorous evaluation of generative models
Implement, analyze, and interpret multiple computational approaches and present results to colleagues in regular update meetings
Work within a collaborative, fast-paced, interdisciplinary team across biology and machine learning to help shape the scientific and strategic vision of the company
Qualification
Experience with conceiving of
Required
PhD (or equivalent industry experience) in Computer Science, Machine Learning, Natural Language Processing, Applied Math, Computational Biology, Statistics, or a related field
Experience with conceiving of, implementing, and evaluating novel machine learning and pretraining large scale LLMs or other models in biomolecular domain
Publications at major machine learning conferences (NeurIPS, ICML, ICLR) or scientific journals (Nature, Science, Nature Biotech, Nature Methods, PNAS) or experience pre-training LLMs at frontier AI/ML labs
Experience with modern deep learning frameworks such as Pytorch or Jax
Familiarity with foundational biology of proteins and nucleic acids