Machine Learning Scientist, Biological Foundation Models
We are looking for someone who wants to build new machine-learning architectures from first principles around biological sequence and experimental data. This is not a role for wrapping existing language models around a protein database. You will work directly with large, internally generated datasets linking sequence to expression, binding, stability and other measured phenotypes, and develop models that can learn from the structure of those experiments.
What you will do
- Design and implement new neural architectures for protein sequence, experimental measurements and iterative design data.
- Build models that learn jointly from sequence, assay context, measurement uncertainty and repeated experimental rounds.
- Develop active-learning and experiment-selection strategies for directed evolution and high-throughput protein optimization.
- Work directly with wet-lab scientists to determine what data should be generated next, not just how existing data should be modeled.
- Establish rigorous evaluation methods for generalization across proteins, targets, assays and experimental campaigns.
What we are looking for
- Deep experience developing modern machine-learning models in PyTorch, JAX or an equivalent framework.
- Strong grounding in representation learning, generative modeling, transformers, diffusion, graph methods or related architectures.
- Comfort building models from scratch rather than relying exclusively on pretrained APIs.
- Experience with biological sequence data is valuable, but exceptional ML researchers from adjacent fields are encouraged to apply.