Abdul Muntakim Rafi

PhD Candidate, University of British Columbia

To benefit from scaling laws in biology, we need to synthesize the right data with the goal of training models.

I am a PhD candidate in Biomedical Engineering at the University of British Columbia, in the de Boer Lab. Almost every model of gene regulation is trained on data that was generated for some other reason: an atlas, a consortium characterisation, whatever happened to get measured.

I am working on addressing that gap. I design experiments whose purpose is to synthesize the most suitable data for training models, at a throughput worth training on. The goal is to scale.

Scaling informative data is not all of my research. Every experiment and dataset carries its own bias, and a model will fit that bias as readily as the biology, so the rest of my work is on ensuring the models learn causal features rather than the correlations an assay or dataset structure has. I focus heavily on a faithful reporting of the model’s performance, and on knowing when to trust a prediction and even the interpretation we draw from it. If you can’t trust, you can’t use it.

I am always looking for students to work with. I have supervised six co-op and PhD students in the de Boer Lab, and motivated undergraduates and high-school students are welcome to get in touch.

Vancouver, Canada

Abdul Muntakim Rafi
Position
PhD candidate, Biomedical Engineering
Lab
de Boer Lab, School of Biomedical Engineering, UBC
Since
2021
Before
MASc, University of Windsor · BSc, BUET

Current work

What I work on.

My work spans from generating data for models to applying the models to understand biology: how training data is generated, how models are trained and evaluated, and when their predictions and interpretations can be trusted.

How do we generate the data?

Most data used to train models of gene regulation were generated to characterise biology, not to train models. I work on developing technologies to synthesize sequence libraries and to design experiments whose purpose is to train models.

How do we train the best models?

Model performance depends on the architecture, the training strategy, and how the training data are processed and different data points are weighted.

What have the models learned?

Standard evaluations can overstate what a model has learned, and models can learn spurious correlations and dataset structure that are predictive without being causal. I work on detecting these and accounting for them, and on evaluations robust enough to probe the different aspects of the underlying problem. Separating learned regulatory logic from memorization is a step towards causal modelling.

Can we trust a single prediction?

Model performance is usually reported over a whole held-out test set, capturing different aspects of the underlying problem. That does not tell us how much to trust an individual prediction, such as a single variant effect. I work on estimating the reliability of individual predictions.

Can we trust how we read them?

Most biological conclusions from these models are drawn through interpretation methods such as attribution and perturbation, which are themselves not extensively tested. I work on assessing how faithful different interpretation methods are.

Topics I am interested in

17 entries

  • Applied machine learning
  • Computational biology
  • Sequence-to-expression models
  • Regulatory genomics
  • Large-scale DNA synthesis
  • Massively parallel reporter assays
  • Active learning
  • Lab-in-the-loop experiments
  • Experimental automation
  • Sequence design
  • Deep learning
  • Model interpretation
  • Simulation of cis-regulation
  • Model reliability
  • Variant effect prediction
  • Data leakage
  • Benchmark design

Output

Publications, talks and supervision.

Peer-reviewed papers
11
Preprints
4
Talks given
23
Co-op and PhD students supervised
6

All publications