Abdul Muntakim Rafi

PhD Candidate, University of British Columbia

To benefit from scaling laws in biology, we need to synthesize the right data with the goal of training models.

I am a PhD candidate in Biomedical Engineering at the University of British Columbia, in the de Boer Lab. Almost every model of gene regulation is trained on data that was generated for some other reason: an atlas, a consortium characterisation, whatever happened to get measured. The field has built remarkable models on top of leftovers, and has rarely built the experiment for the model.

I am working on addressing that gap. I design experiments whose purpose is to synthesize the most suitable data for training models, at a throughput worth training on.

Better data is only half of it. Every experiment and dataset carries its own bias, and a model will fit that bias as readily as the biology, so the rest of my work is on ensuring the models learn causal structure rather than the correlations an assay left behind, on a faithful reporting of the model’s performance, and on knowing when to trust a prediction and the interpretation we draw from it. Trust is what makes them worth using.

I am always looking for students to work with. I have supervised six co-op and PhD students in the de Boer Lab, and motivated undergraduates and high-school students are welcome to get in touch.

Vancouver, Canada

Abdul Muntakim Rafi
Position
PhD candidate, Biomedical Engineering
Lab
de Boer Lab, School of Biomedical Engineering, UBC
Since
2021
Before
MASc, University of Windsor · BSc, BUET

Current work

What I work on.

My work spans the path from data to application: how training data is generated, how models are trained and evaluated, and when their predictions and interpretations can be trusted.

How do we generate the data?

Most data used to train models of gene regulation were generated to characterise biology, not to train models. I work on developing technologies to synthesize sequence libraries and to design experiments whose purpose is to train models.

How do we train the best models?

Model performance depends on the architecture, the training strategy and how the training data are handled. I work on all three, focusing primarily on the data.

What have the models learned?

Standard evaluations overstate what a model has learned when related sequences appear in both the training and test sets. I work on detecting this homology and accounting for it, so that evaluation separates learned regulatory logic from memorization.

Can we trust a single prediction?

Aggregate benchmark scores say little about whether an individual prediction, such as a single variant effect, is correct. I work on estimating the reliability of individual predictions.

Can we trust how we read them?

Most biological conclusions from these models are drawn through interpretation methods such as attribution and perturbation, which are themselves rarely tested. I work on assessing how faithfully they report what a model has learned.

Topics I am interested in

17 entries

  • Applied machine learning
  • Computational biology
  • Sequence-to-expression models
  • Regulatory genomics
  • Large-scale DNA synthesis
  • Massively parallel reporter assays
  • Active learning
  • Lab-in-the-loop experiments
  • Experimental automation
  • Sequence design
  • Deep learning
  • Model interpretation
  • Simulation of cis-regulation
  • Model reliability
  • Variant effect prediction
  • Data leakage
  • Benchmark design

Output

Publications, talks and supervision.

Peer-reviewed papers
11
Preprints
4
Talks given
21
Co-op and PhD students supervised
6

All publications