Abdul Muntakim Rafi

Research

Learning the cis-regulatory code from sequence.

How do we learn the cis-regulatory code from sequence, and how do we know when to trust what a model has learned?

Progress on cis-regulation is limited by data rather than architecture. The genome offers too few examples, they are too correlated with one another, and its homology structure makes standard evaluations dishonest. Almost all of that data was also generated for a different purpose: characterising a system and training a model on it are different objectives, and technology built deliberately to produce training data is rare.

So the work runs in two directions at once. One builds the data, using synthetic sequence to escape the genome’s ceilings and designing it for information content rather than count. The other keeps the resulting models honest: evaluated without leakage, reported faithfully, and able to say how far an individual prediction can be trusted.

Themes

The questions the work is organised around.

How do we generate the data?

Experiments built for the model, not inherited from someone else’s

A model of regulation is bounded by the sequences it was trained on. The genome supplies too few of them and they are too correlated, and synthetic libraries that lift that ceiling saturate in turn, because most of a dataset’s value arrives early and volume alone stops buying accuracy. I have worked on random libraries that remove the homology and sample-size limits, on chromosome-scale sequence that removes fixed context, and now on designing libraries around information content rather than count.

How do we train the best models?

Which design decisions actually matter

A trained model is the product of dozens of choices made at once. Architecture, objective, augmentation and data handling all move the result, and a paper comparing two models cannot say which of them was responsible. Holding the data fixed and letting training vary across many independent attempts separated the factors for one setting, where it was the trainer rather than the architecture that carried the gain.

What have the models learned?

Causal structure, or the shape of the training set

A model that predicts well is only informative if what it learned was regulation. Homologous sequences land on both sides of a standard split, so a model that recalls its neighbours scores like one that understands them, and every assay leaves biases a model will fit as readily as the biology. I built tools that detect that homology and partition data around it, and am mapping it across the genome so the correction does not have to be recomputed each time.

Can we trust a single prediction?

Reliability per prediction, not per benchmark

Anyone applying these models cares about one variant, not an average. Aggregate correlations say nothing about that case, and the usual workaround of thresholding on predicted effect size discards the low-magnitude variants where most GWAS signal is expected to act. I built a meta-model that scores how far an individual prediction can be trusted, and that says which features drive the score.

Can we trust how we read them?

The interpretation tools are instruments too

Most biological claims from these models arrive through an interpretation method rather than from the model itself. Attribution and perturbation methods are models in their own right, far less tested than the networks they are pointed at, and a distorted lens turns an accurate model into a wrong conclusion without any benchmark catching it.

Projects

Projects, published and in progress.

hashFrag — homology, leakage and memorization

Preprint

Neither chromosomal nor random train/test splits account for homology within a species, so standard evaluations of genome-trained models are inflated. We measured how far, showed that the dependence on training-set similarity is not monotonic, and released hashFrag, which detects homology and partitions data at roughly a hundredth of the compute of exhaustive alignment. Its recommendation is to stratify a test set rather than build a fully orthogonal one, because an orthogonal split hides the bias instead of exposing it.

pairFrag — genome-wide homology mapping

In preparation

Making homology-aware evaluation something any group can do without repeating the computation.

Random Promoter DREAM Challenge

Published

Random sequence removes the homology and sample-size ceilings at once. I designed and ran an open challenge on 6.7 million random promoters measured in yeast, with more than 110 teams and 28 final models; nineteen beat the previous state of the art. The winner had the fewest parameters and three of the top five used no transformer, so I built a framework that recombines the entrants’ modules across architectures and trainers to find out which choice carried the gain. It was the trainer. Models tuned on random yeast sequence then transferred to other species and assays.

Chromosome-scale sequence from outside the host

Ongoing

Short oligos sit in one fixed context, so they cannot report on promoter–gene distance, chromatin or real transcripts. Sequence carried on yeast artificial chromosomes behaves like an extra chromosome and has never been under selection in the organism reading it. We annotated a single YAC in Luthra et al., then increased the data tenfold to train models on it, which became Yorzoi, built with Timon Schneider and Tom Ellis at Imperial College London. We are now scaling the data tenfold again for Yakformer.

High-information-content libraries

In progress

Designing sequence synthesis approaches that create high-information-content sequence libraries.

nextFrag (active learning)

Preprint

If every sequence has to be paid for, each one should be chosen to be informative. We benchmarked six selection strategies across architectures, datasets and configurations, simulated on pools that had already been measured, so the benchmark itself needed no new experiment. All beat random sampling, uncertainty-based methods did best while being cheapest to compute, and most of the gain from many small acquisition rounds survives with fewer, larger ones — which is what makes lab-in-the-loop practical. Selected sequences look distinctive, but selecting directly on those properties never matched active learning: informativeness is a property of the model’s ignorance, not of the sequence. Building on this, we are extending the work to large-scale experimental data, to report how active learning is best done in genomics.

gRely — reliability of individual predictions

Preprint

A meta-model that estimates the probability an individual variant-effect prediction is correct, from features of the variant, gene, tissue and model. Its top-scoring fifth reaches 97% sign concordance against 54% in the bottom fifth, and it stays discriminative among the low-magnitude variants that effect-size filtering discards, which is where most GWAS signal is expected to act. It transfers zero-shot to other architectures, so reliability looks like a property of the locus rather than of the model. Begun during an internship at Genentech.

Interested in collaborating on any of these? Reach out .

Funding

Grants and compute.

Fellowships, grants and compute awards, with the role I held on each.

Four Year Doctoral Fellowship (4YF)

University of British Columbia

96,000 CAD over four years. UBC’s flagship doctoral fellowship, awarded with tuition to top PhD students.

2021 – 2025

Continual improvement of gene regulatory models

The Digital Research Alliance of Canada · PI: Carl de Boer

Priority access to GPUs. Co-wrote the proposal.

2025 – 2026

Evaluation, optimization and continual improvement of sequence-based cis-regulatory models

Advanced Research Computing, UBC · PI: Carl de Boer

20,000 CAD in Microsoft Azure credit. Co-applicant; wrote the proposal.

Aug 2023 – Jun 2024

Lossless preprocessing of the sequence and expression space to improve sequence-based gene regulatory models

School of Biomedical Engineering, UBC · PI: Carl de Boer

6,000 CAD. Co-applicant; wrote the proposal and hired a Co-op student through the grant.

May – Aug 2023

Random Promoter DREAM Challenge 2022

TPU Research Cloud, Google · PI: Carl de Boer (UBC), Pablo Meyer (IBM Research), Jake Albrecht (Sage Bionetworks)

50 TPU quotas. Sole graduate student on the organising committee.

May – Jul 2022

Identifying selection on human gene expression with gene regulatory models

The Digital Research Alliance of Canada · PI: Carl de Boer

Priority access to GPUs. Co-wrote the proposal.

2022 – 2024

Efficient edge inference benchmarking for AI-driven applications

Mitacs Accelerate · PI: Jonathan Wu

15,000 CAD. Sole co-applicant; wrote the proposal.

Nov 2020 – Mar 2021

Spatio-temporal human activity recognition on manufacturing floors

Mitacs Accelerate · PI: Jonathan Wu

22,500 CAD. Co-applicant; assisted with proposal writing.

Oct 2019 – Apr 2020