BiologyComputer ScienceMedicine

Gennady Gorin, L. Pachter

2021.7.31Biophysical Reports

DOI: 10.1016/j.bpr.2022.100097

tlooto Summary

It is found that long pre-mRNA transcripts are over-represented in sequencing data, and the mechanistic implications are explored, leading to a length-based model of capture bias.

Abstract

Single-molecule pre-mRNA and mRNA sequencing data can be modeled and analyzed using the Markov chain formalism to yield genome-wide insights into transcription. However, quantitative inference with such data requires careful assessment and understanding of noise sources. We find that long pre-mRNA transcripts are over-represented in sequencing data, and explore the mechanistic implications. A biological explanation for this phenomenon within our modeling framework requires unrealistic transcriptional parameters, leading us to posit a length-based model of capture bias. We provide solutions for this model, and use them to find concordant and mechanistically plausible parameter trends across data from multiple single-cell RNA-seq experiments in several species.

Citation format

GORIN, Gennady; PACHTER, L. Length biases in single-cell RNA sequencing of pre-mrna. Biophysical Reports, 2021, 3.