Please use this identifier to cite or link to this item:
https://www.um.edu.mt/library/oar/handle/123456789/148117| Title: | Advancing microRNA target site identification via bias-corrected chimeric datasets for machine learning approaches |
| Authors: | Sammut, Stephanie (2025) |
| Keywords: | Small interfering RNA Machine learning Neural networks (Computer science) |
| Issue Date: | 2025 |
| Citation: | Sammut, S. (2025). Advancing microRNA target site identification via bias-corrected chimeric datasets for machine learning approaches (Master's dissertation). |
| Abstract: | microRNAs (miRNAs) are small, non-coding RNA molecules that regulate gene expression post-transcriptionally. These ∼22-nucleotide long RNA molecules are loaded onto a protein of the Argonaute (AGO) family guiding it to specific RNA target transcripts, which are consequently inhibited from translation to protein. Despite years of research, the precise mechanisms that determine miRNA target recognition remain unclear. Given that a single miRNA can potentially target any RNA transcript, experimental validation of all possible interactions is impractical. To this end, with the recent availability of volumes of data from high-throughput CLASH experiments that capture interacting RNA molecules mediated by a protein, many miRNA target site prediction methods employing data-driven approaches have been developed. However, despite substantial efforts to produce more accurate models, currently, no standardised framework for benchmarking miRNA target site prediction methods exists, and various strategies have been adopted, hindering fair and reliable comparison. Consequently, this undermines the validity of performance claims. Recently, another high-throughput experimental technique, called chimeric eCLIP, was developed, leading to a 70-fold increase in the recovery of miRNA– target site interactions. Here, we identify a unique opportunity to leverage this immense resource of publicly available raw data, to curate benchmark datasets for miRNA target site prediction. Throughout this work, we additionally uncover a miRNA frequency class bias that arises as a result of the method used to generate negative examples in silico. Such methods are required when modeling miRNA target site prediction as a supervised binary classification problem, due to the absence of experimentally confirmed non-binding miRNA–target pairs. To this end, we develop a novel method for generating negatives that mitigates the identified bias. Leveraging data from these high-throughput experiments and applying this new method for generating negatives, we curate three collections of datasets: one novel dataset containing almost three million examples, and bias-corrected versions of two smaller, published datasets. We contribute these datasets to the publicly available and easy-to-use miRBench Python package, providing a framework for benchmarking miRNA target site prediction methods. We benchmark six state-of-the-art deep learning models on these datasets and train simple models to establish a baseline. Retraining a convolutional neural network, that was originally trained on a biased dataset (Average Precision Score (APS): 0.71), on a larger, bias-corrected dataset improves its performance (APS: 0.81), surpassing the previous state of the art (TargetScanCnn, APS: 0.76). This highlights the advantages of well-curated, unbiased benchmarks, facilitating the development of more accurate miRNA target site prediction models thus enabling more reliable downstream analyses. |
| Description: | M.Sc.(Melit.) |
| URI: | https://www.um.edu.mt/library/oar/handle/123456789/148117 |
| Appears in Collections: | Dissertations - CenMMB - 2025 |
Files in This Item:
| File | Description | Size | Format | |
|---|---|---|---|---|
| 2519MMBMMB520000005034_2.PDF | 9.57 MB | Adobe PDF | View/Open |
Items in OAR@UM are protected by copyright, with all rights reserved, unless otherwise indicated.
