A benchmark study of scoring methods for non-coding mutations

Damien Drubay, Daniel Gautheret, Stefan Michiels

    Résultats de recherche: Contribution à un journalArticleRevue par des pairs

    17 Citations (Scopus)

    Résumé

    Motivation Detailed knowledge of coding sequences has led to different candidate models for pathogenic variant prioritization. Several deleteriousness scores have been proposed for the non-coding part of the genome, but no large-scale comparison has been realized to date to assess their performance. Results We compared the leading scoring tools (CADD, FATHMM-MKL, Funseq2 and GWAVA) and some recent competitors (DANN, SNP and SOM scores) for their ability to discriminate assumed pathogenic variants from assumed benign variants (using the ClinVar, COSMIC and 1000 genomes project databases). Using the ClinVar benchmark, CADD was the best tool for detecting the pathogenic variants that are mainly located in protein coding gene regions. Using the COSMIC benchmark, FATHMM-MKL, GWAVA and SOM liver outperformed the other tools for pathogenic variants that are typically located in lincRNAs, pseudogenes and other parts of the non-coding genome. However, all tools had low precision, which could potentially be improved by future non-coding genome feature discoveries. These results may have been influenced by the presence of potential benign variants in the COSMIC database. The development of a gold standard as consistent as ClinVar for these regions will be necessary to confirm our tool ranking. Availability and implementation The Snakemake, C++ and R codes are freely available from https://github.com/Oncostat/BenchmarkNCVTools and supported on Linux.

    langue originaleAnglais
    Pages (de - à)1635-1641
    Nombre de pages7
    journalBioinformatics
    Volume34
    Numéro de publication10
    Les DOIs
    étatPublié - 15 mai 2018

    Contient cette citation