17 References
Bankevich, Anton, Sergey Nurk, Dmitry Antipov, et al. 2012.
“SPAdes: A New Genome Assembly Algorithm and Its
Applications to Single-Cell Sequencing.” Journal of
Computational Biology 19 (5): 455–77. https://doi.org/10.1089/cmb.2012.0021.
Bolger, Anthony M., Marc Lohse, and Bjoern Usadel. 2014.
“Trimmomatic: A Flexible Trimmer for Illumina
Sequence Data.” Bioinformatics 30 (15): 2114–20. https://doi.org/10.1093/bioinformatics/btu170.
Bonfield, James K., John Marshall, Petr Danecek, et al. 2021.
“HTSlib: C Library for Reading/Writing
High-Throughput Sequencing Data.” GigaScience 10 (2):
giab007. https://doi.org/10.1093/gigascience/giab007.
Bray, Nicolas L., Harold Pimentel, Páll Melsted, and Lior Pachter. 2016.
“Near-Optimal Probabilistic RNA-seq
Quantification.” Nature Biotechnology 34 (5): 525–27. https://doi.org/10.1038/nbt.3519.
Chen, Shifu, Yanqing Zhou, Yaru Chen, and Jia Gu. 2018. “Fastp: An
Ultra-Fast All-in-One FASTQ Preprocessor.”
Bioinformatics 34 (17): i884–90. https://doi.org/10.1093/bioinformatics/bty560.
Cheng, Haoyu, Gregory T. Concepcion, Xiaowen Feng, Haowen Zhang, and
Heng Li. 2021. “Haplotype-Resolved de Novo Assembly Using Phased
Assembly Graphs with Hifiasm.” Nature Methods 18 (2):
170–75. https://doi.org/10.1038/s41592-020-01056-5.
Cleary, John G., Ross Braithwaite, Kurt Gaastra, et al. 2015.
“Comparing Variant Call Files for Performance Benchmarking of
Next-Generation Sequencing Variant Calling Pipelines.”
bioRxiv, ahead of print. https://doi.org/10.1101/023754.
Cock, Peter J. A., Tiago Antao, Jeffrey T. Chang, et al. 2009.
“Biopython: Freely Available Python Tools for
Computational Molecular Biology and Bioinformatics.”
Bioinformatics 25 (11): 1422–23. https://doi.org/10.1093/bioinformatics/btp163.
Cooke, Daniel P., David C. Wedge, and Gerton Lunter. 2021. “A
Unified Haplotype-Based Method for Accurate and Comprehensive Variant
Calling.” Nature Biotechnology 39 (7): 885–92. https://doi.org/10.1038/s41587-021-00861-3.
Crusoe, Michael R., Sanne Abeln, Alexandru Iosup, et al. 2022.
“Methods Included: Standardizing Computational Reuse and
Portability with the Common Workflow Language.”
Communications of the ACM 65 (6): 54–63. https://doi.org/10.1145/3486897.
Danecek, Petr, Adam Auton, Goncalo Abecasis, et al. 2011. “The
Variant Call Format and VCFtools.”
Bioinformatics 27 (15): 2156–58. https://doi.org/10.1093/bioinformatics/btr330.
Danecek, Petr, James K. Bonfield, Jennifer Liddle, et al. 2021.
“Twelve Years of SAMtools and
BCFtools.” GigaScience 10 (2): giab008. https://doi.org/10.1093/gigascience/giab008.
De Coster, Wouter, and Rosa Rademakers. 2023.
“NanoPack2: Population-Scale Evaluation of Long-Read
Sequencing Data.” Bioinformatics 39 (5): btad311. https://doi.org/10.1093/bioinformatics/btad311.
Di Tommaso, Paolo, Maria Chatzou, Evan W. Floden, Pablo Prieto Barja,
Emilio Palumbo, and Cedric Notredame. 2017. “Nextflow Enables
Reproducible Computational Workflows.” Nature
Biotechnology 35 (4): 316–19. https://doi.org/10.1038/nbt.3820.
Dobin, Alexander, Carrie A. Davis, Felix Schlesinger, et al. 2013.
“STAR: Ultrafast Universal RNA-seq Aligner.” Bioinformatics
29 (1): 15–21. https://doi.org/10.1093/bioinformatics/bts635.
Ewels, Philip, Måns Magnusson, Sverker Lundin, and Max Käller. 2016.
“MultiQC: Summarize Analysis Results for Multiple
Tools and Samples in a Single Report.” Bioinformatics 32
(19): 3047–48. https://doi.org/10.1093/bioinformatics/btw354.
Ferragina, Paolo, and Giovanni Manzini. 2000. “Opportunistic Data
Structures with Applications.” Proceedings 41st Annual
Symposium on Foundations of Computer Science (FOCS), 390–98. https://doi.org/10.1109/SFCS.2000.892127.
Garrison, Erik, and Gabor Marth. 2012. Haplotype-Based Variant
Detection from Short-Read Sequencing. https://arxiv.org/abs/1207.3907.
Gurevich, Alexey, Vladislav Saveliev, Nikolay Vyahhi, and Glenn Tesler.
2013. “QUAST: Quality Assessment Tool for Genome
Assemblies.” Bioinformatics 29 (8): 1072–75. https://doi.org/10.1093/bioinformatics/btt086.
Hsi-Yang Fritz, Markus, Rasko Leinonen, Guy Cochrane, and Ewan Birney.
2011. “Efficient Storage of High Throughput DNA
Sequencing Data Using Reference-Based Compression.” Genome
Research 21 (5): 734–40. https://doi.org/10.1101/gr.114819.110.
Huang, Neng, and Heng Li. 2023. “Compleasm: A Faster and More
Accurate Reimplementation of BUSCO.”
Bioinformatics 39 (10): btad595. https://doi.org/10.1093/bioinformatics/btad595.
Kim, Daehwan, Joseph M. Paggi, Chanhee Park, Christopher Bennett, and
Steven L. Salzberg. 2019. “Graph-Based Genome Alignment and
Genotyping with HISAT2 and HISAT-genotype.” Nature
Biotechnology 37 (8): 907–15. https://doi.org/10.1038/s41587-019-0201-4.
Kim, Sangtae, Konrad Scheffler, Aaron L. Halpern, et al. 2018.
“Strelka2: Fast and Accurate Calling of Germline and Somatic
Variants.” Nature Methods 15 (8): 591–94. https://doi.org/10.1038/s41592-018-0051-x.
Kolmogorov, Mikhail, Jeffrey Yuan, Yu Lin, and Pavel A. Pevzner. 2019.
“Assembly of Long, Error-Prone Reads Using Repeat Graphs.”
Nature Biotechnology 37 (5): 540–46. https://doi.org/10.1038/s41587-019-0072-8.
Koren, Sergey, Brian P. Walenz, Konstantin Berlin, Jason R. Miller,
Nicholas H. Bergman, and Adam M. Phillippy. 2017. “Canu: Scalable
and Accurate Long-Read Assembly via Adaptive k-Mer Weighting and Repeat
Separation.” Genome Research 27 (5): 722–36. https://doi.org/10.1101/gr.215087.116.
Köster, Johannes, and Sven Rahmann. 2012. “Snakemake — a Scalable
Bioinformatics Workflow Engine.” Bioinformatics 28 (19):
2520–22. https://doi.org/10.1093/bioinformatics/bts480.
Langmead, Ben, and Steven L. Salzberg. 2012. “Fast Gapped-Read
Alignment with Bowtie 2.” Nature Methods 9 (4): 357–59.
https://doi.org/10.1038/nmeth.1923.
Law, Charity W., Yunshun Chen, Wei Shi, and Gordon K. Smyth. 2014.
“Voom: Precision Weights Unlock Linear Model Analysis Tools for
RNA-seq Read Counts.” Genome
Biology 15 (2): R29. https://doi.org/10.1186/gb-2014-15-2-r29.
Lawrence, Michael, Robert Gentleman, and Vincent Carey. 2009.
“Rtracklayer: An R Package for Interfacing with
Genome Browsers.” Bioinformatics 25 (14): 1841–42. https://doi.org/10.1093/bioinformatics/btp328.
Li, Bo, and Colin N. Dewey. 2011. “RSEM: Accurate
Transcript Quantification from RNA-Seq Data with or Without
a Reference Genome.” BMC Bioinformatics 12 (1): 323. https://doi.org/10.1186/1471-2105-12-323.
Li, Heng. 2011. “Tabix: Fast Retrieval of Sequence Features from
Generic TAB-Delimited Files.”
Bioinformatics 27 (5): 718–19. https://doi.org/10.1093/bioinformatics/btq671.
Li, Heng. 2013. Aligning Sequence Reads, Clone Sequences and
Assembly Contigs with BWA-MEM. https://arxiv.org/abs/1303.3997.
Li, Heng. 2018. “Minimap2: Pairwise Alignment for Nucleotide
Sequences.” Bioinformatics 34 (18): 3094–100. https://doi.org/10.1093/bioinformatics/bty191.
Li, Heng, and Richard Durbin. 2009. “Fast and Accurate Short Read
Alignment with Burrows–Wheeler Transform.”
Bioinformatics 25 (14): 1754–60. https://doi.org/10.1093/bioinformatics/btp324.
Li, Heng, Bob Handsaker, Alec Wysoker, et al. 2009. “The Sequence
Alignment/Map Format and SAMtools.” Bioinformatics 25
(16): 2078–79. https://doi.org/10.1093/bioinformatics/btp352.
Liao, Yang, Gordon K. Smyth, and Wei Shi. 2014. “featureCounts: An Efficient General Purpose
Program for Assigning Sequence Reads to Genomic Features.”
Bioinformatics 30 (7): 923–30. https://doi.org/10.1093/bioinformatics/btt656.
Love, Michael I., Wolfgang Huber, and Simon Anders. 2014.
“Moderated Estimation of Fold Change and Dispersion for RNA-seq Data with DESeq2.”
Genome Biology 15 (12): 550. https://doi.org/10.1186/s13059-014-0550-8.
Manni, Mosè, Matthew R. Berkeley, Mathieu Seppey, Felipe A. Simão, and
Evgeny M. Zdobnov. 2021. “BUSCO Update: Novel and
Streamlined Workflows Along with Broader and Deeper Phylogenetic
Coverage for Scoring of Eukaryotic, Prokaryotic, and Viral
Genomes.” Molecular Biology and Evolution 38 (10):
4647–54. https://doi.org/10.1093/molbev/msab199.
Martin, Marcel. 2011. “Cutadapt Removes Adapter Sequences from
High-Throughput Sequencing Reads.” EMBnet.journal 17
(1): 10. https://doi.org/10.14806/ej.17.1.200.
McKenna, Aaron, Matthew Hanna, Eric Banks, et al. 2010. “The
Genome Analysis Toolkit: A MapReduce Framework
for Analyzing Next-Generation DNA Sequencing Data.”
Genome Research 20 (9): 1297–303. https://doi.org/10.1101/gr.107524.110.
Mölder, Felix, Kim Philipp Jablonski, Brice Letcher, et al. 2025.
“Sustainable Data Analysis with Snakemake.”
F1000Research 10: 33. https://doi.org/10.12688/f1000research.29032.3.
Needleman, Saul B., and Christian D. Wunsch. 1970. “A General
Method Applicable to the Search for Similarities in the Amino Acid
Sequence of Two Proteins.” Journal of Molecular Biology
48 (3): 443–53. https://doi.org/10.1016/0022-2836(70)90057-4.
Neph, Shane, M. Scott Kuehn, Alex P. Reynolds, et al. 2012.
“BEDOPS: High-Performance Genomic Feature
Operations.” Bioinformatics 28 (14): 1919–20. https://doi.org/10.1093/bioinformatics/bts277.
Patro, Rob, Geet Duggal, Michael I. Love, Rafael A. Irizarry, and Carl
Kingsford. 2017. “Salmon Provides Fast and Bias-Aware
Quantification of Transcript Expression.” Nature Methods
14 (4): 417–19. https://doi.org/10.1038/nmeth.4197.
Poplin, Ryan, Pi-Chuan Chang, David Alexander, et al. 2018. “A
Universal SNP and Small-Indel Variant Caller Using Deep
Neural Networks.” Nature Biotechnology 36 (10): 983–87.
https://doi.org/10.1038/nbt.4235.
Quinlan, Aaron R., and Ira M. Hall. 2010. “BEDTools:
A Flexible Suite of Utilities for Comparing Genomic Features.”
Bioinformatics 26 (6): 841–42. https://doi.org/10.1093/bioinformatics/btq033.
Rautiainen, Mikko, Sergey Nurk, Brian P. Walenz, et al. 2023.
“Telomere-to-Telomere Assembly of Diploid Chromosomes with
Verkko.” Nature Biotechnology 41 (10):
1474–82. https://doi.org/10.1038/s41587-023-01662-6.
Rhie, Arang, Brian P. Walenz, Sergey Koren, and Adam M. Phillippy. 2020.
“Merqury: Reference-Free Quality, Completeness, and Phasing
Assessment for Genome Assemblies.” Genome Biology 21
(1): 245. https://doi.org/10.1186/s13059-020-02134-9.
Ritchie, Matthew E., Belinda Phipson, Di Wu, et al. 2015. “Limma
Powers Differential Expression Analyses for RNA-Sequencing
and Microarray Studies.” Nucleic Acids Research 43 (7):
e47. https://doi.org/10.1093/nar/gkv007.
Roberts, Michael, Wayne Hayes, Brian R. Hunt, Stephen M. Mount, and
James A. Yorke. 2004. “Reducing Storage Requirements for
Biological Sequence Comparison.” Bioinformatics 20 (18):
3363–69. https://doi.org/10.1093/bioinformatics/bth408.
Robinson, Mark D., Davis J. McCarthy, and Gordon K. Smyth. 2010.
“edgeR: A Bioconductor
Package for Differential Expression Analysis of Digital Gene Expression
Data.” Bioinformatics 26 (1): 139–40. https://doi.org/10.1093/bioinformatics/btp616.
Sahlin, Kristoffer. 2022. “Strobealign: Flexible Seed Size Enables
Ultra-Fast and Accurate Read Alignment.” Genome Biology
23 (1): 260. https://doi.org/10.1186/s13059-022-02831-7.
Sena Brandine, Guilherme de, and Andrew D. Smith. 2019. “Falco:
High-Speed FastQC Emulation for Quality Control of
Sequencing Data.” F1000Research 8: 1874. https://doi.org/10.12688/f1000research.21142.1.
Shen, Wei, Shuai Le, Yan Li, and Fuquan Hu. 2016.
“SeqKit: A Cross-Platform and Ultrafast Toolkit for
FASTA/Q File Manipulation.” PLOS ONE 11
(10): e0163962. https://doi.org/10.1371/journal.pone.0163962.
Smith, Temple F., and Michael S. Waterman. 1981. “Identification
of Common Molecular Subsequences.” Journal of Molecular
Biology 147 (1): 195–97. https://doi.org/10.1016/0022-2836(81)90087-5.
Smith, Tom, Andreas Heger, and Ian Sudbery. 2017. “UMI-tools: Modeling Sequencing Errors in
Unique Molecular Identifiers to Improve Quantification
Accuracy.” Genome Research 27 (3): 491–99. https://doi.org/10.1101/gr.209601.116.
Soneson, Charlotte, Michael I. Love, and Mark D. Robinson. 2016.
“Differential Analyses for RNA-seq:
Transcript-Level Estimates Improve Gene-Level Inferences.”
F1000Research 4: 1521. https://doi.org/10.12688/f1000research.7563.2.
Tan, Adrian, Gonçalo R. Abecasis, and Hyun Min Kang. 2015.
“Unified Representation of Genetic Variants.”
Bioinformatics 31 (13): 2202–4. https://doi.org/10.1093/bioinformatics/btv112.
Vasimuddin, Md., Sanchit Misra, Heng Li, and Srinivas Aluru. 2019.
“Efficient Architecture-Aware Acceleration of BWA-MEM
for Multicore Systems.” 2019 IEEE International Parallel and
Distributed Processing Symposium (IPDPS), 314–24. https://doi.org/10.1109/IPDPS.2019.00041.
Wingett, Steven W., and Simon Andrews. 2018. “FastQ
Screen: A Tool for Multi-Genome Mapping and Quality
Control.” F1000Research 7: 1338. https://doi.org/10.12688/f1000research.15931.2.
Zheng, Zhenxian, Shumin Li, Junhao Su, Amy Wing-Sze Leung, Tak-Wah Lam,
and Ruibang Luo. 2022. “Symphonizing Pileup and Full-Alignment for
Deep Learning-Based Long-Read Variant Calling.” Nature
Computational Science 2 (12): 797–803. https://doi.org/10.1038/s43588-022-00387-x.
Zook, Justin M., Jennifer McDaniel, Nathan D. Olson, et al. 2019.
“An Open Resource for Accurately Benchmarking Small Variant and
Reference Calls.” Nature Biotechnology 37 (5): 561–66.
https://doi.org/10.1038/s41587-019-0074-6.