Academic Journal
TOFU-MAaPO: fast, scalable and reproducible analysis of large metagenome sequence data from the Sequence Read Archive.
| Τίτλος: | TOFU-MAaPO: fast, scalable and reproducible analysis of large metagenome sequence data from the Sequence Read Archive. |
|---|---|
| Συγγραφείς: | Wacker EM; Institute of Clinical Molecular Biology, Kiel University, Kiel, Germany., Rühlemann MC; Institute of Clinical Molecular Biology, Kiel University, Kiel, Germany.; Institute for Medical Microbiology and Hospital Epidemiology, Hannover Medical School, Hannover, Germany., Franke A; Institute of Clinical Molecular Biology, Kiel University, Kiel, Germany., Ellinghaus D; Institute of Clinical Molecular Biology, Kiel University, Kiel, Germany. d.ellinghaus@ikmb.uni-kiel.de. |
| Πηγή: | Nature communications [Nat Commun] 2026 Jun 11; Vol. 17 (1). Date of Electronic Publication: 2026 Jun 11. |
| Τύπος έκδοσης: | Journal Article; Research Support, Non-U.S. Gov't |
| Γλώσσα: | English |
| Στοιχεία περιοδικού: | Publisher: Nature Pub. Group Country of Publication: England NLM ID: 101528555 Publication Model: Electronic Cited Medium: Internet ISSN: 2041-1723 (Electronic) Linking ISSN: 20411723 NLM ISO Abbreviation: Nat Commun Subsets: MEDLINE |
| Imprint Name(s): | Original Publication: [London] : Nature Pub. Group |
| Ιατρικοί όροι (MeSH): | Metagenome*/genetics , Metagenomics*/methods , Software*, Sequence Analysis, DNA/methods ; Humans ; Reproducibility of Results ; High-Throughput Nucleotide Sequencing ; Shotgun Sequencing ; Workflow |
| Περίληψη: | Metagenomic shotgun sequencing data from over 600,000 metagenomes are publicly available in repositories such as NCBI's Sequence Read Archive (SRA). Technically advanced and easy-to-use best-practice metagenome software workflows for raw data pre-processing, assembly of metagenome-assembled genomes, and taxonomic and functional annotation of metagenome-assembled genomes are needed for reproducible analysis and harmonization of large-scale metagenomic datasets. We introduce TOFU-MAaPO (Taxonomic Or FUnctional Metagenomic Assembly and PrOfiling), a portable, automated single-command Nextflow pipeline for large-scale analysis of metagenomic short-read sequencing data. It analyzes metagenome files locally or directly from the SRA using accession or study IDs. In a benchmark against three established metagenome software pipelines, the TOFU-MAaPO workflow yielded 12%, 42% to 77% more high-quality metagenome-assembled genomes, likely reflecting the integration of multiple complementary binning tools with a unified refinement strategy. Using its assembly-free taxonomic abundance profiling module, we also automatically downloaded 16,462 uniquely identifiable and accessible human gut metagenome samples from the SRA and taxonomically annotated them against the Genome Taxonomy Database on a high-performance cluster in less than 55 hours, including download time. TOFU-MAaPO makes large metagenome projects more accessible to individual research groups and is freely available at https://github.com/ikmb/TOFU-MAaPO . (© 2026. The Author(s).) |
| Competing Interests: | Competing interests: The authors declare no competing interests. |
| References: | Almeida, A. et al. A unified catalog of 204,938 reference genomes from the human gut microbiome. Nat. Biotechnol. 39, 105–114 (2021). (PMID: 3269097310.1038/s41587-020-0603-3) Rühlemann M. et al. Disease signatures in the gut metagenome of a prospective family cohort of inflammatory bowel disease. Preprint at medRxiv https://doi.org/10.1101/2023.12.10.23299783 . Wang, J. & Jia, H. Metagenome-wide association studies: fine-mining the microbiome. Nat. Rev. Microbiol. 14, 508–522 (2016). (PMID: 2739656710.1038/nrmicro.2016.83) MGI: The Million Microbiome of Humans Project (MMHP) Officially Launched to Build the World’s Largest Human Microbiome Database-MGI Tech Website-Leading Life Science Innovation (accessed 20 February 2025); https://en.mgi-tech.com/News/info/id/96 (2019). Di Tommaso, P. et al. Nextflow enables reproducible computational workflows. Nat. Biotechnol. 35, 316–319 (2017). (PMID: 2839831110.1038/nbt.3820) Köster, J. & Rahmann, S. Snakemake-a scalable bioinformatics workflow engine. Bioinformatics 28, 2520–2522 (2012). (PMID: 2290821510.1093/bioinformatics/bts480) Knight, R. et al. Best practices for analysing microbiomes. Nat. Rev. Microbiol. 16, 410–422 (2018). (PMID: 2979532810.1038/s41579-018-0029-9) Kurtzer, G. M., Sochat, V. & Bauer, M. W. Singularity: scientific containers for mobility of compute. PLoS ONE 12, e0177459 (2017). (PMID: 28494014542667510.1371/journal.pone.0177459) Quince, C., Walker, A. W., Simpson, J. T., Loman, N. J. & Segata, N. Shotgun metagenomics, from sampling to analysis. Nat. Biotechnol. 35, 833–844 (2017). (PMID: 2889820710.1038/nbt.3935) Krakau, S., Straub, D., Gourlé, H., Gabernet, G. & Nahnsen, S. nf-core/mag: a best-practice pipeline for metagenome hybrid assembly and binning. NAR Genom. Bioinform. 4, lqac007 (2022). (PMID: 35118380880854210.1093/nargab/lqac007) Kieser, S., Brown, J., Zdobnov, E. M., Trajkovski, M. & McCue, L. A. ATLAS: a Snakemake workflow for assembly, annotation, and genomic binning of metagenome sequence data. BMC Bioinforma. 21, 257 (2020). (PMID: 10.1186/s12859-020-03585-4) Uritskiy, G. V., DiRuggiero, J. & Taylor, J. MetaWRAP-a flexible pipeline for genome-resolved metagenomic data analysis. Microbiome 6, 158 (2018). (PMID: 30219103613892210.1186/s40168-018-0541-1) McIver, L. J. et al. bioBakery: a meta’omic analysis environment. Bioinformatics 34, 1235–1237 (2018). (PMID: 29194469603094710.1093/bioinformatics/btx754) Clarke, E. L. et al. Sunbeam: an extensible pipeline for analyzing metagenomic sequencing experiments. Microbiome 7, 46 (2019). (PMID: 30902113642978610.1186/s40168-019-0658-x) Van Damme, R. et al. Metagenomics workflow for hybrid assembly, differential coverage binning, metatranscriptomics and pathway analysis (MUFFIN). PLoS Comput. Biol. 17, e1008716 (2021). (PMID: 33561126789936710.1371/journal.pcbi.1008716) Lee, H. G. et al. metaFun: an analysis pipeline for metagenomic big data with fast and unified functional searches. Gut Microbes 18, 2611544 (2026). (PMID: 415309171281882210.1080/19490976.2025.2611544) Rühlemann, M. C., Wacker, E. M., Ellinghaus, D. & Franke, A. MAGScoT: a fast, lightweight and accurate bin-refinement tool. Bioinformatics 38, 5430–5433 (2022). (PMID: 36264141975010110.1093/bioinformatics/btac694) Nagy-Szakal, D. et al. Fecal metagenomic profiles in subgroups of patients with myalgic encephalomyelitis/chronic fatigue syndrome. Microbiome 5, 44 (2017). (PMID: 28441964540546710.1186/s40168-017-0261-y) Sieber, C. M. K. et al. Recovery of genomes from metagenomes via a dereplication, aggregation and scoring strategy. Nat. Microbiol. 3, 836–843 (2018). (PMID: 29807988678697110.1038/s41564-018-0171-1) Chklovski, A., Parks, D. H., Woodcroft, B. J. & Tyson, G. W. CheckM2: a rapid, scalable and accurate tool for assessing microbial genome quality using machine learning. Nat. Methods 20, 1203–1212 (2023). (PMID: 3750075910.1038/s41592-023-01940-w) Pasolli, E. et al. Accessible, curated metagenomic data through ExperimentHub. Nat. Methods 14, 1023–1024 (2017). (PMID: 29088129586203910.1038/nmeth.4468) Parks, D. H. et al. A complete domain-to-species taxonomy for Bacteria and Archaea. Nat. Biotechnol. 38, 1079–1086 (2020). (PMID: 3234156410.1038/s41587-020-0501-8) Malmstrom, R. R. Quality MAGnified. Nat. Rev. Microbiol. 21, 771 (2023). (PMID: 3778907510.1038/s41579-023-00981-4) Han, H., Wang, Z. & Zhu, S. Benchmarking metagenomic binning tools on real datasets across sequencing platforms and binning modes. Nat. Commun. 16, 2865 (2025). (PMID: 401285351193369610.1038/s41467-025-57957-6) Huerta-Cepas, J. et al. eggNOG 5.0: a hierarchical, functionally and phylogenetically annotated orthology resource based on 5090 organisms and 2502 viruses. Nucleic Acids Res. 47, D309–d14 (2019). (PMID: 30418610632407910.1093/nar/gky1085) Cantalapiedra, C. P., Hernández-Plaza, A., Letunic, I., Bork, P. & Huerta-Cepas, J. eggNOG-mapper v2: functional annotation, orthology assignments, and domain prediction at the metagenomic scale. Mol. Biol. Evol. 38, 5825–5829 (2021). (PMID: 34597405866261310.1093/molbev/msab293) conda: a system-level, binary package and environment manager running on all major operating systems and platforms. contributors c (25.1.1) https://docs.conda.io/projects/conda/%20https:/github.com/conda/conda . Andrews, S. et al. FastQC (Version 0.11.9) https://github.com/s-andrews/FastQC/releases/tag/v0.11.9 (2012). Bushnell, B. BBDuk (Version 39.00) https://sourceforge.net/projects/bbmap/ (2022). Chen, S., Zhou, Y., Chen, Y. & Gu, J. fastp: an ultra-fast all-in-one FASTQ preprocessor. Bioinformatics 34, i884–i890 (2018). (PMID: 30423086612928110.1093/bioinformatics/bty560) Langmead, B. & Salzberg, S. L. Fast gapped-read alignment with Bowtie 2. Nat. Methods 9, 357–359 (2012). (PMID: 22388286332238110.1038/nmeth.1923) Ewels, P., Magnusson, M., Lundin, S. & Käller, M. MultiQC: summarize analysis results for multiple tools and samples in a single report. Bioinformatics 32, 3047–3048 (2016). (PMID: 27312411503992410.1093/bioinformatics/btw354) Beghini, F. et al. Integrating taxonomic, functional, and strain-level profiling of diverse microbial communities with bioBakery 3. eLife. https://doi.org/10.7554/eLife.65088 (2021). Li, D., Liu, C. M., Luo, R., Sadakane, K. & Lam, T. W. MEGAHIT: an ultra-fast single-node solution for large and complex metagenomics assembly via succinct de Bruijn graph. Bioinformatics 31, 1674–1676 (2015). (PMID: 2560979310.1093/bioinformatics/btv033) Li, H. Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics 34, 3094–3100 (2018). (PMID: 29750242613799610.1093/bioinformatics/bty191) Kang, D. D. et al. MetaBAT 2: an adaptive binning algorithm for robust and efficient genome reconstruction from metagenome assemblies. PeerJ. 7, e7359 (2019). (PMID: 31388474666256710.7717/peerj.7359) Wang, Z. et al. Effective binning of metagenomic contigs using contrastive multi-view representation learning. Nat. Commun. 15, 585 (2024). (PMID: 382333911079420810.1038/s41467-023-44290-z) Alneberg, J. et al. Binning metagenomic contigs by coverage and composition. Nat. Methods 11, 1144–1146 (2014). (PMID: 2521818010.1038/nmeth.3103) Wu, Y. W., Simmons, B. A. & Singer, S. W. MaxBin 2.0: an automated binning algorithm to recover genomes from multiple metagenomic datasets. Bioinformatics 32, 605–607 (2016). (PMID: 2651582010.1093/bioinformatics/btv638) Pan, S., Zhao, X. M. & Coelho, L. P. SemiBin2: self-supervised contrastive learning leads to better MAGs for short- and long-read sequencing. Bioinformatics 39, i21–i29 (2023). (PMID: 373871711031132910.1093/bioinformatics/btad209) Nissen, J. N. et al. Improved metagenome binning and assembly using deep variational autoencoders. Nat. Biotechnol. 39, 555–560 (2021). (PMID: 3339815310.1038/s41587-020-00777-4) Hyatt, D. et al. Prodigal: prokaryotic gene recognition and translation initiation site identification. BMC Bioinforma. 11, 119 (2010). (PMID: 10.1186/1471-2105-11-119) Johnson, L. S., Eddy, S. R. & Portugaly, E. Hidden Markov model speed heuristic and iterative HMM search procedure. BMC Bioinforma. 11, 431 (2010). (PMID: 10.1186/1471-2105-11-431) Mistry, J., Finn, R. D., Eddy, S. R., Bateman, A. & Punta, M. Challenges in homology search: HMMER3 and convergent evolution of coiled-coil regions. Nucleic Acids Res. 41, e121 (2013). (PMID: 23598997369551310.1093/nar/gkt263) Chaumeil, P. A., Mussig, A. J., Hugenholtz, P. & Parks, D. H. GTDB-Tk v2: memory friendly classification with the genome taxonomy database. Bioinformatics 38, 5315–5316 (2022). (PMID: 36218463971055210.1093/bioinformatics/btac672) Chaumeil, P. A., Mussig, A. J., Hugenholtz, P. & Parks, D. H. GTDB-Tk: a toolkit to classify genomes with the Genome Taxonomy Database. Bioinformatics 36, 1925–1927 (2019). (PMID: 31730192770375910.1093/bioinformatics/btz848) Parks, D. H., Imelfort, M., Skennerton, C. T., Hugenholtz, P. & Tyson, G. W. CheckM: assessing the quality of microbial genomes recovered from isolates, single cells, and metagenomes. Genome Res. 25, 1043–1055 (2015). (PMID: 25977477448438710.1101/gr.186072.114) Blanco-Míguez, A. et al. Extending and improving metagenomic taxonomic profiling with uncharacterized species using MetaPhlAn 4. Nat. Biotechnol. 41, 1633–1644 (2023). (PMID: 368233561063583110.1038/s41587-023-01688-w) Wood, D. E., Lu, J. & Langmead, B. Improved metagenomic analysis with Kraken 2. Genome Biol. 20, 257 (2019). (PMID: 31779668688357910.1186/s13059-019-1891-0) Lu, J., Breitwieser, F. P., Thielen, P. & Salzberg, S. L. Bracken: estimating species abundance in metagenomics data. PeerJ. Comput. Sci. 3, e104 (2017). (PMID: 402714381201628210.7717/peerj-cs.104) Shaw, J. & Yu, Y. W. Rapid species-level metagenome profiling and containment estimation with sylph. Nat. Biotechnol. https://doi.org/10.1038/s41587-024-02412-y (2024). Patro, R., Duggal, G., Love, M. I., Irizarry, R. A. & Kingsford, C. Salmon provides fast and bias-aware quantification of transcript expression. Nat. Methods 14, 417–419 (2017). (PMID: 28263959560014810.1038/nmeth.4197) Wood, D. E. & Salzberg, S. L. Kraken: ultrafast metagenomic sequence classification using exact alignments. Genome Biol. 15, R46 (2014). (PMID: 24580807405381310.1186/gb-2014-15-3-r46) manzik/cmdbench: quick and easy resource usage monitoring and benchmarking for any command’s CPU, memory, disk usage and runtime (Version 0.1.13) https://github.com/manzik/cmdbench%20https:/github.com/manzik/cmdbench . Wickham, H. et al. Welcome to the Tidyverse. J. Open Source Softw. 4, 1686 (2019). (PMID: 10.21105/joss.01686) Baglama, J., Reichel, L. & Lewis, B. W. Fast Truncated Singular Value Decomposition and Principal Components Analysis for Large Dense and Sparse Matrices [R Package Irlba Version 2.3.5.1] https://cran.r-project.org/package=irlba%20https:/cran.r-project.org/package=irlba (2022). Hadley, W. ggplot2: elegant graphics for data analysis. J. R. Stat. Soc. Ser. A: Stat. Soc. 174, 245–246 (2016). Wacker, E. M., Rühlemann, M., Franke, A. & Ellinghaus, D. ikmb/TOFU-MAaPO. https://doi.org/10.5281/zenodo.20123041 (2026). |
| Grant Information: | EL 831/5-1 Deutsche Forschungsgemeinschaft (German Research Foundation); EXC 2167/2 - 390884018 Deutsche Forschungsgemeinschaft (German Research Foundation) |
| Molecular Sequence: | SRA SRP102150 |
| Entry Date(s): | Date Created: 20260611 Date Completed: 20260614 Latest Revision: 20260726 |
| Update Code: | 20260726 |
| PubMed Central ID: | PMC13260335 |
| DOI: | 10.1038/s41467-026-74033-9 |
| PMID: | 42277027 |
| Βάση Δεδομένων: | MEDLINE |
| ISSN: | 2041-1723 |
|---|---|
| DOI: | 10.1038/s41467-026-74033-9 |