nf-core/ampliseq: Ampliseq Version 3.0.0
- Daniel Straub1
- Daniel Lundin2
- Jeanette Tångrot3
- Alexander Peltzer4
- nf-core bot
- Matthew5
- Pieter Provoost6
- emnilsson
- Sateesh_Peri7
- Adam Bennett8
- Tom Brekke
- 4less
- John Sundh9
- DiegoBrambilla
- Sofoklis Keisaris
- Lokeshwaran Manoharan10
- Till Englert11
- Asaf Peer12
- Aaron O'Brien13
- Maxime U Garcia
- Nicolas Henry14
- Thomas Weber15
- Chandini
- Matthias Hörtenhuber16
- Sam Minot17
- Thom Griffioen18
- Zack Mudge19
- Dan Clayton20
- Gisela Gabernet
- Venkat Malladi
- 1. Quantitative Biology Center (QBiC)
- 2. Linnaeus University & Stockholm University
- 3. National Bioinformatics Infrastructure Sweden (NBIS)
- 4. Boehringer Ingelheim
- 5. Hartree Centre
- 6. UNESCO
- 7. PeriMatrix IT Consultancy
- 8. OceanOmics Centre at the University of Western Australia
- 9. @NBISweden
- 10. Lund University
- 11. QBiC @qbicsoftware
- 12. HUJI
- 13. Universidad Andrés Bello
- 14. Analysis and Bioinformatics for Marine Science (ABiMS)
- 15. Data Science Centre, EMBL
- 16. @ScilifelabDataCentre
- 17. Cirro Bio
- 18. @Syngenta
- 19. APHL-CDC Public Health Laboratory Fellow (Bioinformatics)
- 20. STFC Hartree Centre
Description
nf-core/ampliseq version 3.0.0 - 2026-10-09
Summary of changes
This is a major version update, several parameters and settings of previous versions are not compatible any more, so make sure to read this information and the current pipeline documentation carefully!
Most interesting additions:
- Nanopore/ONT analysis support (R10.4 sequencing preferably with SUP basecalling) with Savont ASV calling
- Multi-database taxonomy classification with DADA2 with consolidation options via
--consolidate_taxonomies - Additional long-format summary tables to simplify downstream analysis
- Comparison of expected to computed sequences (
--expected_sequences), abundances (--expected_abundances), and taxonomic profile (--expected_abundances) and related parameter
Breaking usage changes:
- Folder input removed (
--input_folder,--extension,--multiple_sequencing_runs) - Legacy sample sheet support (see documentation) and SBDI export removed (
--sbdiexport) - Choose the sequencing raw data type with
--sequencing_type(default:illumina_pe) - Choose the ASV calling tool with
--asv_calling(default:auto, selects recommended tool accordint to data type) - Primer input now via
--primer_fwd&--primer_revinstead of--FW_primer&--RV_primer - Default for
--sample_inferencewas changed fromindependenttopooled, that means more sensitive ASV calling but also higher resource requirements, set to "independent" for previous behaviour - Reverse strand mapping for DADA2 classification method with
--dada_taxonomy_rcisnt enforced with IonTorrent and PacBio any more and needs to be set explicitly now
Detailed changes
Added
- #1021,#1023,#1025,#1048,#1054,#1094,#1102,#1104 - Comparison of observed ASVs and taxonomic profiles against expected data is now available with
--expected_*parameters (by @d4straub, reviewed by @erikrikarddaniel) - #1022 - added new GloSED database for classification of Eukaryotic ITS. Not support for SH numbers as of yet. (by @tom-brekke)
- #1026,#1046,#1049,#1079,#1103 - Add support for Oxford Nanopore Technology (ONT) R10.4 sequencing (preferably with SUP basecalling) with Savont (by @d4straub & @erikrikarddaniel)
- #1042,#1045,#1050,#1108 - DADA2 taxonomy tables (
ASV_tax.*.tsv,ASV_tax_species.*.tsv) now include one<rank>_confidencecolumn per rank, in addition to the existing overallconfidencecolumn; they are left out of the phyloseq and TreeSummarizedExperiment objects (by @erikrikarddaniel) - #1042 - Added CI test coverage for
--addsh, which previously had none (test_pacbio_itsnow also runs DADA2 taxonomy against the UNITE database it already uses for SINTAX) (by @erikrikarddaniel) - #1052,#1094,#1116 - Added
summary_tables/ampliseq.counts.tsv.gz, ASV counts in long format with consistent, lower-case column names, ready for analysis in R, Python or similar without pipeline-specific parsing; also written as Parquet by default, skip with--skip_parquet_summary(by @erikrikarddaniel) - #1056 -
--dada_ref_taxonomynow accepts a comma-separated list of databases (e.g.gtdb,silva), running DADA2 taxonomic classification against each; one full set of output files is published per listed database, and the first-listed database feeds every downstream step that expects a single taxonomy (consolidating multiple databases into one is planned as a future addition) (by @erikrikarddaniel) - #1062,#1105,#1112 - Added
--consolidate_taxonomies(firstby default,most-specificorscore) to pick a per-ASV winning database across multiple--dada_ref_taxonomydatabases, instead of always using the first-listed one; the winning result is published asdada2/ASV_tax.consolidated.<method>.tsv(by @erikrikarddaniel) - #1063,#1107 - Added
summary_tables/ampliseq.taxonomy.<classifier>.<database>.tsv.gz, one taxonomy table per classifier and database, with per-ASV barrnap, decontam and filter annotations (documentation) (by @erikrikarddaniel) - #1065 - New
--max_ee_roverrides--max_eefor reverse reads only (paired-end Illumina data), for runs where read quality drops on the reverse read only (fixes #1037) (by @erikrikarddaniel) - #1066 - Added
--sintax_cutoffparameter (maintaining the default: 0.8). (by @pieterprovoost) - #1070 - VSEARCH LCA now also publishes the raw
--blast6outhit table (ASV_tax_vsearch_lca.*.txt/ASV_ITS_tax_vsearch_lca.*.txt) alongside the existing LCA and taxonomy TSV outputs (by @pieterprovoost) - #1078,#1094 - Added DADA2 reference taxonomic databases SILVA 144 (
silva&silva=144) and Kraken2 from June 2026 (standard&standard=20260626) (by @d4straub)
Changed
- #1018 - Change version to 2.19.0dev (by @d4straub)
- #1027,#1094,#1114 - Changed parameters, including the default of
--sample_inferencefromindependenttopooledand--dada_taxonomy_rcisnt enforced with IonTorrent and PacBio any more (by @d4straub)
| previously | now | comment |
| -------------------------- | ----------------------------- | ------------------------------------------ |
| --nanopore | --sequencing_type nanopore | |
| --pacbio | --sequencing_type pacbio | |
| --iontorrent | --sequencing_type iontorrent | |
| --single_end | --sequencing_type illumina_se | |
| default | --sequencing_type illumina_pe | default |
| | --asv_calling auto | default |
| default | --asv_calling dada2 | new |
| | --asv_calling savont | new |
| --illumina_pe_its | --illumina_pe_readthrough | |
| --extension | | removed |
| --input_folder | | removed, use --input |
| --multiple_sequencing_runs | | removed, use the run column of --input |
| --FW_primer | --primer_fwd | |
| --RV_primer | --primer_rev | |
| --classifier | --qiime_classifier | |
| --sbdiexport | | removed |
- #1028 - Multi-region sample sheet via
--multiregionhad its header changed fromFW_primerandRV_primertoprimer_fwdandprimer_rev, respectively. (by @d4straub) - #1032 - Refactor the pipeline's parameter handling and initialization (by @erikrikarddaniel).
- #1051 - Continue #1032's parameter-handling refactor:
qiime2_ancomandqiime2_diversitysubworkflows now take all their flags (ancom,ancombc,ancombc2,qiime_adonis_formula) as explicit arguments instead of reading some directly fromparams.*; no behaviour change (by @erikrikarddaniel). - #1074 -
--max_ee_ris now covered by the existingtest_novaseqprofile instead of a dedicated test, removing the extra CI run added in #1065 (by @erikrikarddaniel) - #1084 - The top taxonomic rank is now called
Domainrather thanKingdomwherever the database holds a domain (Bacteria, Archaea, Eukaryota) there (fixes #1059) (by @erikrikarddaniel)
| affects | previously | now |
| ---------------------------------------------------------------------------------------- | -------------------------------------------- | ----------------------------------------------------------------------------- |
| --dada_ref_taxonomy GTDB, SILVA 132/138/138.2, RDP, Greengenes2 | Kingdom | Domain |
| --dada_ref_taxonomy SBDI-GTDB | Domain, Kingdom (domain repeated) | Domain |
| --dada_ref_taxonomy SILVA 144, PR2 | Domain, Kingdom | unchanged |
| UNITE, COIDB and the other SINTAX and VSEARCH-LCA databases | Kingdom | unchanged, their top rank is a kingdom |
| --dada_ref_tax_custom without --dada_assign_taxlevels | Kingdom | unchanged |
| DADA2 ASV_tax*.tsv and the phyloseq and TreeSummarizedExperiment objects built from it | Kingdom column | the database's own top rank, as above |
| summary tables ampliseq.taxonomy.*.tsv.gz | kingdom column for every database | the database's own ranks in lower case; domain for QIIME2 and pplace |
| summary_report.html, QIIME2 section | "classified at Kingdom level" | "classified at Domain level" |
| --consolidate_taxonomies | output named after the first-listed database | unchanged, Domain/Kingdom still treated as one rank when matching columns |
- #1087 - Replaced the workflow overview figure with a simplified metro map. (by @d4straub)
- #1088,#1094 - Local modules and subworkflows moved from
<name>.nfto<name>/main.nf, following the nf-core directory structure; every local module now has ameta.ymland anenvironment.yml, the latter a symlink for modules sharing an environment inmodules/local/envs/; the unused QIIME2 2024.10 environment file is removed (by @erikrikarddaniel) - #1090 - Change version to 3.0.0 (by @d4straub)
- #1092 -
RENAME_RAW_DATA_FILESnow runs in a Seqera container with bash and coreutils, so-profile condaand container profiles report the same bash version; trimmed code comments and corrected two typos in the documentation (by @erikrikarddaniel)
Fixed
- #1019 - Improve channel assignment & improve some version reporting (by @d4straub)
- #1034 - Fix reading taxonomy tables when taxon names contain
#(by @pieterprovoost) - #1038 - Ensure that the ASV count matrix in exported R objects is consistently stored as integer regardless of the pipeline parameters (by @hindrek)
- #1042 - DADA2's
confidencecolumn was incorrectly included as an extra, spurious rank when building the taxonomy string imported into QIIME2; no longer included (by @erikrikarddaniel) - #1057 - Fixed DADA2 read number tracking table sequence (reported by @Malytherin, fixed by @d4straub)
- #1058 - Fixed QIIME2 caching (reported by @luciazifcakova, fixed by @d4straub)
- #1077 - Template update for nf-core/tools version 4.1.0 (by @d4straub)
- #1080 -
SUMMARY_REPORTno longer fails with "input file name collision" when--report_abstract,--metadata,--inputor--input_fastashare a file name (fixes #1073) (by @erikrikarddaniel) - #1082 - Fixed two kinds of non-rank content in the taxonomy string imported into QIIME2: a rank that is empty for every ASV was rendered as the literal taxon
NA, and the UNITESHand COIDBBOLD_binidentifiers were treated as a taxonomic rank, adding a spurious extra level to the barplots and the rank-collapsed abundance tables (fixes #1075) (by @erikrikarddaniel) - #1085 -
QIIME2_EXPORT_RELTAX,QIIME2_FEATURETABLE_GROUPandQIIME2_INTREEasked for 1 GB, less than QIIME2 needs to start, so their first attempt was killed for running out of memory; they now get 3 GB like the other QIIME2 processes (by @erikrikarddaniel) - #1086 - The processes that reformat reference databases used a legacy Singularity image whose pull could hang indefinitely; they now use a Seqera container declaring every tool their scripts call (fixes #1081) (by @erikrikarddaniel)
- #1111 - Pinned nf-schema 2.7.2 in AWS full-test launch config (by @d4straub)
Dependencies
- #1035 - ITSxRust 0.2.2 to 0.3.0, reducing peak memory approximately six-fold on large inputs; extraction output is unchanged (by @ayobi)
- #1077 - MultiQC 1.34 to 1.35 (by @d4straub)
- #1089 - Updated nf-core modules and subworkflows, and removed the stale
mafftpin from.nf-core.yml(by @erikrikarddaniel)
| software | previously | now | | ------------ | ---------- | ---------- | | ITSxRust | 0.2.2 | 0.3.0 | | Savont | | 0.7.0 | | MultiQC | 1.34 | 1.35 | | gappa | 0.8.0 | 0.9.0 | | HMMER | 3.3.2 | 3.4 | | Nextflow | 25.10.4 | 26.04.0 | | Chopper | | 0.12.0b | | Porechop_ABI | | 0.5.0post1 | | DuckDB | | 1.5.5 | | sed | 4.7 | 4.10 | | bash | 5.0.17 | 5.2.37 |
Removed
- #1028 - Removed support for the legacy sample sheet to simplify parsing (by @d4straub)
- #1083 - Removed
--sbdiexportand theSBDI/output files; submission files for the Swedish Biodiversity Infrastructure will be produced by a separate pipeline instead. The SBDI-GTDB reference database is unaffected and remains the default for--dada_ref_taxonomy(closes #1055) (by @erikrikarddaniel) - #1086 - Removed the unused
taxref_reformat_phytoref.sh; PhytoRef is included in PR2 (by @erikrikarddaniel)
Files
nf-core/ampliseq-3.0.0.zip
Additional details
Related works
- Is supplement to
- Software: https://github.com/nf-core/ampliseq/tree/3.0.0 (URL)
Software
- Repository URL
- https://github.com/nf-core/ampliseq