Skip to content
Merged
Show file tree
Hide file tree
Changes from 13 commits
Commits
Show all changes
26 commits
Select commit Hold shift + click to select a range
2723d4c
Merge pull request #1014 from nf-core/dev
d4straub Jun 17, 2026
84c38dd
adding glosed to nf-core/ampliseq
tom-brekke Jun 19, 2026
5ed586f
Merge branch 'nf-core:master' into GloSED_2
tom-brekke Jun 19, 2026
5c28e3e
bumped up the memory for glosed
tom-brekke Jun 19, 2026
036b0c1
bumped up the memory for glosed
tom-brekke Jun 19, 2026
ef61940
trying a slightly different taxonomy format
tom-brekke Jun 22, 2026
e62fa6c
updated citations to include GloSED database
tom-brekke Jun 24, 2026
99f8390
updated usage doc to include the glosed database
tom-brekke Jun 24, 2026
b18e57b
updated changelog for GloSED addition
tom-brekke Jun 24, 2026
9b8afb4
added glosed.nf.test.snap
tom-brekke Jun 24, 2026
98604f3
added a versioned glosed as well as a most-up-to-date one into the re…
tom-brekke Jun 24, 2026
4b4e760
updated changelog
tom-brekke Jun 24, 2026
7090485
Merge branch 'dev' into GloSED_2
tom-brekke Jun 24, 2026
fc4609b
Update nextflow_schema.json
tom-brekke Jun 26, 2026
39630b2
Update conf/ref_databases.config
tom-brekke Jun 26, 2026
68c013d
[automated] Fix code linting
nf-core-bot Jun 26, 2026
e522174
Merge branch 'dev' into GloSED_2
tom-brekke Jul 1, 2026
4fa332f
Merge branch 'dev' into GloSED_2
tom-brekke Jul 10, 2026
a4a5f11
updated test_glosed.config to use the scaled-down glosed database hos…
tom-brekke Jul 10, 2026
efd305e
moved the includeConfig conf/ref_databases.config line up above the p…
tom-brekke Jul 10, 2026
8edf453
updated test_glosed.config to specify the test-databases
tom-brekke Jul 10, 2026
ee35794
updated paths and reformatting for the test dataset of glosed
tom-brekke Jul 15, 2026
01ccdb9
[automated] Fix code linting
nf-core-bot Jul 15, 2026
aa6d88a
Fix typo in test for glosed 1k link
d4straub Jul 15, 2026
72aea64
Re-add appropriate resource limits
d4straub Jul 15, 2026
6fc8371
updated glosed test.snap file
tom-brekke Jul 15, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 3 additions & 1 deletion CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,8 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

### `Added`

- [[#1022](https://github.com/nf-core/ampliseq/issues/1022)] - added new GloSED database for classification of Eukaryotic ITS. Not support for SH numbers as of yet. (by @tom-brekke)

### `Changed`

- [#1018](https://github.com/nf-core/ampliseq/pull/1018) - Change version to 2.19.0dev (by @d4straub)
Expand All @@ -19,7 +21,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

### `Removed`

## nf-core/ampliseq version 2.18.0 - 2026-06-17
## nf-core/ampliseq version 2.18.0 - 2026-06-18

### `Added`

Expand Down
4 changes: 4 additions & 0 deletions CITATIONS.md
Original file line number Diff line number Diff line change
Expand Up @@ -65,6 +65,10 @@

> Kõljalg U, Larsson KH, Abarenkov K, Nilsson RH, Alexander IJ, Eberhardt U, Erland S, Høiland K, Kjøller R, Larsson E, Pennanen T, Sen R, Taylor AF, Tedersoo L, Vrålstad T, Ursing BM. UNITE: a database providing web-based methods for the molecular identification of ectomycorrhizal fungi. New Phytol. 2005 Jun;166(3):1063-8. doi: 10.1111/j.1469-8137.2005.01376.x. PMID: 15869663.

- [GloSED - Global standardised soil eukaryome dataset](https://www.nature.com/articles/s41597-026-07315-y)

> Mikryukov V, Dulya O, Abarenkov K, Anslan S, Hagh-Doust N, Prins V, Panksep K, Põlme S, Ibrahim KS, Bahram M, Adamson K, Agan A, Ahmed T, Alatalo JM, Albornoz FE, Al-Hatmi AM, Alkahtani S, Alvarez-Manjarrez J, Ankuda J, Antonelli A, Ariyan M, Armolaitis K, Aslani F, Barrio IC, Bauters M, Biersma EM, Bitenieks K, Bonito G, Brearley FQ, Bråthen KA, Buegger F, Butterbach-Bahl K, Bálint M, Cameron EK, Canini F, Casique-Valdés R, Corrales A, Davydov EA, De Crop E, De Kesel A, Djeugap JF, Drenkhan R, Duarte Ritter C, Dudov SV, Espenberg M, Fanuel O, Fedosov VE, Florence L, Furneaux BR, Furtado ANM, Färkkilä S, Gamova NS, Garibay-Orijel R, Geml J, Ghosh S, Godoy R, Gohar D, Gryzenhout M, Hasan AH, Hashem AH, Heilmann-Clausen J, Henkel TW, Hiiesalu I, Hiiesalu I, Hosseyni Moghaddam MS, Hyde KD, Inostroza KK, Kariman K, Karimullina E, Kepfer-Rojas S, Khalid AN, Klavina D, Kohout P, Korotkov YN, Kupagme JY, Kurina O, Lamit LJ, Lateef AA, Ledoux NA, Lim YW, Maciá-Vicente JG, Makovskis K, Martínez S, Marín C, Meidl P, Mortimer PE, Mundra S, Naluyange V, Netherway T, Newsham KK, Nouhra E, Nyamukondiwa C, Nteziryayo V, Ochieno DMW, Oja J, Onipchenko VG, Otsing E, Owaid MN, Piepenbring M, Pochekutova P, Pombo MM, Pritsch K, Puusepp R, Pärn J, Põldmaa K, Rahimlou S, Rinaldi AC, Rojas O, Roslin T, Runnel K, Rähn E, Saba M, Saitta A, Salih TS, Sarapuu J, Serrano E, Serrano O, Sharmah D, Sharp C, Skalska-Tuomi MW, Tchan KI, Truong C, van der Merwe H, Vanié-Léabo LLP, Vasco-Palacios AM, Verbeken A, Vlk L, Wijayawardene NN, Wood JL, Yasanthika WAE, Yorou NS, Zahn G, Zettur I, Zucconi L, Kõljalg U, Tedersoo L. Global dataset of soil eukaryotic communities created with a uniform protocol and long read sequencing. Sci Data. 2026 May 5. doi: 10.1038/s41597-026-07315-y.

- [MIDORI2 - a collection of reference databases](https://doi.org/10.1002/edn3.303/)

> Leray, M., Knowlton, N., & Machida, R. J. (2022). MIDORI2: A collection of quality controlled, preformatted, and regularly updated reference databases for taxonomic assignment of eukaryotic mitochondrial sequences. Environmental DNA, 4, 894– 907. doi: https://doi.org/10.1002/edn3.303.
Expand Down
82 changes: 82 additions & 0 deletions bin/taxref_reformat_glosed.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,82 @@
#!/bin/sh

# Reformat GloSED references for assignTaxonomy and addSpecies.
# Inputs (in current directory by default):
# - GloSED__OTU_sequences.fasta.gz
# - GloSED__Taxonomy.tsv.zip
# Outputs:
# - assignTaxonomy.fna
# - addSpecies.fna

set -eu

SEQ_GZ="${1:-GloSED__OTU_sequences.fasta.gz}"
TAX_ZIP="${2:-GloSED__Taxonomy.tsv.zip}"

[ -f "$SEQ_GZ" ] || { echo "ERROR: missing $SEQ_GZ" >&2; exit 1; }
[ -f "$TAX_ZIP" ] || { echo "ERROR: missing $TAX_ZIP" >&2; exit 1; }

TMPDIR_LOCAL="$(mktemp -d glosed_reformat.XXXXXX)"
trap 'rm -rf "$TMPDIR_LOCAL"' EXIT INT TERM

TAX_TSV="${TMPDIR_LOCAL}/glosed_taxonomy.tsv"
META_TSV="${TMPDIR_LOCAL}/glosed_meta.tsv"

unzip -p "$TAX_ZIP" > "$TAX_TSV"

awk -F '\t' 'BEGIN { OFS="\t" }
NR==1 {
for (i = 1; i <= NF; i++) {
if ($i == "OTU") otu = i
else if ($i == "Kingdom") kingdom = i
else if ($i == "Phylum") phylum = i
else if ($i == "Class") classcol = i
else if ($i == "Order") ordercol = i
else if ($i == "Family") family = i
else if ($i == "Genus") genus = i
else if ($i == "Species") species = i
}
next
}
{
id = $otu
k = $kingdom; p = $phylum; c = $classcol; o = $ordercol; f = $family; g = $genus; s = $species
gsub(/ /, "_", k); gsub(/ /, "_", p); gsub(/ /, "_", c); gsub(/ /, "_", o); gsub(/ /, "_", f); gsub(/ /, "_", g); gsub(/ /, "_", s)

tax = ""
if (k != "." && k != "") {
tax = k ";"
if (p != "." && p != "") {
tax = tax p ";"
if (c != "." && c != "") {
tax = tax c ";"
if (o != "." && o != "") {
tax = tax o ";"
if (f != "." && f != "") {
tax = tax f ";"
if (g != "." && g != "") {
tax = tax g ";"
if (s != "." && s != "") {
tax = tax s ";"
}
}
}
}
}
}
}
print id, tax, g, s
}
' "$TAX_TSV" > "$META_TSV"

gzip -dc "$SEQ_GZ" | awk -F '\t' 'NR==FNR { tax[$1] = $2; next }
/^>/ { id = substr($0,2); if (id in tax && tax[id] != "") print ">" tax[id]; else print ">" id; next }
{ print }
' "$META_TSV" - > assignTaxonomy.fna

gzip -dc "$SEQ_GZ" | awk -F '\t' 'NR==FNR { if ($3 != "." && $3 != "" && $3 != "NA" && $4 != "." && $4 != "" && $4 != "NA") addsp[$1] = $3 " " $4; next }
/^>/ { id = substr($0,2); keep = (id in addsp); if (keep) print ">" id " " addsp[id]; next }
{ if (keep) print }
' "$META_TSV" - > addSpecies.fna

echo "Created files: assignTaxonomy.fna addSpecies.fna"
14 changes: 14 additions & 0 deletions conf/ref_databases.config
Original file line number Diff line number Diff line change
Expand Up @@ -95,6 +95,20 @@ params {
fmtscript = "taxref_reformat_gtdb.sh"
dbversion = "GTDB R05-RS95 (https://data.gtdb.ecogenomic.org/releases/release95/95.0/)"
}
'glosed' {
Comment thread
tom-brekke marked this conversation as resolved.
title = "GloSED - Global standardised Soil Eukaryome Dataset"
file = [ "https://zenodo.org/records/17827890/files/GloSED__OTU_sequences.fasta.gz", "https://zenodo.org/records/17827890/files/GloSED__Taxonomy.tsv.zip" ]
citation = "Sundh J, Larsson E, Nilsson RH, et al. GloSED fungal ITS database. Zenodo. https://zenodo.org/records/17827890"
fmtscript = "taxref_reformat_glosed.sh"
dbversion = "GloSED (https://zenodo.org/records/17827889)"
Comment thread
tom-brekke marked this conversation as resolved.
Outdated
}
'glosed=1.0.0' {
title = "GloSED - Global standardised Soil Eukaryome Dataset"
file = [ "https://zenodo.org/records/17827890/files/GloSED__OTU_sequences.fasta.gz", "https://zenodo.org/records/17827890/files/GloSED__Taxonomy.tsv.zip" ]
citation = "Sundh J, Larsson E, Nilsson RH, et al. GloSED fungal ITS database. Zenodo. https://zenodo.org/records/17827890"
fmtscript = "taxref_reformat_glosed.sh"
dbversion = "GloSED (https://zenodo.org/records/17827890)"
}
'midori2-co1' {
title = "MIDORI2 - CO1 Taxonomy Database - Release GB250"
file = [ "https://reference-midori.info/download/Databases/GenBank250_2022-06-13/DADA2_sp/uniq/MIDORI2_UNIQ_NUC_SP_GB250_CO1_DADA2.fasta.gz" ]
Expand Down
33 changes: 33 additions & 0 deletions conf/test_glosed.config
Original file line number Diff line number Diff line change
@@ -0,0 +1,33 @@
/*
========================================================================================
Nextflow config file for running minimal tests with GloSED database
========================================================================================
Defines input files and everything required to run a fast and simple pipeline test
using the GloSED (Global Standardised Soil Eukaryome Dataset) reference database.

Use as follows:
nextflow run nf-core/ampliseq -profile test_glosed,<docker/singularity> --outdir <OUTDIR>

----------------------------------------------------------------------------------------
*/

Comment thread
d4straub marked this conversation as resolved.
process {
resourceLimits = [
cpus: 4,
memory: '256.GB',
time: '2.h'
]
}

params {
config_profile_name = 'Test profile for GloSED database'
config_profile_description = 'Minimal test dataset to verify pipeline function with GloSED reference database'

// Input data
FW_primer = "GTGYCAGCMGCCGCGGTAA"
RV_primer = "GGACTACNVGGGTWTCTAAT"
input = params.pipelines_testdata_base_path + "ampliseq/samplesheets/Samplesheet.tsv"
dada_ref_taxonomy = "glosed"

skip_qiime = true
}
1 change: 1 addition & 0 deletions docs/usage.md
Original file line number Diff line number Diff line change
Expand Up @@ -274,6 +274,7 @@ Pre-configured reference taxonomy databases are:
| greengenes | - | - | + | (+)³ | - | - | 16S rRNA |
| greengenes2 | + | - | - | + | - | - | 16S rRNA |
| pr2 | + | - | - | - | - | - | 18S rRNA |
| GloSED | + | - | - | - | - | - | eukaryotic nuclear ribosomal ITS region
| unite-fungi | + | + | - | - | + | - | eukaryotic nuclear ribosomal ITS region |
| unite-alleuk | + | + | - | - | + | - | eukaryotic nuclear ribosomal ITS region |
| coidb | + | + | - | - | + | - | eukaryotic Cytochrome Oxidase I (COI) |
Expand Down
1 change: 1 addition & 0 deletions nextflow.config
Original file line number Diff line number Diff line change
Expand Up @@ -322,6 +322,7 @@ profiles {
test_pacbio_its { includeConfig 'conf/test_pacbio_its.config' }
test_iontorrent { includeConfig 'conf/test_iontorrent.config' }
test_fasta { includeConfig 'conf/test_fasta.config' }
test_glosed { includeConfig 'conf/test_glosed.config' }
test_failed { includeConfig 'conf/test_failed.config' }
test_full { includeConfig 'conf/test_full.config' }
aws_scratch { includeConfig 'conf/aws_scratch.config' }
Expand Down
6 changes: 4 additions & 2 deletions nextflow_schema.json
Original file line number Diff line number Diff line change
Expand Up @@ -455,12 +455,13 @@
"properties": {
"dada_ref_taxonomy": {
"type": "string",
"help_text": "Choose any of the supported databases, and optionally also specify the version. Database and version are separated by an equal sign (`=`, e.g. `silva=138`) . This will download the desired database, format it to produce a file that is compatible with DADA2's assignTaxonomy and another file that is compatible with DADA2's addSpecies.\n\nThe following databases are supported:\n- GTDB - Genome Taxonomy Database - 16S rRNA\n- SBDI-GTDB, a Sativa-vetted version of the GTDB 16S rRNA\n- PR2 - Protist Reference Ribosomal Database - 18S rRNA\n- RDP - Ribosomal Database Project - 16S rRNA\n- SILVA ribosomal RNA gene database project - 16S rRNA\n- UNITE - eukaryotic nuclear ribosomal ITS region - ITS\n- COIDB - eukaryotic Cytochrome Oxidase I (COI) from The Barcode of Life Data System (BOLD) - COI\n\nGenerally, using `gtdb`, `pr2`, `rdp`, `sbdi-gtdb`, `silva`, `coidb`, `unite-fungi`, or `unite-alleuk` will select the most recent supported version.\n\nPlease note that commercial/non-academic entities [require licensing](https://www.arb-silva.de/silva-license-information) for SILVA v132 database (non-default) but not from v138 on (default).",
"help_text": "Choose any of the supported databases, and optionally also specify the version. Database and version are separated by an equal sign (`=`, e.g. `silva=138`) . This will download the desired database, format it to produce a file that is compatible with DADA2's assignTaxonomy and another file that is compatible with DADA2's addSpecies.\n\nThe following databases are supported:\n- GTDB - Genome Taxonomy Database - 16S rRNA\n- SBDI-GTDB, a Sativa-vetted version of the GTDB 16S rRNA\n- PR2 - Protist Reference Ribosomal Database - 18S rRNA\n- RDP - Ribosomal Database Project - 16S rRNA\n- SILVA ribosomal RNA gene database project - 16S rRNA\n- UNITE - eukaryotic nuclear ribosomal ITS region - ITS\n- COIDB - eukaryotic Cytochrome Oxidase I (COI) from The Barcode of Life Data System (BOLD) - COI\n- GloSED - Global Standardised Soil Eukaryome Dataset - fungal ITS (Note: SH support not available)\n\nGenerally, using `gtdb`, `pr2`, `rdp`, `sbdi-gtdb`, `silva`, `coidb`, `unite-fungi`, or `unite-alleuk` will select the most recent supported version.\n\nPlease note that commercial/non-academic entities [require licensing](https://www.arb-silva.de/silva-license-information) for SILVA v132 database (non-default) but not from v138 on (default).",
"description": "Name of supported database, and optionally also version number",
"default": "sbdi-gtdb=R11-RS232-1",
"enum": [
"coidb",
"coidb=221216",
"glosed",
Comment thread
tom-brekke marked this conversation as resolved.
"greengenes2",
"greengenes2=2024.09",
"gtdb",
Expand Down Expand Up @@ -742,7 +743,8 @@
},
"addsh": {
"type": "boolean",
"description": "If ASVs should be assigned to UNITE species hypotheses (SHs). Only relevant for ITS data."
"description": "If ASVs should be assigned to UNITE species hypotheses (SHs). Only relevant for ITS data.",
"help_text": "Requires a DADA2 reference database with precomputed SH lookup files. Currently supported for UNITE reference databases. Not currently supported for `--dada_ref_taxonomy glosed`."
},
"cut_its": {
"type": "string",
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -315,7 +315,7 @@ def validateInputParameters() {
validDBs += " " + db
}
}
error("UNITE species hypothesis information is not available for the selected reference database, please use the option `--dada_ref_taxonomy` to select an appropriate database. Currently, the option `--addsh` can only be used together with the following UNITE reference databases:\n" + validDBs + ".")
error("Species hypothesis (SH) lookup files are not available for `--dada_ref_taxonomy ${params.dada_ref_taxonomy}`. This currently includes `glosed`. The option `--addsh` can only be used with databases that provide precomputed SH lookup files (currently UNITE reference databases):\n" + validDBs + ".")
}

if (params.addsh && params.cut_its == "none") {
Expand Down
47 changes: 47 additions & 0 deletions tests/glosed.nf.test
Original file line number Diff line number Diff line change
@@ -0,0 +1,47 @@
nextflow_pipeline {

name "Test pipeline with GloSED"
script "../main.nf"
tag "pipeline"
profile "test_glosed"

test("-profile test_glosed") {

when {
params {
outdir = "$outputDir"
}
}

then {
// stable_name: All files + folders in ${params.outdir}/ with a stable name
def stable_name = getAllFilesFromDir(params.outdir, relative: true, includeDir: true, ignore: ['pipeline_info/*.{html,json,txt}'])

assert workflow.success
assertAll(
{ assert snapshot(
// pipeline versions.yml file for multiqc from which Nextflow version is removed because we test pipelines on multiple Nextflow versions
removeNextflowVersion("$outputDir/pipeline_info/software_versions.yml"),
// All stable path name, with a relative path
stable_name,
// Manually chosen files for content checks
path("$outputDir/overall_summary.tsv"),
path("$outputDir/barrnap/rrna.arc.gff"),
path("$outputDir/barrnap/rrna.bac.gff"),
path("$outputDir/barrnap/rrna.euk.gff"),
path("$outputDir/barrnap/rrna.mito.gff"),
path("$outputDir/cutadapt/cutadapt_summary.tsv"),
path("$outputDir/dada2/ASV_seqs.fasta"),
path("$outputDir/dada2/ASV_table.tsv"),
path("$outputDir/dada2/ref_taxonomy.glosed.txt"),
path("$outputDir/dada2/DADA2_stats.tsv"),
path("$outputDir/dada2/DADA2_table.tsv"),
path("$outputDir/input/Samplesheet.tsv"),
path("$outputDir/multiqc/multiqc_data/multiqc_fastqc.txt"),
path("$outputDir/multiqc/multiqc_data/multiqc_general_stats.txt"),
path("$outputDir/multiqc/multiqc_data/multiqc_cutadapt.txt")
).match() }
)
}
}
}
Loading
Loading