Repository navigation
Manual
- Introduction
- System Requirements
- Installation
- Eukfinder Database
- Using Eukfinder to classifiy short read data
- Using Eukfinder to classifiy long read data or assembled contigs
- Custom Databases
- Masking Low-Complexity Sequences
Eukfinder is designed for the classification of eukaryotic sequences in metagenomic data. It processes Illumina short reads (Eukfinder_short) and assemblies/long-read data (Eukfinder_long) through automated classification steps. Users can also apply an additional manual binning workflow to refine nuclear and mitochondrial genomes.
Features:
- reference-independent and cultivation-free
- separate reads or contigs into five groups: Bacteria, Archaeal, Eukaryotic, Viral, and Unknown.
- generate a fasta file with Euk and unknown contigs for binning
Warning! Eukfinder is not deadlock-safe, please avoid running parallel instances!
The resource requirements for this pipeline will vary greatly based on the amount of data being processed, but due to large memory requirements of many software used (Centrifuge and metaSPAdes to name a few), we recommend the following minimal setup.
Minimum required: ~200 GB
Recommended: ≥ 300 GB if using full Centrifuge and PLAST databases
Breakdown:
Centrifuge database: 100–150 GB
PLAST database: 6–15 GB
Temporary files (assembly, classification, binning): varies depending on sample size (expect 5–50 GB per sample)
Output folders for each sample (classified reads, bins, intermediate files): varies
Minimum required: 32 GB
Recommended: 64–128 GB for larger metagenomic samples
Some components like metaSPAdes and Centrifuge benefit greatly from higher RAM, especially during classification and assembly steps.
All required software can be installed via the Eukfinder conda package (see Installation Instructions).
✅ Recommended: Install using conda install -c bioconda eukfinder to automatically handle all dependencies.
Manual Installation (Not Recommended) If you prefer to install the environment manually, you’ll need to install the following tools and libraries individually:
- python==3.12.8
- numpy
- pandas
- joblib
- pyqt=5
- spades
- seqkit
- trimmomatic
- centrifuge
- bowtie2
- pip
- plast
- ete3==3.1.3
MacOS NOTE: MacOS and other non-Linux operating systems are not explicitly supported by the developers.
Required to:
-
Download databases from the Eukfinder server (e.g., Centrifuge, PLAST, host genome)
-
Install packages via conda and pip
-
Optional: access NCBI or other public taxonomic databases during custom DB building
Tip: Use a high-speed connection to download the databases efficiently. The full set may take several hours on slower networks.
To begin using Eukfinder, you will first need to install it, and then either download or create a database.
Anaconda or miniconda required
Anaconda or Miniconda must be installed to run this script.
If you don’t already have Anaconda or Miniconda installed, you can follow these links to download and install them:
conda create -n eukfinder -c bioconda eukfinder
Default reference databases can be downloaded from Eukfinder Databases
- Plast Database
- Centrifuge Database
- Human Genome for read decontamination
- Read Adapters for Illumina sequencing
eukfinder download_db
Users can flexibly customize the reference data (see here)
If you have Conda 4.4 or later:
conda init
conda activate eukfinder
After this, you can run Eukfinder commands.
If you have Conda prior to 4.4:
source activate eukfinder
Then run Eukfinder commands.
Eukfinder requires two separate databases, one for Centrifuge and one for PLAST.
Default reference databases can be downloaded from Eukfinder Databases
For more spcecifics see Download or build databases section.
Quick start guide, for details on variables and how to run see Eukfinder_with_Illumina_short_reads page
eukfinder read_prep
Eukfinder read_prep [-h] --r1 R1 --r2 R2 -n THREADS -i ILLUMINA_CLIP
--hcrop HCROP -l LEADING_TRIM -t TRAIL_TRIM --wsize WSIZE
--qscore QSCORE --mlen MLEN --mhlen MHLEN
--hg HG -o OUT_NAME --cdb CDB
eukfinder short_seqs
Eukfinder short_seqs [-h] --r1 R1 --r2 R2 --un UN -o OUT_NAME -n NUMBER_OF_THREADS
-z NUMBER_OF_CHUNKS -t TAXONOMY_UPDATE
-p PLAST_DATABASE -m PLAST_ID_MAP --cdb CDB
-e E_VALUE --pid PID --cov COV --max_m MAX_M
--mhlen MHLEN --pclass PCLASS --uclass UCLASS
Quick start guide, for details on variables and how to run see Eukfinder_with_Long_read_or_contig_data page
eukfinder long_seqs
Eukfinder long_seqs [-h] -l LONG_SEQS -o OUT_NAME --mhlen MHLEN --cdb CBD
-n NUMBER_OF_THREADS -z NUMBER_OF_CHUNKS
-t TAXONOMY_UPDATE -p PLAST_DATABASE -m PLAST_ID_MAP
-e E_VALUE --pid PID --cov COV
We realize the standard database may not suit everyone's needs.
Eukfinder also allows creation of customized databases.
Instructions on how to Build custom DBs can be found here