Skip to content
Jason Shao edited this page Oct 16, 2025 · 16 revisions

Table of Contents

  1. Introduction
  2. System Requirements
  3. Installation
  4. Eukfinder Database
  5. Using Eukfinder to classifiy short read data
  6. Using Eukfinder to classifiy long read data or assembled contigs
  7. Custom Databases
  8. Masking Low-Complexity Sequences

Introduction

Eukfinder is designed for the classification of eukaryotic sequences in metagenomic data. It processes Illumina short reads (Eukfinder_short) and assemblies/long-read data (Eukfinder_long) through automated classification steps. Users can also apply an additional manual binning workflow to refine nuclear and mitochondrial genomes.

Features:

  • reference-independent and cultivation-free
  • separate reads or contigs into five groups: Bacteria, Archaeal, Eukaryotic, Viral, and Unknown.
  • generate a fasta file with Euk and unknown contigs for binning

Warning! Eukfinder is not deadlock-safe, please avoid running parallel instances!

System Requirements

The resource requirements for this pipeline will vary greatly based on the amount of data being processed, but due to large memory requirements of many software used (Centrifuge and metaSPAdes to name a few), we recommend the following minimal setup.

Disk Space

Minimum required: ~200 GB

Recommended: ≥ 300 GB if using full Centrifuge and PLAST databases

Breakdown:

Centrifuge database: 100–150 GB

PLAST database: 6–15 GB

Temporary files (assembly, classification, binning): varies depending on sample size (expect 5–50 GB per sample)

Output folders for each sample (classified reads, bins, intermediate files): varies

Memory (RAM)

Minimum required: 32 GB

Recommended: 64–128 GB for larger metagenomic samples

Some components like metaSPAdes and Centrifuge benefit greatly from higher RAM, especially during classification and assembly steps.

Dependencies

All required software can be installed via the Eukfinder conda package (see Installation Instructions).

✅ Recommended: Install using conda install -c bioconda eukfinder to automatically handle all dependencies.

Manual Installation (Not Recommended) If you prefer to install the environment manually, you’ll need to install the following tools and libraries individually:

  • python==3.12.8
  • numpy
  • pandas
  • joblib
  • pyqt=5
  • spades
  • seqkit
  • trimmomatic
  • centrifuge
  • bowtie2
  • pip
  • plast
  • ete3==3.1.3

⚠️ Manual installation may lead to version compatibility issues. Use with caution.

MacOS NOTE: MacOS and other non-Linux operating systems are not explicitly supported by the developers.

Network connectivity

Required to:

  • Download databases from the Eukfinder server (e.g., Centrifuge, PLAST, host genome)

  • Install packages via conda and pip

  • Optional: access NCBI or other public taxonomic databases during custom DB building

Tip: Use a high-speed connection to download the databases efficiently. The full set may take several hours on slower networks.

Installation

To begin using Eukfinder, you will first need to install it, and then either download or create a database.

Anaconda or miniconda required

Anaconda or Miniconda must be installed to run this script.

If you don’t already have Anaconda or Miniconda installed, you can follow these links to download and install them:

1. Created environment and install eukfinder

conda create -n eukfinder -c bioconda eukfinder

2. Download or build databases

Default reference databases can be downloaded from Eukfinder Databases

  • Plast Database
  • Centrifuge Database
  • Human Genome for read decontamination
  • Read Adapters for Illumina sequencing
eukfinder download_db

Users can flexibly customize the reference data (see here)

3. Activate eukfinder environment before running the command

If you have Conda 4.4 or later:

conda init
conda activate eukfinder

After this, you can run Eukfinder commands.

If you have Conda prior to 4.4:

source activate eukfinder

Then run Eukfinder commands.

Eukfinder Databases

Eukfinder requires two separate databases, one for Centrifuge and one for PLAST.
Default reference databases can be downloaded from Eukfinder Databases

For more spcecifics see Download or build databases section.

Using Eukfinder to classify short read data

Quick start guide, for details on variables and how to run see Eukfinder_with_Illumina_short_reads page

eukfinder read_prep

    Eukfinder read_prep [-h] --r1 R1 --r2 R2 -n THREADS -i ILLUMINA_CLIP
                             --hcrop HCROP -l LEADING_TRIM -t TRAIL_TRIM --wsize WSIZE
                             --qscore QSCORE --mlen MLEN --mhlen MHLEN
                             --hg HG -o OUT_NAME --cdb CDB

eukfinder short_seqs

 Eukfinder short_seqs [-h] --r1 R1 --r2 R2 --un UN -o OUT_NAME -n NUMBER_OF_THREADS
                             -z NUMBER_OF_CHUNKS -t TAXONOMY_UPDATE 
                             -p PLAST_DATABASE -m PLAST_ID_MAP --cdb CDB
                             -e E_VALUE --pid PID --cov COV --max_m MAX_M
                             --mhlen MHLEN --pclass PCLASS --uclass UCLASS

Using Eukfinder to classify long read data or assembled contigs

Quick start guide, for details on variables and how to run see Eukfinder_with_Long_read_or_contig_data page

eukfinder long_seqs

 Eukfinder long_seqs [-h] -l LONG_SEQS -o OUT_NAME --mhlen MHLEN --cdb CBD
                            -n NUMBER_OF_THREADS -z NUMBER_OF_CHUNKS 
                            -t TAXONOMY_UPDATE -p PLAST_DATABASE -m PLAST_ID_MAP
                            -e E_VALUE --pid PID --cov COV

Custom Databases

We realize the standard database may not suit everyone's needs.
Eukfinder also allows creation of customized databases.

Instructions on how to Build custom DBs can be found here