Accelerated computation of seismic probabilistic power spectral densities (PPSD) over long time periods using Python multiprocessing.
Multiprocessing spectra is a Python-based workflow designed to automate and accelerate the computation of seismic spectra over large datasets spanning multiple days or months of continuous seismic recordings. The project was developed to address a common bottleneck in seismic noise characterization: the processing of large volumes of waveform data for Power Spectral Density (PSD) and Probabilistic Power Spectral Density (PPSD) analyses. By leveraging Python's multiprocessing capabilities, the workflow distributes independent computations across multiple CPU cores, significantly reducing execution time compared to a sequential implementation.
Characterizing the seismic environment is a fundamental task in many scientific and engineering applications, including:
- Site characterization for large-scale scientific infrastructures;
- Long-term seismic noise monitoring;
- Environmental noise assessment;
- Seismological studies;
- Detector site selection and qualification.
Probabilistic Power Spectral Density (PPSD) analysis provides a statistical description of seismic noise levels as a function of frequency and time and is widely used within the seismological community. Given a seismic time series (x(t)), the power spectral density is estimated through Fourier-based techniques and accumulated over many time intervals to obtain a probabilistic representation of the seismic background. For large archives containing months or years of continuous data, the computational cost becomes significant. This repository addresses that challenge through parallel processing.
- Multiprocessing-based execution.
- Automatic distribution of independent spectral calculations across available CPU cores.
- Efficient processing of long-duration seismic datasets.
- Waveform acquisition utilities.
- Support for structured seismic archives.
- Batch processing of multiple days of data.
- Power Spectral Density (PSD) estimation.
- Probabilistic Power Spectral Density (PPSD) generation.
- Percentiles of PPSD
- Long-term seismic noise characterization.
- Automated processing pipeline.
- Configurable input parameters.
- Structured output products.
- Download of seismic data files in batched from servers
flowchart TD
A[Download Waveforms]
B[Load Metadata]
C[Split Data by Day]
D[Compute PSD Segments]
E[Parallel Processing]
F[Build PPSD Statistics]
G[Generate Outputs]
A --> B
B --> C
C --> D
D --> E
E --> F
F --> G
Since seismic data files needed for spectra calculation are too heavy for being uploaded on a Git Repository, this repository includes utilities for waveform retrieval.
Run python scripts/data_download.py to download seismic data in the data_sds folder already organized according to the SDS protocol needed to read them.
The parameters of the time period and seismic stations can be changed inside the file.
Default values are already set.
Run the parallel processing workflow, e.g.:
python ./scripts/python PPSD_daily_spectra_parallel_v_2.3.3.py -source local -client /archive/AMIATA -net ET -sta P2 -loc 01 -datetime 2023-07-25T00:00 -days 2 -T_fft 128 -n_cpus 12
The script automatically distributes calculations across multiple processes and stores the resulting spectra in the output directory.
The workflow requires:
- Waveform Data
- Continuous seismic recordings stored in an SDS archive structure.
- Station Metadata (Instrument response information describing: Station code; Network code; Sensor characteristics; Response functions).
All input data are provided by scripts/data_download.py.
Depending on the selected configuration, the workflow generates:
- Output files with the percentiles of the PPSD distributions;
- Frequency-dependent seismic noise levels;
- PPSD plots (deprecated);
Outputs are stored in the output/ directory.
The project relies on the scientific Python ecosystem and typical seismological analysis libraries. Main dependencies include:
- Python ≥ 3.x
- NumPy
- SciPy
- Matplotlib
- ObsPy
- Multiprocessing