PDDL planning benchmark suites — classical, numeric, and profiling — published as the pypddl-datasets
Python package. The package itself is small: benchmark data is downloaded on
first use from the matching GitHub release and cached locally.
pip install pypddl-datasetsimport pypddl_datasets as pb
pb.list_suites() # ['autoscale-agile-strips', ..., 'ipc-optimal-strips', ...]
task = pb.fetch_task("classical/tests/gripper/test-1.pddl")
task.domain_path # .../gripper/domain.pddl (correct also where instances
task.task_path # .../gripper/test-1.pddl carry their own domain files)
task.domain, task.problem # "classical-tests-gripper", "test-1.pddl" — lab-safe display names
domain = pb.fetch_domain("classical/downward-benchmarks/gripper")
domain.path # the domain directory
domain.tasks # list[Task], downloaded once and cached
suite = pb.fetch_suite("ipc-optimal-strips") # Suite(path, domains)
for domain in suite.domains:
for task in domain.tasks:
run_planner(task.domain_path, task.task_path)Most suites have a -test companion (e.g. "ipc-optimal-strips-test") whose
entries are one representative task per domain — a cheap smoke run before
committing to a full suite. pb.export_suite(suite, dest) materializes a
suite as a plain directory tree for non-Python tools.
Domains can be filtered by their declared PDDL requirements — supported is
a capability ceiling (keep what your planner handles), requires a feature
floor (keep what exercises a feature). The data declares exactly the atomic
requirements each file uses (strict-validated; aggregates like :adl never
appear), and all queries are metadata-only (no download):
from pypddl_datasets import Requirement as R
SUPPORTED = {R.STRIPS, R.TYPING, R.ACTION_COSTS, R.NEGATIVE_PRECONDITIONS}
pb.task_requirements("classical/tests/gripper/test-1.pddl") # frozenset({R.STRIPS})
pb.domain_requirements("classical/tests/gripper") # union over the domain's tasks
pb.find_tasks(requires={R.CONDITIONAL_EFFECTS}) # task names, per-task precision
pb.find_domains(suite="ipc-satisficing-strips", supported=SUPPORTED)
pb.find_suites(supported=SUPPORTED) # suites runnable in full
pb.fetch_suite("ipc-satisficing-strips", supported=SUPPORTED) # filtered fetchpb.list_domains() lists every individually fetchable domain. The cache
location defaults to the platform cache dir and can be overridden with the
PYPDDL_DATASETS_CACHE environment variable. On machines without internet
access, set PYPDDL_DATASETS_DATA to a local checkout's data/ directory
and domains resolve there without downloading.
src/pypddl_datasets/— the package: fetch API, suite definitions, and the instance generators (make_problem+ CLI per domain; choosing train/valid/test splits is left to the user).data/— all benchmark data, organized as<formalism>/<collection>/<domain>(classical/,numeric/). Not shipped in the package; released as a single archive ondata-v*GitHub releases, downloaded and unpacked once per machine on first use.pypddl_datasets.scripts— repository tooling, importable in a checkout but never shipped in the wheel:package_data(byte-reproducibledata.tar.gz),extract_requirements(regenerates the committedrequirements.{tasks,domains,suites}.json),strict_clean(mechanical requirements-declaration repair). Run withpython -m pypddl_datasets.scripts.<name>.pypddl_datasets.validation— data checks; must pass for a data release to go out:python -m pypddl_datasets.validationchains the layout check (flat domain directories, every problem pairs), the suite-configuration check (every SUITES entry resolves,-testsuites select from their base, benchmark suites have a-testcompanion — a one-instance-per-domain miniature for dry-running experiment pipelines), and the PDDL content check (parses everything with pypddl);validation.requirementsguards metadata freshness at release time.
The dev dependency group pins the test and lint tools (pytest, pypddl,
pyright, pylint). Install it into .venv (for example uv sync --group dev),
then run what CI runs:
.venv/bin/pytest
.venv/bin/pyright # strict, pyrightconfig.json
.venv/bin/pylint src/pypddl_datasets tests # config in pyproject.tomlGenerator changes must keep the output byte-identical for every seed unless the change is meant to alter the distribution.
Releases run from the Actions "release" workflow (Run workflow); tags are
outputs of the workflow, never triggers — pushing v* or data-v* tags by
hand publishes nothing.
- scope —
packagepublishes a new package version.data-and-packagefirst validates the data (layout, strict PDDL content, metadata freshness), uploads the byte-reproducibledata.tar.gzto a new immutabledata-v<N>GitHub release, and commits the pin (DATA_VERSIONandDATA_SHA256insrc/pypddl_datasets/fetching.py) to main. - bump —
patchorminor. The workflow bumps__version__insrc/pypddl_datasets/__init__.py(the single version source; pyproject reads it dynamically), independently re-verifies the pinned data release against the actual GitHub asset, builds, checks the wheel contents, commits + tagsv<version>, and publishes to PyPI via trusted publishing. A final job installs the published package on a clean runner and fetches a task through a fresh cache — the full user path, end to end. - dry_run — rehearses all gates, packaging, and builds with no tags, commits, uploads, or publishing.
Data releases are permanent: published package versions pin them by tag and
sha256, so never delete a data-v* release.
GPL-3.0-or-later (LICENSES/GPL-3.0-or-later.txt),
except seventeen generator modules ported from Jörg Hoffmann's FF domain
collection. They keep their original notice and may be used for non-commercial
research only (LICENSES/LicenseRef-Freiburg.txt);
each carries SPDX-License-Identifier: LicenseRef-Freiburg, and
CREDITS.md lists them together with
the authors of all domains and generators. The benchmark data is not part of
the package and keeps the terms of its sources.
Oversized PDDL files (>= 50 MiB) are committed as gzipped .pddl.gz twins
and materialized locally (the plain files are gitignored). After cloning:
pip install -e .
python -m pypddl_datasets.scripts.large_files unpackWhen adding files that large, python -m pypddl_datasets.scripts.large_files pack creates the twins and updates the managed .gitignore block — the
layout validation refuses anything oversized left unpacked, so CI will tell
you.
Domain directories must contain their .pddl files directly, with no
subdirectories (that is how discovery and domain/problem pairing work), and
domain names flattened with / → - must stay unique (Task.domain;
test-guarded).
Before opening a pull request, run the same checks the CI and the data release gate run:
pip install 'pypddl>=1.0.27,<1.1' -e .
python -m pypddl_datasets.validation --root data --strict # layout + PDDL content, same as the CI gate
pytest tests # suite definitions stay consistentData pull requests do not need to update the shared requirements.*.json
metadata. Before a data release, regenerate and commit it with
python -m pypddl_datasets.scripts.extract_requirements --data-root data;
the release workflow only verifies that it is fresh.