Skip to content

fs: per-file fsync in sparse-aware copyFile causes 17x regression for many-files workloads #303

Description

@pengubco

Problem

Commit a424ba1 ("Linux: Make copyFile sparse aware") introduced tgt.Sync() at the end of the new copyFile implementation in fs/copy_linux.go. While the sparse-aware copy logic itself is valuable, the per-file fsync causes a severe performance regression when copying directories with many files.

In my use case, this caused significant container start latency when using Docker-style volumes (VOLUME directive in Dockerfile). The container runtime's volumeCopyUp path uses fs.CopyDir to copy image content into the volume mount at container start. When the image has thousands of files in the volume path, the per-file fsync serializes all I/O and adds more than one minute of startup delay. This delay has caused availability issue and I had to bump the timeout to mitigate.

Benchmark Results

50,000 files × 10 KiB each, ext4, cold page cache:

Method Time
Old behavior (io.Copy, no fsync) 4.4s
New sparse-aware, no fsync 4.8s
New sparse-aware, single sync at end 4.9s
New sparse-aware, per-file fsync (current) 1m 16s

The per-file fsync makes the copy 17x slower. The sparse-aware logic itself (SEEK_DATA/SEEK_HOLE + copy_file_range) adds only ~10% overhead vs the old io.Copy path.

Full writeup and reproduction code: https://github.com/pengubco/ad-hoc-continuity-0.5.0-fsync-regression

Impact

This affects any code path that uses fs.CopyDir or fs.CopyFile for directories with many files:

  • Volume copy-up — containerd CRI plugin copies image content into volume mounts at container start (volumeCopyUp). Related: #8639 in containerd/containerd.
  • nerdctl cp — copying files into/out of containers.
  • Snapshot operations — any directory tree replication using continuity/fs.

Why per-file fsync is unnecessary

The old openAndCopyFile (io.Copy) never called fsync. For the primary callers of CopyDir:

  • Volume copy-up: if the machine crashes, the container never started — the runtime redoes the copy from the immutable image layer on restart.
  • Image extraction: interrupted extractions are discarded and re-extracted from the content store.

None of these paths require per-file durability. The data source is always available for replay.

Suggestion

Remove tgt.Sync() from copyFile. If durability before return is desired for some callers, a single sync() at the end of CopyDir provides equivalent crash safety with <1% overhead (vs 17x with per-file). Alternatively, make it opt-in via a CopyDirOpt.

Environment

  • Linux (ext4)
  • Go 1.26
  • continuity at commit a424ba1
  • AWS EC2 with EBS volumes

Activity

  1. ayush-panta commented on Jul 20, 2026

    @ayush-panta
    Contributor

    I'm apprehensive to removing the sync. This PR links the change to the blockfile snapshotter, and it appears the blockfile in containerd depends on copyFileWithSync(), which uses a file-level sync.

    @AkihiroSuda @dcantah I'd like your opinion on this. Is there a need for granular file or dir-level sync? If so, I have an implementation with options for both of these.

  2. AkihiroSuda commented on Jul 20, 2026

    @AkihiroSuda
    Member

    For compatibility we should keep the current behavior by default but new options can be added to https://pkg.go.dev/github.com/containerd/continuity@v0.5.0/fs#CopyDirOpt

    There should be CopyFileOpt too.

  3. ayush-panta commented on Jul 20, 2026

    @ayush-panta
    Contributor

    I'd like to confirm whether sync should be the default.

    It appears that in v0.4.5 and earlier Linux used openAndCopyFile(), which does not have tgt.Sync(). The change in v0.5.0 effectively added sync to every fs.copyFile, which can cause increased latency for general file operations. If we make sync default, then we have to somehow make users of this function explicitly opt-out.

    My current implementation uses a variadic copyFile with CopyDirOpt and CopyFileOpt, both assuming no sync by default, allowing most code paths to go back to how they were before v0.4.5. We can update the blockfile copy invocation in containerd to opt-in to sync.

    Please let me know if I'm misunderstanding something, and I can flip my implementation.

  4. pengubco commented on Jul 21, 2026

    @pengubco
    ContributorAuthor

    I confirm that v0.4.5 and earlier, there is no fsync. That's why my workload worked on v0.4.5 but timeout after upgrade to v0.5.0.

  5. ayush-panta commented on Jul 27, 2026

    @ayush-panta
    Contributor

    @AkihiroSuda sorry to ping again, but could we get more opinions on this? PR is linked above. This is related to some performance regressions in AWS EC2 with EBS volumes, as Peng showed previously.

  6. AkihiroSuda commented on Jul 28, 2026

    @AkihiroSuda
    Member

    Approved #310

  7. ayush-panta commented on Jul 29, 2026

    @ayush-panta
    Contributor

    Is there a plan for a continuity release soon? I'd like these changes to be available so that we may test containerd 2.3 internally, as 2.3.x currently requires continuity 0.5.0.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions