Commit Graph

2 Commits

Author SHA1 Message Date
ignacionelson c77d80309e Harden the background orphan import from #1809
Found in review, none of them reachable in our shipped setups but each
cheap to close:

- Two chunks could adopt the same path when more than one worker runs
  the default queue: a run that stalls unblocks a new one after five
  minutes, and the old chain can resume beside it. Two rows on one set of
  bytes means deleting either deletes the other's file. Each path is now
  claimed under a cache lock and checked for a row inside it, so a path
  another chunk holds is left to it. A lock around the whole chunk was
  tried first and dropped: a chunk queues the next one while it still
  holds the lock, so the next one was discarded and the run died.
- A chunk now checks that the account that started the run is still
  active, still staff and still holds import_orphans. A run can outlast
  that access, and every chunk adopts files in that person's name.
- A failure shows a plain sentence and sends the exception to the log.
  A storage error can name a bucket, an endpoint or a path.
2026-10-05 02:27:18 -03:00
veenone 4b30849a88 Let "Import all" adopt every orphan the search matches, in a background job
The header checkbox on Import orphan files selected only the 25 rows on
screen, so an install with thousands of stray files had to import them a
page at a time. Once a whole page is ticked, the selection bar now offers
"Select all N matching files", and "Import all" takes every orphan the
search matches, on every page.

The import runs in a queued job because it is too slow for a request.
Each file is hashed in full and written in three commits, so 5,000 files
of 4 MB take about four minutes, and PHP stops a request after 30 s of
CPU, around file 1,100. ImportOrphanFilesJob works on the default queue in
chunks of about 45 s: each chunk rescans, imports what is still orphaned
and queues the next one. That keeps every job inside the worker's 60 s
timeout and the queue's 90 s retry_after, so no extra worker is needed,
and mail queued in the meantime goes out between chunks. If a run dies
part way, the next one picks up what is left.

Only one run can be active at a time. OrphanImportProgress keeps its state
in the cache and starts a run under a lock. While a run is active, every
other import is refused, the per-row button included, so no file is
adopted twice. The page polls files/orphans/import-status every 3 s and
shows the run as running, finished, failed with the reason, or stalled
after 5 minutes without progress, which usually means no worker is
listening.

Bulk delete still works one page at a time. The adoption itself moved to
OrphanFileImporter so the request and the job share it, and the rule for
what can be imported now lives in OrphanFileScanner::importable().
2026-10-05 06:35:46 +07:00