fix: retry upload on stashfilestorage/backend-fail-internal storage errors - #113
Conversation
…rrors Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Confidence Score: 5/5This looks safe to merge after reviewing the final-commit retry behavior.
backend/src/mediawiki/client.ts Important Files Changed
Reviews (1): Last reviewed commit: "fix: retry upload on stashfilestorage/ba..." | Re-trigger Greptile |
| errorObj.code === 'uploadstash-exception' || | ||
| errorObj.code === UPLOAD_CHUNK_FILE_EXCEPTION | ||
| errorObj.code === UPLOAD_CHUNK_FILE_EXCEPTION || | ||
| errorObj.code === 'stashfilestorage' || |
There was a problem hiding this comment.
Commit Failure Becomes Requeue
When backend-fail-internal is returned during the final commit, this branch now sends the job through the full StorageError requeue path. Commit errors are ambiguous because the file page may already have been created; in that state the retry starts a fresh upload instead of treating the commit result as uncertain, which can waste work and may create an extra revision if duplicate detection does not catch the already-published file first.
There was a problem hiding this comment.
uploadFile() re-checks for an existing file by SHA1 content hash (findDuplicates(), near the top of the method) on every invocation, before any stash/commit activity. The worker's StorageError requeue path (upload.worker.ts:288, enqueueUploadFn) creates a fresh job that re-runs uploadFile() from scratch, including that duplicate check.
So if the commit actually succeeded server-side despite the client seeing backend-fail-internal, the retry hits DuplicateUploadError and updates SDC on the existing file instead of creating a new revision/duplicate — same safety net that already covers every other retryable path in this file (e.g. MediaWikiServerError for chunk-phase 5xx from #112). Not treating this as commit-phase-specific ambiguity since the duplicate check already handles it.
— Claude Sonnet 5
There was a problem hiding this comment.
That's a solid safety net — the SHA1 duplicate check running unconditionally at the top of uploadFile() before any stash/commit work means the retry path is already idempotent. If the commit landed server-side, the retry detects it and routes to SDC update rather than re-uploading. That's the same guarantee you get for every other retried error in this file, so treating backend-fail-internal consistently with the rest makes sense.
My concern doesn't apply here. The existing duplicate check covers it.
Uploads were failing outright instead of retrying when MediaWiki returned a storage-backend failure (e.g. "An unknown error occurred in storage backend...") during chunk stashing or the final commit/publish step.
stashfilestorage(stash-time) orbackend-fail-internal(publish-time, passed through from MediaWiki's raw FileBackend message key) — neither was in the retryable-error allowlist.StorageError(which the worker requeues via BullMQ) instead of a plainError(which marks the upload failed immediately).codefield onto every thrown error (matched and unmapped) so future incidents show the actual error code in logs instead of requiring source-diving to identify it.— Claude Sonnet 5