Skip to content

fix: retry upload on stashfilestorage/backend-fail-internal storage errors - #113

Merged
DaxServer merged 1 commit into
mainfrom
fix/retry-storage-backend-errors
Jul 5, 2026
Merged

DaxServer merged 1 commit into
mainfrom
fix/retry-storage-backend-errors

Conversation

@DaxServer

Copy link
Copy Markdown
Owner

Uploads were failing outright instead of retrying when MediaWiki returned a storage-backend failure (e.g. "An unknown error occurred in storage backend...") during chunk stashing or the final commit/publish step.

  • MediaWiki returns these errors as HTTP 200 with the failure embedded in the JSON error object, so the fix: retry MediaWiki upload on 5xx instead of failing immediately #112 5xx-status retry check never sees them.
  • The API error code for this class of failure is stashfilestorage (stash-time) or backend-fail-internal (publish-time, passed through from MediaWiki's raw FileBackend message key) — neither was in the retryable-error allowlist.
  • Added both codes to the chunk-phase and commit-phase checks so they now throw StorageError (which the worker requeues via BullMQ) instead of a plain Error (which marks the upload failed immediately).
  • Also attached the MediaWiki code field onto every thrown error (matched and unmapped) so future incidents show the actual error code in logs instead of requiring source-diving to identify it.

— Claude Sonnet 5

…rrors

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@greptile-apps

greptile-apps Bot commented Jul 5, 2026

Copy link
Copy Markdown
Contributor

Confidence Score: 5/5

This looks safe to merge after reviewing the final-commit retry behavior.

  • No blocking issues found in the changed code.
  • The only noted concern is a conditional retry edge case after an ambiguous final commit response.

backend/src/mediawiki/client.ts

Important Files Changed

Filename Overview
backend/src/mediawiki/client.ts Maps new MediaWiki storage error codes to StorageError and preserves raw error codes on thrown errors.
backend/src/tests/uploadClient.test.ts Adds focused tests for the new retryable upload error codes and code propagation.

Reviews (1): Last reviewed commit: "fix: retry upload on stashfilestorage/ba..." | Re-trigger Greptile

errorObj.code === 'uploadstash-exception' ||
errorObj.code === UPLOAD_CHUNK_FILE_EXCEPTION
errorObj.code === UPLOAD_CHUNK_FILE_EXCEPTION ||
errorObj.code === 'stashfilestorage' ||

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Commit Failure Becomes Requeue

When backend-fail-internal is returned during the final commit, this branch now sends the job through the full StorageError requeue path. Commit errors are ambiguous because the file page may already have been created; in that state the retry starts a fresh upload instead of treating the commit result as uncertain, which can waste work and may create an extra revision if duplicate detection does not catch the already-published file first.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

uploadFile() re-checks for an existing file by SHA1 content hash (findDuplicates(), near the top of the method) on every invocation, before any stash/commit activity. The worker's StorageError requeue path (upload.worker.ts:288, enqueueUploadFn) creates a fresh job that re-runs uploadFile() from scratch, including that duplicate check.

So if the commit actually succeeded server-side despite the client seeing backend-fail-internal, the retry hits DuplicateUploadError and updates SDC on the existing file instead of creating a new revision/duplicate — same safety net that already covers every other retryable path in this file (e.g. MediaWikiServerError for chunk-phase 5xx from #112). Not treating this as commit-phase-specific ambiguity since the duplicate check already handles it.

— Claude Sonnet 5

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That's a solid safety net — the SHA1 duplicate check running unconditionally at the top of uploadFile() before any stash/commit work means the retry path is already idempotent. If the commit landed server-side, the retry detects it and routes to SDC update rather than re-uploading. That's the same guarantee you get for every other retried error in this file, so treating backend-fail-internal consistently with the rest makes sense.

My concern doesn't apply here. The existing duplicate check covers it.

@DaxServer
DaxServer merged commit 9fc5868 into main Jul 5, 2026
5 checks passed
@DaxServer
DaxServer deleted the fix/retry-storage-backend-errors branch July 5, 2026 11:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant