Workflow improvement
Export a visual Google Docs review bundle.
Proposed implementation
Add an executable companion under examples/docs-review-bundle/ using existing gws subprocesses, no separate auth. Create a new nonexisting user-named directory within CWD. Export Docs JSON with all tabs, native PDF, DOCX and Markdown via Drive API, optionally comments (include only if retrieval succeeds; explicitly mark unavailable). Use relative export filenames with gws cwd inside bundle to work with existing path restrictions. Check source revision before/after exports and label mixed-version output instead of claiming atomic snapshot. Extract raster assets from DOCX zip safely with total/member size caps and without extractall/traversal; map relationships to document order, available alt text and nearby paragraph text; do not invent native Docs object IDs or exact caption mapping. Build a self-contained escaped index.html referencing only local safe assets; include native paragraph/table outline and source tab info, figure thumbnails, nearby text, artifact links and embedded PDF. If pdftoppm exists, generate page PNG previews via argument-vector subprocess with timeout and check success; otherwise embed PDF and clearly identify absence of raster previews. Readable text can come from native Markdown export; preserve raw source JSON. Include manifest recording versions, capabilities, asset mapping confidence and artifact hashes; never store tokens or signed temporary URLs in display HTML, and never automatically fetch image contentUri or document hyperlinks. Atomic/publish-ready semantics: mark complete only after successful export/index write, leave explicit failed state on partial work, no overwrite existing dirs, no partial page list labeled complete. Do not contact non-Google domains from content. Optional offline rendering from captured fixture files should allow hermetic end-to-end testing.
Acceptance criteria and tests
Hermetic fixture JSON/DOCX with styles/tables/missing image URIs; real zip/XML extraction; malicious zip paths/symlinks/member/total limits; HTML/script injection escapes; untrusted links never fetched; duplicate image basenames; table and tab traversal; missing pdftoppm fallback; renderer failure and timeouts reported; inconsistent revisions marked; missing revisions not claimed consistent; export failure yields no complete manifest; existing output dir refuses; no parent/symlink escape; native image metadata mapping reports uncertainty; gws stub integration exercises export command/cwd and outputs.
Upstream coordination
Related: New companion workflow; related googleworkspace#743 and googleworkspace#856.
This fork issue tracks one independent contribution from our document-workflow improvement effort. Existing upstream issues remain the canonical reports; the resulting PR will target googleworkspace/cli and reference them.
Delivery
- Separate branch:
feat/docs-review-bundle.
- Tests first, independent code review, required checks and changeset.
- Synthetic fixtures only; no personal documents or credentials in public artifacts.
Implementation PR: googleworkspace#934
Workflow improvement
Export a visual Google Docs review bundle.
Proposed implementation
Add an executable companion under examples/docs-review-bundle/ using existing gws subprocesses, no separate auth. Create a new nonexisting user-named directory within CWD. Export Docs JSON with all tabs, native PDF, DOCX and Markdown via Drive API, optionally comments (include only if retrieval succeeds; explicitly mark unavailable). Use relative export filenames with gws cwd inside bundle to work with existing path restrictions. Check source revision before/after exports and label mixed-version output instead of claiming atomic snapshot. Extract raster assets from DOCX zip safely with total/member size caps and without extractall/traversal; map relationships to document order, available alt text and nearby paragraph text; do not invent native Docs object IDs or exact caption mapping. Build a self-contained escaped index.html referencing only local safe assets; include native paragraph/table outline and source tab info, figure thumbnails, nearby text, artifact links and embedded PDF. If pdftoppm exists, generate page PNG previews via argument-vector subprocess with timeout and check success; otherwise embed PDF and clearly identify absence of raster previews. Readable text can come from native Markdown export; preserve raw source JSON. Include manifest recording versions, capabilities, asset mapping confidence and artifact hashes; never store tokens or signed temporary URLs in display HTML, and never automatically fetch image contentUri or document hyperlinks. Atomic/publish-ready semantics: mark complete only after successful export/index write, leave explicit failed state on partial work, no overwrite existing dirs, no partial page list labeled complete. Do not contact non-Google domains from content. Optional offline rendering from captured fixture files should allow hermetic end-to-end testing.
Acceptance criteria and tests
Hermetic fixture JSON/DOCX with styles/tables/missing image URIs; real zip/XML extraction; malicious zip paths/symlinks/member/total limits; HTML/script injection escapes; untrusted links never fetched; duplicate image basenames; table and tab traversal; missing pdftoppm fallback; renderer failure and timeouts reported; inconsistent revisions marked; missing revisions not claimed consistent; export failure yields no complete manifest; existing output dir refuses; no parent/symlink escape; native image metadata mapping reports uncertainty; gws stub integration exercises export command/cwd and outputs.
Upstream coordination
Related: New companion workflow; related googleworkspace#743 and googleworkspace#856.
This fork issue tracks one independent contribution from our document-workflow improvement effort. Existing upstream issues remain the canonical reports; the resulting PR will target googleworkspace/cli and reference them.
Delivery
feat/docs-review-bundle.Implementation PR: googleworkspace#934