Skip to content

Read Google Docs as compact structured content #2

Description

@ratovarius

Workflow improvement

Read Google Docs as compact structured content.

Proposed implementation

Implement gws docs +read --document ID with --params JSON passthrough for API options, using includeTabsContent=true and suggestionsViewMode=SUGGESTIONS_INLINE unless explicitly supported alternatives are documented. Normalize fetched Docs structure to compact structured JSON: document ID/title/revision/view, recursive tab ID/title/parent, ordered blocks with original start/end UTF-16 indices, paragraphs/headings and text runs with relevant style/link/suggestion metadata, nested tables/cells, inline object references and image metadata, footnote/reference markers. Preserve markers for unknown elements so content is not silently erased. Never claim complete content when a restrictive fields mask omitted it: reject incompatible fields masks or validate expected content. Use the existing global json/table/yaml/csv formatter; do not add a global markdown mode or a flag for every response field. The reader removes API noise, exposes a searchable outline and allows downstream jq to select a section. Request construction plus meaningful transformation justifies the helper. Use existing auth/executor capture and sanitize infrastructure, no independent raw unauthenticated HTTP. Use returned API indices, never Python/Unicode character counts. Preserve existing +write behavior. Support dry-run as request plan without fetching content or acquiring credentials.

Acceptance criteria and tests

Command registration and required args; simple body and legacy body fallback; recursive tabs/childTabs; emoji/non-BMP indices unchanged; styled links across text runs; tables with nested cell content; images missing contentUri; suggestions insertion/deletion metadata; unknown blocks preserved; no content if forbidden partial mask; server failures propagated; dry-run is side-effect free; formatter selection honored.

Upstream coordination

Related: googleworkspace#726.

This fork issue tracks one independent contribution from our document-workflow improvement effort. Existing upstream issues remain the canonical reports; the resulting PR will target googleworkspace/cli and reference them.

Delivery

  • Separate branch: feat/docs-structured-read.
  • Tests first, independent code review, required checks and changeset.
  • Synthetic fixtures only; no personal documents or credentials in public artifacts.

Implementation PR: googleworkspace#932

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions