Skip to content

fix: group transcript captions into readable sentences - #2417

Open
pavzagor wants to merge 11 commits into
CapSoftware:mainfrom
pavzagor:codex/fix-shitty-transcription
Open

pavzagor wants to merge 11 commits into
CapSoftware:mainfrom
pavzagor:codex/fix-shitty-transcription

Conversation

@pavzagor

@pavzagor pavzagor commented Oct 5, 2026 •

Copy link
Copy Markdown

The Transcript tab currently renders every subtitle cue as a separate row, splitting speech at commas, short pauses, and every eight words. This groups adjacent cues into sentence rows while preserving speaker boundaries and exact outer timestamps. Existing transcripts, translated captions, and provisional live transcripts benefit without retranscription.

Subtitle files retain their original timing. Owners can select Edit transcript to expose the original cues for precise corrections; edits never save a merged sentence under a single cue ID. Timestamped copying uses the readable sentence rows. Grouping also stops at long silence, overlaps, and bounded length/duration when punctuation is missing.

Depends on #2416 (AssemblyAI diarization). This branch is rebased on #2416's current head (50ae1be), which is rebased on current main; merge #2416 first. Until then GitHub's diff includes the prerequisite. Review only this follow-up's four files.

English before/after proof

A fresh AssemblyAI transcription of NASA's public JFK archival clip produces 12 caption fragments before → 3 sentence rows after, preserving all 72 words. Selecting the matching row on either side seeks the embedded source video to 12.850 seconds.

Watch/download the before/after MP4 · GitHub video page · Source clip · All proof files and methodology

English transcript before and after

The proof uses the actual old/new React components and real ASR output, with storage/auth hooks mocked. It is not production footage or proof of a persisted backend edit. Full local share-page verification was blocked by unavailable MySQL at 127.0.0.1:3306. NASA's clip splices two speeches; A/B are the unmodified ASR labels.

Validation

  • 91 focused tests passed: sentence grouping/UI, diarization, VTT, text formatting, translation, caption generation, and transcript editing.
  • Scoped Biome check and git diff --check passed.
  • Web tsc --noEmit passed in the existing development checkout; the isolated checkout reused dependencies and was unsuitable for the workspace reference type check.
  • Browser checked sentence seeking, fragment editing controls, and 360px/393px layouts without horizontal overflow. Timestamped copying is covered by the UI tests.
  • No migration, new environment variable, or paid reprocessing is required by this change.

Review fix: multilingual sentence endings

Addressed the Arabic question-mark finding in decfd40 with Unicode Sentence_Terminal. Regression tests cover Arabic, quoted Arabic and Hindi, plus original/translated Arabic rendering and exact seeking. All 91 focused tests, TypeScript, scoped Biome and whitespace checks pass. The approved English output remains byte-for-byte equivalent as parsed sentence entries (12 fragments → 3 sentences).

Review-fix walkthrough · Validation log · English parity check

Arabic and Hindi sentence boundaries before and after

This additional proof uses the actual component with a synthetic multilingual fixture and mocked storage/auth.

October 5 review fixes

Single-letter labels such as Option A. now end their sentence. Standalone and consecutive initials, titles, and ellipses still remain grouped. A lone embedded initial remains ambiguous with a label and is treated as terminal. Switching between cue editing and sentence reading clears cue selection explicitly, so a selected non-leading cue cannot leave an inconsistent sentence highlight. The legacy literal-angle-bracket fix from #2416 is included.

Validation on e261911a: 81 focused transcript/parser/edit/export tests passed; scoped Biome passed. Actual-component fixtures verify sentence boundaries, complete 2 < 3, precise seeking to 2 seconds, and selection clearing at desktop and 390px mobile widths. Auth/storage/network are mocked in these fixtures; this is not authenticated full-page proof. The attempted full web typecheck did not pass: the isolated checkout lacks built shared-project declarations and reuses dependency links from the primary checkout. Fresh CI remains subject to upstream workflow approval.

Mobile reading fixture · After Done editing · Fixture methodology

Security review: the extra canonical full-recording pass belongs to prerequisite #2416 and is required for recording-wide speaker identity. The pre-feature workflow already fell back to a full pass after chunks when live promotion failed. The existing canonical database claim remains covered by scheduling tests. Per-owner usage budgets are not introduced by this sentence-grouping follow-up; the broader cost-control concern needs a separately defined quota policy.

October 7 update

Rebased on #2416's current head. The visual-evidence commit was removed from the branch as requested; the linked screenshots remain available at their original commits. e.g. and i.e. (any case) no longer end a sentence row, with regression tests.

Validation on 8341885: 52 focused sentence/diarization/agent tests pass; typecheck and full web suite pass on the #2416 base. Greptile 5/5 on 8341885.

RetriggerConfidence Score: 5/5

The PR appears safe to merge based on the reviewed changes.

Summary

The PR groups caption cues into readable sentence rows while retaining original cues for editing and preserving speaker information. Since the previous review, it also adjusts agent VTT parsing and makes provisional live transcription opt-in.

  • Both previous findings were manually resolved; neither is reposted.
  • No new actionable issue was established.

Reviews (4) · Last reviewed commit: "fix: keep e.g. and i.e. inside transcrip..." · Reviewed by Greptile

@superagent-security superagent-security Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Superagent found 1 security concern(s).

Comment thread apps/web/workflows/live-transcribe.ts
Comment thread apps/web/lib/transcript-sentences.ts Outdated
Comment thread apps/web/app/s/[videoId]/_components/tabs/Transcript.tsx
@richiemcilroy

Copy link
Copy Markdown
Member

Thanks. This needs #2416 to land first, and we require Greptile 5/5 on the latest commit plus green CI before review. Please rebase on current main, remove the .github/assets evidence files from the branch, and trigger a Greptile re-review.

@pavzagor
pavzagor force-pushed the codex/fix-shitty-transcription branch from 6f3e5a7 to 240d2ac Compare October 7, 2026 15:33
@pavzagor

pavzagor commented Oct 7, 2026

Copy link
Copy Markdown
Author

@greptileai review

@pavzagor

pavzagor commented Oct 7, 2026

Copy link
Copy Markdown
Author

@richiemcilroy done

@pavzagor
pavzagor force-pushed the codex/fix-shitty-transcription branch from 240d2ac to 7a8b38b Compare October 7, 2026 15:53
@pavzagor
pavzagor force-pushed the codex/fix-shitty-transcription branch from 7a8b38b to 0e20d45 Compare October 7, 2026 16:01
@pavzagor
pavzagor force-pushed the codex/fix-shitty-transcription branch from 0e20d45 to 8341885 Compare October 7, 2026 16:26
@pavzagor

pavzagor commented Oct 7, 2026

Copy link
Copy Markdown
Author

@greptileai review

pavzagor commented Oct 7, 2026

Copy link
Copy Markdown
Author

@greptileai review


Generated by Claude Code

@pavzagor

pavzagor commented Oct 8, 2026

Copy link
Copy Markdown
Author

@richiemcilroy

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants