Skip to content

chore: sync main with upstream kenn-io/msgvault (aa411818) - #2

Merged
arcaputo3 merged 125 commits into
mainfrom
sync/upstream-2026-10-05
Oct 5, 2026
Merged

arcaputo3 merged 125 commits into
mainfrom
sync/upstream-2026-10-05

Conversation

@arcaputo3

Copy link
Copy Markdown

Fast-forwards the fork's main to upstream aa41181 (122 commits since d2895b6, 2026-09-17). Merge with 'Create a merge commit' so upstream history is preserved.

rodboev and others added 30 commits September 17, 2026 18:23
The daemon checks `draft.create` before an agent's `draft-reply` request can run or make other work wait. The command handler and operation gate share this check. The draft command also checks that the token covers the requested account.

Only `draft.create` is accepted. Tests cover allowed and rejected requests and drafting for the correct account.

Refs kenn-io#666.


Co-authored-by: Rod Boev <rodboev@users.noreply.github.com>
Shorten Web UI links by omitting defaults, keyboard focus, scroll position, and inactive workspace choices. Ordinary tabs now use readable links such as `?workspace=files&mode=full_text`; filters, layout changes, and selected items use a smaller exploration payload when needed.

Keep the complete session state in browser history so Back and Forward restore focus and choices from other tabs. Search mode stays explicit so a shared link does not pick up a different browser's saved preference.


Co-authored-by: Wes McKinney <wesm@users.noreply.github.com>
Retrieve, edit, and delete managed IMAP reply drafts through the CLI, with the selected daemon owning provider access and local persistence.

The daemon tracks draft ownership and revisions in SQLite or PostgreSQL. An edit saves replacement content before APPEND, publishes the new revision, then removes the exact old UID. Conditional deletion conflicts stop removal. Reused mailbox IDs preserve previously archived content.

Deletion failures before a remote write leave the draft retryable. When remote removal is confirmed, a retry can finish the local operation. Uncertain writes stay pending and retain content and known IMAP copies; explicit recovery remains a separate follow-up. Discarded drafts retain local content for `draft-get`.

Lifecycle commands remain owner-only. Committed changes refresh analytics even when cleanup or response delivery fails.

Refs kenn-io#666, slice 3b.


Co-authored-by: Rod Boev <rodboev@users.noreply.github.com>
The IMAP setup and folder filtering links now work from both GitHub and the published documentation. Both pages use relative Markdown paths, which the docs build converts to the existing site routes.

Closes kenn-io#872

Co-authored-by: Rod Boev <rodboev@users.noreply.github.com>
Reject unsupported CardDAV providers in config files. A typo such as `provider = "googl"` previously selected password authentication and could reuse an existing bound credential; only `""` and `"google"` are now accepted, matching the API.

Credential reuse also requires an explicit provider match. Follow-up to kenn-io#820.


Co-authored-by: Wes McKinney <wesm@users.noreply.github.com>
Explain when archive queries run out of memory or temporary disk space, with steps to narrow the results or adjust the server's query limits. Everything, Files, and other views using the shared query error handler receive this guidance; other unexpected failures direct the operator to the server logs.

Files now offers Retry after an initial failure and hides its count while loading or unavailable, so a failed request does not appear to mean the archive has no files. The configuration guide explains how to check resources, change the limits, and restart the daemon.


Co-authored-by: Wes McKinney <wesm@users.noreply.github.com>
Install `libsqlite3-dev` in the `frontend` and `web-e2e` Playwright containers so their daemon builds can compile the default `sqlite_vec` dependency. Both jobs currently fail with `sqlite3.h: No such file or directory` before the browser tests run.

The PR dispatcher loads this reusable workflow from `main`, so the dependency must land separately before it can unblock the browser checks on kenn-io#841. The change adds the package to the two existing apt install lists.

Refs kenn-io#841


Co-authored-by: Rusty Shackleford <salmonumbrella@users.noreply.github.com>
…#879)

Adds `msgvault eval` to measure whether search changes help users find relevant messages. It compares keyword (`fts`), semantic (`vector`), and combined (`hybrid`) search against queries and relevance ratings you provide.

```sh
msgvault eval --qrels qrels.txt --topics topics.tsv --modes fts,vector,hybrid -n 100
```

`topics.tsv` contains query IDs and search text, separated by tabs. `qrels.txt` uses `query_id 0 document_id grade`; grades of 1 or higher mean relevant.

- Reports relevant results found, their ranking, and query timings as a table or `--json`.
- Records model, search settings, and archive/index sizes for comparison.
- Scores messages or conversations with `--doc-key`, counting each conversation once.
- Reports skipped queries, missing ratings, and incomplete rankings; rejects conflicting ratings and ambiguous IDs.

Vector modes require a local SQLite archive, a configured embedding provider, and an existing index. Compare runs using the same query file.

Supersedes kenn-io#649 because maintainers cannot update its fork branch. Preserves Frederic Masi's authorship and rebases the contribution onto main. Follow-up to kenn-io#367.


Co-authored-by: Wes McKinney <wesm@users.noreply.github.com>
PST import now treats a message flagged with attachments but carrying an empty attachment table as having no attachments. It no longer increments the import error count for that case.

go-pst wraps `ErrTableContextNoRows` inside the attachment iterator error, while `ReadAttachments` recognized only `ErrAttachmentsNotFound`. It now matches either sentinel by identity. Other attachment table, read, and iterator failures remain counted, and source identity and message content inputs stay unchanged.

On the EDRM Enron PST from kenn-io#886 the import error count drops from 3 to 1; the remaining error comes from a search-folder entry that a separate change handles.

Refs kenn-io#886

Co-authored-by: Rod Boev <rodboev@users.noreply.github.com>
`msgvault import-pst` now skips PST search folders. On the EDRM Enron file from kenn-io#886, it still imports all 2,178 messages, reports 2 errors instead of 3, and avoids creating an empty `All Messages` label.

The PST folder walker documented this behavior but never checked the folder type. It passed `Search Root/All Messages` with a stored count of 2,178 to go-pst, which refuses to read search folders. The walker now filters `IdentifierTypeSearchFolder`, matching go-pst's check.

The two remaining errors on that file come from unreadable attachment tables and are handled in a separate change.

Refs kenn-io#886


Co-authored-by: Rod Boev <rodboev@users.noreply.github.com>
`draft-recover` resumes an interrupted IMAP draft edit or delete from its saved receipt. It can publish a known replacement or finish cleanup after the original copy is gone, without appending another copy.

The daemon checks the caller's source-scoped `draft.edit` or `draft.delete` grant before returning delegated revision or policy results. It then verifies the revision, source identity, mailbox generation, exact UID, and draft flags. Delegated responses include lifecycle metadata while omitting draft content, raw MIME, and candidate bytes.

Recovery uses only saved receipts and bytes. Unknown UIDs, moved copies, mailbox searches, and sync-archived replacements remain unresolved, so recovery cannot adopt an unrelated message.

Refs kenn-io#666 (comment), slice 3c

Closes kenn-io#862

Co-authored-by: Rod Boev <rodboev@users.noreply.github.com>
…chments/ directory (kenn-io#883)

Fixes kenn-io#878.

Apple Mail keeps some attachments outside `.partial.emlx` files, in a sibling `Attachments/` directory. This change restores cached top-level attachments when importing those messages and reports how many it restored.

Re-importing adds newly cached attachments to the existing message. Its ID and fallback conversation key come from the original MIME bytes, so restoration does not create duplicates. Updates retain existing labels and fill placeholders in the stored MIME, preserving attachments even if Apple Mail later removes their cached files.

Restored parts use base64 with a matching Content-Transfer-Encoding header. Missing files keep their placeholders; unreadable files or directories produce warnings. Attachment reads are bounded by the remaining message budget, and their actual base64 size is checked before restoring a part. Filename validation and the single-file fallback remain in place.

Restoration supports direct children of the outer multipart only. Nested parts are not restored, and the documentation states that limit. A sanitized Apple Mail directory listing is needed to verify the numbering of nested parts before adding support. LF and CRLF messages retain their line endings; mixed line endings can leave placeholders unchanged.


Co-authored-by: Thomas Heinrichsdobler <fucx@users.noreply.github.com>
Update the documentation for 0.20.0 with the final changelog, upgrade
guidance, and acknowledgements for all 16 contributors identified
between the release tags.

- Fill gaps in the CLI reference, including search evaluation, visual
search, contact activity, subset export, and browser message links.
- Correct verification limits, IMAP behavior, provider consent, SQL
access, recovery advice, Windows builds, and container version
selection.
- Rewrite the product and lifecycle pages around user outcomes, keep
their Markdown companions aligned, and place browser screenshots beside
the tasks they illustrate.
- Replace prerelease notices with links to the dated 0.20.0 changelog
while preserving the detailed upgrade notes and older release history.

The ten refreshed browser screenshots are published directly on
`docs-assets` at commit cba6c00. The
captures use the existing reviewed Enron fixture. Existing terminal
illustrations retain their visible v0.19.0 label; the two obsolete IMAP
deletion diagrams remain excluded from public pages.

---------

Co-authored-by: Codex <noreply@openai.com>
…o#893)

Replace vague website claims with concrete descriptions of saving messages, managing contacts, searching attachments, and reviewing deletions. Update the homepage, guide, page descriptions, and Markdown companions together, including the SQLite backup limit.

Move Changelog directly below Setup in the documentation navigation so release changes and upgrade notes are easier to find. This PR changes copy and navigation only; the deployment and Markdown publishing fixes are in kenn-io#892.


Co-authored-by: Wes McKinney <wesm@users.noreply.github.com>
Repair the Vercel deployment that failed because uploading only `docs/` omitted the sibling `website/` sources. Preview and production targets now build locally and deploy the complete output. Legacy documentation redirects also match trailing slashes, so links such as `/setup/` reach `/docs/setup/`.

Publish Markdown companions for all public content pages using the same sanitized sources as the HTML build. `llms.txt` lists those pages, and each HTML page advertises its Markdown URL. The existing built-site check now verifies coverage, alternate links, and source content.

These changes are already deployed on [msgvault.io](https://msgvault.io).


Co-authored-by: Wes McKinney <wesm@users.noreply.github.com>
## What changed

- Give Directory results and person details their own scroll areas so long records remain reachable.
- Show populated attributes near the top of the person overview, with typed values, choice labels, and sensitive values concealed.
- Hide empty attribute fields by default, with a control to reveal them when adding values. The summary's Edit attributes control jumps to the full field section.

## Why

Long person records can extend past the visible detail pane, and useful attributes sit below several other sections. This makes the Directory harder to use for people with more profile data.

## Usage

Open a person in Directory. Use Edit attributes from the overview to reach the full attribute section, and Show empty fields when adding a value to an unused field.

Refs kenn-io#901


Co-authored-by: Rusty Shackleford <salmonumbrella@users.noreply.github.com>
…es (kenn-io#907)

Archives with qualifying messages dated before 1970 can fail analytics cache builds because annual relationship temperature rows violate the validator's year floor.

The annual rollup now filters its input at the same 1970 floor used by validation, while current scores and daily activity still include older messages. A shared constant keeps the two checks aligned.

Regression coverage proves cache builds succeed with pre-1970 messages, keeps 1970 annual data, retains older activity in current and daily projections, and rejects invalid stored annual rows.

Closes kenn-io#897


Co-authored-by: Rod Boev <rodboev@users.noreply.github.com>
…n-io#913)

IMAP synchronization can lose the identity of already archived messages whose `Message-ID` uses historical, non-RFC syntax, such as `123456789` or `<[legacy-token==@example.test]>`. MIME import preserves these identifiers, but `rawMIMEMessageID` returned an empty string when the stricter mail-header parser rejected them. After a filtered Inbox import, a full sync can consequently report `inconclusive IMAP identity returned a dedup stub` and leave folder state unpublished.

Keep the existing strict parser as the first choice and fall back to the existing `mime.ParseMessageIDs` normalizer when it rejects an identifier. This keeps standard identifiers and comment handling intact while allowing header enumeration, label refresh, and raw-message deduplication to recognize the same legacy IDs. Malformed angle-bracket structures and identifier-shaped body text remain rejected.

The regression tests cover legacy IDs across IMAP fetch paths and repeated full syncs after a filtered Inbox import, checking that one archived message retains both folder memberships, both saved folder cursors, and its original raw MIME bytes.


Co-authored-by: Elie BRUNO <eliemada@users.noreply.github.com>
## What changed

- Let the relationship calendar use the available card width while keeping narrow layouts reachable.
- Show each day's activity summary on pointer hover or keyboard focus, with edge tooltips kept inside the calendar.
- Hide year arrows for single-year history. For multi-year history, the unavailable direction is visibly disabled.

## Why

The calendar's fixed cell sizing wastes room in wider cards and can clip in narrow ones. Day details and year controls are also easy to miss or misread.

## Usage

Hover over or focus a calendar day to read its summary. Use the year arrows when the relationship spans multiple years.

Refs kenn-io#901


Co-authored-by: Rusty Shackleford <salmonumbrella@users.noreply.github.com>
## What changed

- Collapse linked identities by default, with a count in the disclosure. Keep it open during unlink confirmation or an error, and reset it when switching people.
- Add an explicit Messages/Files choice. Opening another person starts in Messages; browser Back restores the earlier person's choice.
- Show populated person attributes in the relationship header and open that person directly in Directory.

## Why

Linked identities can crowd out the relationship detail, and the Files view is easy to lose while moving between people. Populated attributes also require a separate trip through Directory to find.

## Usage

Use Messages or Files in the relationship header, expand Identities when needed, and use Open in Directory for the selected person.

Depends on kenn-io#902. This branch is stacked on that PR, so its diff includes the Directory change until kenn-io#902 merges.

Refs kenn-io#901


Co-authored-by: Rusty Shackleford <salmonumbrella@users.noreply.github.com>
…-io#899)

## What changed

- Handle Digest challenges during CardDAV requests, including conditional writes and same-origin redirects, within the existing request limits.
- Let the local operator pin private addresses for one exact HTTPS origin. A failed pin does not fall back to DNS, and TLS still verifies the hostname.
- Carry that policy through account test, save, restart, and schedule changes. A concurrent policy edit returns a conflict; if discovery was already persisted, the controller disables the previous service and schedule.

## Why

An address book that requires Digest authentication and sits behind a private network address cannot complete account setup or sync. The container may also be unable to resolve its hostname even when it can reach the address.

## Usage

Set the local daemon config before adding the account:

```toml
[carddav]
trusted_origin = "https://contacts.example:8443"
trusted_addresses = ["10.1.2.3"]
```

Use the same origin and port in the account URL. The account API cannot set these fields. Without this local policy, private destinations remain blocked.

Closes kenn-io#898


Co-authored-by: Rusty Shackleford <salmonumbrella@users.noreply.github.com>
CLI archive commands can now opt out of starting a local background daemon when a supervised msgvault serve owns the archive. Set `[server].daemon_auto_start = false` to use a running daemon or return an actionable error.

The resolver previously spawned a background serve whenever it found no usable runtime. The new setting gates that implicit spawn and treats daemon replacement as disabled too, while explicit daemon lifecycle commands keep their existing behavior. The default remains automatic startup.

The setting addresses the client-side race. Launchd ownership and Full Disk Access behavior remain the supervisor and operating-system concerns described in the report.

Closes kenn-io#895



Co-authored-by: Rod Boev <rodboev@users.noreply.github.com>
…kenn-io#875)

Tests now settle work the test process owns deterministically. Runtime-record heartbeats, idle shutdown, search indexing, vector initialization, lease renewal, and delegated-draft gate waits use `testing/synctest` or direct channel receives. External database, process, network, and DuckDB waits keep their budgets, with names that explain the event.

`make testify-helper-check` now rejects bare testify polling budgets below one second. The development guide documents the rule and the conversion follows the [agentsview precedent](kenn-io/agentsview#1702).

No usage change.

Closes kenn-io#858


Co-authored-by: Rod Boev <rodboev@users.noreply.github.com>
…ics (kenn-io#841)

## What changed

- Export selected meetings as bounded JSON or Markdown context, with archive citations, coverage, and optional transcripts.
- Query source-reported action items and meeting activity across Granola, Circleback, Notion, and generic imports. Keep original action status and distinguish unknown durations from zero.
- Add context export to Web selection and meeting activity to person, domain, Directory, and Relationships views, with links back to archived evidence.
- Expose the same operations through HTTP, the daemon-backed CLI, and read-only MCP tools. Backfill existing SQLite and PostgreSQL archives without a provider resync.

## Why

Meeting notes are already searchable, but using them as agent context still requires parsing provider-specific bodies. Follow-ups and time spent with a person are harder to query across sources. These operations make the archived evidence reusable while keeping source status and missing information visible.

## Usage

In the Web UI, select meetings and choose Export meeting context. Meeting panels show activity and recorded follow-ups for the current scope.

```sh
msgvault meetings context --id 12 --id 34 --format markdown --output context.md
msgvault meetings actions --person-id 5 --status pending --json
msgvault meetings metrics --domain example.com --after 2026-01-01 --json
```

Add `--include-transcript` when exporting full transcript evidence. Action status describes the latest archived source snapshot.

**Web UI** (synthetic meeting data)

Meeting activity with known and unknown duration coverage:

![Meeting activity and duration coverage](https://raw.githubusercontent.com/salmonumbrella/msgvault/535ace0e97e48c6104e9e0ecad762f0d9e55f72e/meeting-overview.png)

Archived action evidence, transcript, and JSON/Markdown context export:

![Archived meeting reader with action evidence and context export](https://raw.githubusercontent.com/salmonumbrella/msgvault/535ace0e97e48c6104e9e0ecad762f0d9e55f72e/source-deleted-reader.png)

Closes kenn-io#840


Co-authored-by: Rusty Shackleford <salmonumbrella@users.noreply.github.com>
Opening Sources now reads sync status without taking a sync lock or changing stored runs. Cancelling a status request also cancels its database work.

Status and Sync now requests time out after 20 seconds. Sync errors offer Refresh to check current status without starting another sync. Polling pauses in hidden tabs and resumes when visible. After eight refresh attempts with a held lock but no active run, including failed requests, polling stops and offers Refresh. Cursor fields are removed from status responses, schemas, and generated clients.

One tradeoff: viewing Sources no longer repairs a run left marked as running after a crash. A daemon restart or another sync that acquires the source lock recovers it. Until recovery, Sync now stays unavailable and status polling continues while the tab is visible. The poll limit applies to held locks without an active run.

Refs kenn-io#901


Co-authored-by: Rusty Shackleford <salmonumbrella@users.noreply.github.com>
…-io#908)

MCP initialization and tool discovery now proceed without waiting for archive statistics.

Authenticated health exposes text and visual lanes independently from schema 2.28.0. Older daemons keep basic MCP tools and reduced semantic guidance, but need an upgrade for full semantic, similar-message, and visual search. Visual search also requires the 2.4.0 route. Readiness stays on each search request.

The concurrent seeded-attribute test barriers exit on release, so late retries can't hang the store CI shard.

Closes kenn-io#896


Co-authored-by: Rod Boev <rodboev@users.noreply.github.com>
)

Canceled SQLite planner maintenance is now logged at debug level, so large archives stop printing repeated warning noise during normal CLI startup and shutdown.

The one-second best-effort maintenance attempt still runs at schema setup, successful sync, and store close. Genuine maintenance failures remain warnings, and successful planner-statistics refreshes keep their current behavior.

The change addresses diagnostic severity without changing the maintenance schedule. The reported behavior came from kenn-io#894.

Closes kenn-io#894



Co-authored-by: Rod Boev <rodboev@users.noreply.github.com>
Slack sync can now exclude one-to-one DMs and group DMs independently. Set `dms = false` or `group_dms = false` under `[slack]` to shorten routine syncs while selected channels keep their current filters.

The importer previously accepted every DM and group DM before applying channel filters. Channel filters remain the existing way to narrow named channels, but they can't select DMs by name.

These settings add the smallest separate configuration and Settings API controls for each conversation kind. They are default-true and use Slack conversation kind flags at the existing selection owner. During incremental syncs, skipped conversations retain resume and thread state for later re-entry. After `--full` resets state, re-enabling a conversation performs a fresh, idempotent walk.

Existing channel selection and full-import behavior remain available. The timings and conversation breakdown came from kenn-io#885.

Closes kenn-io#885



Co-authored-by: Rod Boev <rodboev@users.noreply.github.com>
## What changed

- Add `import-imazing-csv` for iMazing Messages export roots and CSV directories, including comma, tab, and semicolon files.
- Preserve conversations, sender identity, service, delivery evidence, conservative reply links, and available attachments across deterministic reruns.
- Support an explicit IANA timezone and optional vCard enrichment, and document rerun and cross-source overlap behavior.

## Why

Some message histories survive only as iMazing CSV exports after the original Apple Messages database or encrypted backup is gone. Those exports can now join the archive without requiring the original backup.

## Usage

```bash
msgvault import-imazing-csv ~/Downloads/messages-export \
  --me +14155550100 --timezone America/Los_Angeles
```

Pass the export root or its `csv/` directory. `--me` is required; `--timezone` defaults to the local IANA zone, and `--contacts` accepts a vCard file for missing participant names.

Closes kenn-io#832


Co-authored-by: Rusty Shackleford <salmonumbrella@users.noreply.github.com>
## What changed

- Add a SQLite approximate index for semantic, hybrid, and similar-message search, with exact reranking and fallback to exhaustive search.
- Add `msgvault embeddings optimize` to build or resume the index from stored embeddings without calling the embedding provider. `embeddings list` shows its state and progress.
- Keep ready indexes current as embeddings change, and expose query-embedding, retrieval, and hydration timings through the API, CLI, and MCP.
- Leave PostgreSQL search behavior unchanged.

## Why

Large SQLite archives currently scan stored vectors for each semantic query. The accelerator bounds candidate retrieval while keeping exact search available when the index is unavailable or explicitly disabled.

## Usage

After activating an embedding generation:

```bash
msgvault embeddings optimize
msgvault embeddings list
```

The default `[vector.search]` setting, `sqlite_accelerator = "auto"`, uses a ready index. Set it to `"exact"` to return to exhaustive search.

Closes kenn-io#870


Co-authored-by: Rusty Shackleford <salmonumbrella@users.noreply.github.com>
mariusvniekerk and others added 28 commits October 1, 2026 12:41
Use backoff/v7 for SQLite contention and checkpoint maintenance,
Calendar requests, Circleback rate limits, Beeper message-page fetches,
sync identity discovery, IMAP transport reconnects, and daemon busy
responses. Preserve each path’s attempt limit, nominal exponential
intervals, error classification, cancellation behavior, and operation
error returned to callers. Exponential policies use the library’s
default jitter; deliberate full-jitter policies retain their existing
delay ceilings. Disable the library’s default elapsed-time limit where
the existing attempt budget or caller context owns the limit.

Keep provider and embedding retries that advance exponential delay
across `Retry-After` responses: v7 resets the policy after those
responses. Also keep Gmail’s separate quota-cycle and request-timeout
budgets, linear seeded-attribute delays, persisted job schedules, and
immediate optimistic-concurrency reconciliation.

<sup>generated by a clanker</sup>

---------

Co-authored-by: Codex <noreply@openai.com>
…enn-io#1035)

Explore searches and large People/Domain groupings were doing avoidable work, while provider resyncs could invalidate cached Files listings. 

- Resolve full-text candidate IDs in one bounded query instead of repeatedly counting and hydrating pages of messages. Preserve ranking, filters, totals, and candidate limits.
- Build grouping membership from sparse participant edges and one roster per conversation, then deduplicate numeric keys before attaching aggregate data. No cache format or default memory-limit change.
- Fetch fewer exact vector candidates initially and retain widening for duplicate chunks and deleted messages. Unfiltered hybrid search reuses bounded vector retrieval instead of materializing every live message ID.
- Preserve attachment IDs for retained source parts during provider resync and clear thumbnails when content changes. Unchanged attachments keep cached Files listings usable; actual removals still reject stale metadata. This fixes a reproduced source of HTTP 409 responses; the original report did not retain enough detail to establish its cause.
- Keep relationship files visible below expanded overviews, wrap long reader subjects, reveal truncated names and filenames on hover while preserving grid keyboard focus, improve avatar contrast, restore `/` search focus from Deletions, and skip unavailable CardDAV requests. Availability recovery loads address books without replaying consumed conflict focus requests.

Synthetic benchmarks on the same machine show:

| Operation | Before | After |
| --- | ---: | ---: |
| Full-text HTTP search, 5,000 matches | 452 ms | 207 ms |
| Full-text HTTP search, 20,000 matches | 1,698 ms | 384 ms |
| Exact semantic retrieval | 190 ms | 166 ms |
| Hybrid retrieval, rare query | 170 ms | 41 ms |
| Hybrid retrieval, common query | 244 ms | 122 ms |

These are medians of three benchmark runs. HTTP searches use 20,000 messages; retrieval uses 100,000 messages with 64-dimensional vectors and excludes embedding calls and HTTP projection. Rare full-text searches were unchanged. On the 2.56-million-message grouping fixture, People and Domains now complete with the configured 512 MB limit and disk spill, taking about 2.7 s and 1.7 s warm; both previously exhausted that limit.

Real-archive latency still needs measurement. New `Server-Timing` phases separate lexical search, embedding, retrieval, projection, identity hydration, and grouping, with reproducible benchmark recipes in the development guide.

Generated with Codex


Co-authored-by: Wes McKinney <wesm@users.noreply.github.com>
…n-io#1034)

MsgVault now uses Kit v0.29.2 for OpenAI-compatible embedding requests and SQLite accelerator rank fusion. This removes the local OpenAI transport and retry loop, plus the unused hybrid/rrf.go implementation. Model, role, batch and transport settings map to Kit behind the existing embedding interface.

Existing generation fingerprints, stored vectors, chunking and database layouts stay unchanged. A fixture with a pre-migration fingerprint remains searchable, indexing selects the same generation without rebuilding, and model or content-recipe changes still invalidate it. The public api_key_env configuration remains unchanged.

Kit owns retry classification and jitter, with a 200 ms initial delay, a 25.6 s backoff cap before jitter, the existing attempt budget and a one-hour Retry-After cap. HTTP 408 now retries; Retry-After: 0 uses backoff. Malformed completed responses fail directly, while transport read failures retry. Provider errors report status and reason rather than raw response bodies. Consent checks still run before every HTTP attempt and cancellation stops retries.

Retained implementations:

- Document fusion already delegates to Docbank. Its explicit, potentially sparse ranks do not match Kit RRF's contiguous list positions, so it remains to preserve scores, occurrence identity and source authority.
- Full-text SQL retains field weights and PostgreSQL rank normalization that Kit's query builders cannot express. Exact fused SQL also keeps the existing bounded ANN widening and attached vector database.
- GenerationFingerprint remains the stored identity and comparison contract. Kit Descriptor.Generation identifies the vector space separately from the input recipe; switching this storage contract is unnecessary for transport adoption.
- Voyage retains its provider-specific clients and retry helpers. There is no Kit sqlitevec storage migration.

Fusion still orders tied scores by MessageID, represents missing legs locally, and applies subject boosts before the result limit. Kit receives identities rather than MsgVault's NaN score sentinels.

<sup>generated by a clanker</sup>


Co-authored-by: Marius van Niekerk <mariusvniekerk@users.noreply.github.com>
…ations (kenn-io#1013)

This removes 171 lines of code. The diff is +168 because of 339 lines of tests that run people and organizations through the same history, supersede and conflict checks.

People and organizations can both carry custom attributes, like a nickname or an industry, with a history of past values. The code that saves, replaces and clears those values existed twice, once per table, so a fix to one could miss the other.

Both now go through one shared store. People still take the fact-pin locks and record manual pins, and merged organizations still refuse writes. The two clear routes also share one parser for their dry-run, expected value and ordinal options. Routes, JSON bodies and tables are unchanged.

Refs kenn-io#1002, slice 2 (shared attribute store for people and organizations)


Co-authored-by: Rod Boev <rodboev@users.noreply.github.com>
)

Every pane divider in the Web UI is now one 4px kit-ui handle. Before,
Relationships repainted its handle as a 1px hairline. In Everything, the
results table's border sat beside the handle, so that divider was 5px.

- **Hidden handle between stacked panes.** When the reading pane sat
below a list, in Everything or Relationships, its handle was 0px wide.
You could not see it or drag it with a pointer; only the arrow keys
resized the pane. The split wrapper now stretches the handle in both
orientations, and the reading-pane browser test now drags it with the
mouse.
- **No borders beside a handle.** While a reading pane is open, the
results table and grouped table drop their border and rounded corners on
the edge next to the handle. The reading pane's frame already left that
edge open.
- **One source for the size.** SplitPane reserved room for the handle
with its own `4`. It now reads `layout.splitHandleSize` from kit-ui's
`brand.json`, and its size-limit tests compute their expected values
from the same token.

This depends on kenn-io/kit-ui#82, which makes kit-ui the only owner of
the handle thickness. That PR removes the handle's `class` prop and adds
a `kit-ui-check` rule that rejects app CSS that restyles the handle.
`web/package.json` pins kit-ui to that PR's merge commit (`ceff715`).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
## What changed

- Add named CardDAV connections alongside the existing default account, each with its own credential binding, discovery, schedule, retry gate, and sync history.
- Select a connection in the CLI, API, and Settings. Operations synchronizes all enabled connections and reports partial failures with connection names.
- Preserve one global publication target and route writes through its owning connection.
- Upgrade existing SQLite and PostgreSQL archives for named connections while preserving CardDAV data and account references.

## Why

One CardDAV account is not enough when personal and work contacts use different accounts or services. Each connection needs its own credentials and sync settings while sharing the same people directory.

## Usage

```sh
msgvault add-carddav https://contacts.example.com/dav/ you@example.com --connection work
msgvault carddav connections
msgvault carddav books --connection work
msgvault sync-carddav --connection work
```

Settings → CardDAV account also lets you add and switch connections. Omit `--connection` when adding to use `default`; omit it when synchronizing to run all enabled connections. Aggregate partial or failed sync exits nonzero.

Closes kenn-io#990


Co-authored-by: Rusty Shackleford <salmonumbrella@users.noreply.github.com>
…enn-io#1031)

## What changed

- Directory's Primary channel filter now offers Email, Chat, Meeting, and Other.
- The menu choices and labels use one catalog. Regression tests cover selecting and clearing each channel.

## Why

Phone never matched the directory projection, while people whose last contact was a meeting or another activity could not be filtered by channel.

## Usage

Open **Directory → Filters → Primary channel** and choose a channel.

Closes kenn-io#1020


Co-authored-by: Rusty Shackleford <salmonumbrella@users.noreply.github.com>
…1033)

Closes kenn-io#997

The Microsoft setup steps now add `IMAP.AccessAsUser.All` from **Microsoft Graph > Delegated permissions**. Office 365 Exchange Online has no delegated IMAP permission, so users who followed the old step could not find the permission that `add-o365` needs.

No configuration or usage changes.


Co-authored-by: Mike Campbell <exactmike@users.noreply.github.com>
…enn-io#1055)

### What changed

PST imports preserve transport threading headers and fill missing Message-ID, In-Reply-To, and References from MAPI properties. After all folders are imported, msgvault resolves reply parents and reconciles conversations using unique, live identifiers from the same source.

### Why

Partial transport headers discarded identifiers that were still present in the PST. Reply chains could also split when messages arrived before their parents. Reimporting now repairs missing metadata on existing messages while preserving archived content and conflicting stored identifiers.

### Usage

Rerun the same PST with the same account identifier and `--no-resume` to repair an earlier import. Existing messages remain counted as skipped.

Closes kenn-io#690.


Co-authored-by: Rusty Shackleford <salmonumbrella@users.noreply.github.com>
## What changed

- Replace expensive bare-base64 regex scans with linear byte scans, and skip the data-URI matcher when its prefix is absent.
- Preserve preprocessing output, thresholds, padding, Unicode handling, and generation compatibility.
- Add boundary regressions, differential fuzzing, and reproducible preprocessing benchmarks.

## Why

Embedding builds can spend more time preprocessing messages than calling the embedding endpoint. The counted-repeat base64 patterns carry hundreds of regex states through ordinary text, leaving the endpoint idle between batches. This removes that unnecessary CPU work without changing what gets embedded.

## Usage

No usage change. Existing vector generations remain compatible; no rebuild is required.

Refs kenn-io#497



Co-authored-by: Rusty Shackleford <salmonumbrella@users.noreply.github.com>
…enn-io#1060)

Activity projection can complete for archives affected by kenn-io#1014. On upgraded SQLite archives, `last_modified` has no column default. Bodyless imports could leave it NULL after the original backfill had already run, causing every projection pass to abort before saving its reconciled revisions.

The shared message upsert now stamps `last_modified` on inserts and updates. A new one-time migration fills existing NULLs while preserving other timestamps. This requires one messages-table scan on the first startup after upgrade. The missing-timestamp validation error now identifies the offending message.

The affected messages remain eligible for projection. This change does not introduce a general skip or quarantine policy for invalid activity records.

Fixes kenn-io#1014.


Co-authored-by: Wes McKinney <wesm@users.noreply.github.com>
…io#1048)

## What changed

- Start job runtime budgets after admission to the operation gate. Report waiting jobs as queued, admitted jobs as running, and name the actual gate holder in health status.
- Give scheduled activity projection, attachment packing, and daily attachment maintenance one-minute budgets and allow them to yield to queued work. Resume after committed progress; record a timeout and wait for the next trigger when a pass cannot reach a checkpoint.
- Persist activity reconciliation cursors across passes and restarts. Limit scheduled projection to ten batches, including identity-revision and configuration changes, while manual builds retain their full-pass behavior.
- Make attachment discovery cancellable and checkpoint scheduled verification in windows of 128 packed blobs or 32 MiB of raw content. Continue admitting new loose blobs during verification and preserve full-catalog manual unpacking.
- Document status fields, runtime limits, and resumable maintenance behavior.

## Why

kenn-io#956 keeps scheduled sources on cadence under load and already moves packing out of scheduled syncs. Long activity reconciliation and attachment maintenance can still occupy the shared operation gate while syncs wait. Bounded passes release the gate and resume behind queued work, and accurate admission status makes stalled sources visible.

## Usage

Existing schedules apply. Automatic activity projection and attachment maintenance use the new limits without configuration changes. Manual activity builds and attachment commands retain their existing behavior.

Refs kenn-io#956. The scheduler continues to serialize archive work through the operation gate; this change adds bounded yielding rather than concurrent archive writers.


Co-authored-by: Rusty Shackleford <salmonumbrella@users.noreply.github.com>
…coring (kenn-io#929)

Review identity suggestions from the CLI or MCP using the same evidence and review tokens as the Directory queue. Every HTTP accept or reject requires a current token. Changes to evidence or linked cluster membership require fresh review before a new link is applied. Stale dialogs close while preserving note drafts, and duplicate suggestions retain that review check.

This PR also exposes person merge and CardDAV write tools through MCP. Each write capability has its own opt-in flag: `--allow-identity-decisions`, `--allow-person-merges`, `--allow-carddav-writes`, or `--allow-identity-scoring`. HTTP MCP also requires `--http-allow-writes`. The client must obtain user approval for each confirmation.

Optional scoring is disabled by default. `msgvault person scoring status` lists the exact fields sent to the fixed Jev endpoint at `api.typesafe.ai`: contacts’ names, email addresses, phone numbers, bounded raw identifiers and their scopes, and matching evidence. Enable `[people.identity_scoring]`, configure the credential environment variable and retention declaration, then consent to the current disclosure fingerprint. Run `person scoring run` and inspect `history`.

Scoring creates review suggestions and journal entries. Local blockers prevent unnecessary provider requests; a changed input invalidates an in-flight score and permits rescoring. Network requests release the archive write gate, and interrupted batches return completed results with a redacted error. Scores remain advisory and never accept a match or merge people.

API schema 3.0.0 replaces the token-free accept/reject routes. Upgrade clients and daemon together; older daemons are rejected before archive requests. HTTP confirmation-dependent writes require MCP protocol 2026-07-28 or newer.

Closes kenn-io#928


Co-authored-by: Rusty Shackleford <salmonumbrella@users.noreply.github.com>
…1044)

## What changed

- Keep a completed read snapshot when syncs overlap an export. Preserve its watermarks so the next build can append later messages and repair journaled child rows.
- Reuse existing message shards when participant identity links change alongside new messages. Refresh relationship data from retained facts, new facts, and complete recipients.
- Keep full repairs for changed message facts, account attribution, failed syncs, messages that become exportable below the cached ID boundary, and source-deleted messages above it. Archives without the repair journal retain the conservative fallback. Serving defers population scans to cache maintenance.
- If a sync commits while a build is starting a child-row repair, the build finishes with a full rebuild instead of failing. This covers scheduled, automatic, and CLI builds.
- Document builder memory budgets, thread counts, spill space, and staging requirements.

## Why

A sync during a long export could force another full archive export even when an append or child-row repair was enough. Ordinary identity revision changes combined with new messages also replaced all message shards. These follow-up builds consume time, memory, and disk without changing the retained message facts.

## Usage

No new flags or configuration keys. Builder defaults remain a `2GB` memory budget, at most two threads, and `32GB` of spill. DuckDB's budget covers its buffer manager; process RSS can exceed it. The configuration guide explains when to lower the budget or thread count and how much staging space to allow.




Co-authored-by: Rusty Shackleford <salmonumbrella@users.noreply.github.com>
…1063)

## What changed

- Wrap scoring scans back to the beginning before reporting that scoring is complete. Keep the 128-candidate limit and preserve full-sweep progress across calls and restarts.
- Add the sweep marker to SQLite and PostgreSQL, including upgrades of existing archives. Cover changed evidence, expired backoff and leases, scan boundaries, and incomplete-run reporting.

## Why

A scoring run could report `processed: 0` with no error while a suggestion below its saved cursor was ready to score. Operators had to run scoring again to discover the skipped work.

## Usage

No usage change.

Closes kenn-io#1058.

Follows kenn-io#929, which is now merged. This PR contains only the sweep-completion fix.


Co-authored-by: Rusty Shackleford <salmonumbrella@users.noreply.github.com>
Prepare the documentation for 0.21.0 to become the current release when this PR merges. The changelog leads with upgrade actions and credits all 12 contributors. Product pages, Markdown companions, and guides point to the release and make Graph mail, headless sign-in, mail/chat drafts, and Kata agendas easier to find.

The 10 Web UI screenshots are refreshed on [docs-assets](kenn-io@01626d3). They show the current navigation and relationship view at 1920 × 1080, using the same reviewed fixture. The capture script keeps the archive isolated while preserving the normal build toolchain. Existing OAuth images remain intact.

Contributor credits are also included in the [0.21.0 release](https://github.com/kenn-io/msgvault/releases/tag/v0.21.0).

![Everything with the grouped sidebar, global search, and Save view](https://raw.githubusercontent.com/kenn-io/msgvault/01626d305682608d111074ba33caf5ccb985521b/analytical-light-compact-darwin.png)


Co-authored-by: Wes McKinney <wesm@users.noreply.github.com>
…o#1056)

Dedup now uses archived source facts before label count when earlier survivor rules tie. This keeps a copy's native message ID, threading evidence, and RFC822 Message-ID relevant to selection. Sent-copy eligibility, source preference, raw MIME, and content-equivalent payload completeness retain their existing precedence.

Gmail thread evidence comes from stored conversations, so historical and newly synced copies receive the same treatment without extra metadata writes. A conversation ID equal to the message ID earns no thread point: the archive cannot distinguish a generated fallback from a genuine single-message thread. Preserved Google Groups thread markers, archived reply headers, and resolved reply parents also count; threading contributes at most one point.

`identity discover --provider` includes the authenticated Gmail profile using the selected source's existing OAuth app or service-account credentials:

```bash
msgvault identity discover --source-id 14 --provider
msgvault identity discover --source-id 14 --provider --apply
```

Preview leaves ownership unchanged; `--apply` confirms strong evidence. Provider discovery requires the profile to match the selected source. Ordinary sync only refreshes matching identities already confirmed: a profile mismatch logs a warning and archiving continues.

Removing the last confirmed identity now saves `no_default_identity`, preventing later scheduled syncs from restoring it. Beeper repair and iMazing reimport preserve that choice. Use the source's add command with `--no-default-identity=false` to re-enable automatic confirmation. The guides and survivor diagram describe the selection rules and identity workflow.

Closes kenn-io#397


Co-authored-by: Rusty Shackleford <salmonumbrella@users.noreply.github.com>
Rerunning `msgvault import-whatsapp` on an Apple `ChatStorage.sqlite` now writes only new and changed messages. Edits show up however old the message is, and senders or chats that `LID.sqlite` resolves later move to the right person and conversation.

Until now every rerun rewrote the whole archive, so a handful of new messages cost a full rewrite and fresh change signals for every cache and embedding reader. Each message also took four separate commits, so an interrupted run could leave one without its body or search entry.

Each run still reads every Apple row and derives it as a full import would, then skips it when the stored message, body and search entry already match. A changed message commits in one transaction. Reading only new rows would need saved cursors, an edit window and recovery state, and it still missed late edits. Messages deleted in WhatsApp stay archived.

Closes kenn-io#1050


Co-authored-by: Emmanuel Sérié <eserie@users.noreply.github.com>
Plaud users can archive cloud recordings, transcripts and notes directly in msgvault, then search and browse them alongside other meetings.

`add-plaud` checks the existing source owner through the daemon before opening the browser. It confirms the live Plaud account before replacing credentials. Plaud and Circleback share the OAuth implementation while keeping separate credentials and callbacks.

`sync-plaud` reads complete transcripts, speaker labels and all note tabs. It updates archived meetings in place and preserves previously fetched content while Plaud processing is pending. Limited runs start with the newest recordings, then rotate through the least recently attempted recordings. A failed recording still fails that run, but does not prevent later runs from reaching other recordings. The archive keeps one current rotation map per source; run history retains outcomes and counts without duplicating that map.

Enable Cloud Sync and transcribe recordings in Plaud, then configure an account:

```toml
[[plaud]]
identifier = "personal"
account_email = "you@example.com"
schedule = "0 */6 * * *"
enabled = true
```

```sh
msgvault add-plaud personal
msgvault sync-plaud personal
```

Authorization runs on the daemon host using `localhost:8091/callback/plaud`. The daemon also runs scheduled syncs. `--limit` bounds recordings per run, `--full` repairs archived records, and `--after YYYY-MM-DD` filters dates locally. `--probe` prints tool metadata and counts without recording content.

Each selected recording is reread in full to detect text edits. Sync does not download audio, change Plaud data, or propagate source deletions.

Closes kenn-io#1039.


Co-authored-by: Rusty Shackleford <salmonumbrella@users.noreply.github.com>
…nn-io#1098)

Beeper voice notes in archives upgraded from before msgvault recorded attachment state now go to Docbank for transcripts, as long as their bytes are in the archive. Media discovery skipped Beeper audio with an empty state, so those recordings never left the archive.


Co-authored-by: Rod Boev <rodboev@users.noreply.github.com>
## What changed

- Give `query_sql` an explicit object output schema while keeping both row results and accepted cache-build jobs.
- Cover SQL-enabled catalogs under both write policies and validate both result shapes against the advertised schema.

## Why

Claude Desktop rejects the whole MCP server when `query_sql` omits the required top-level object type from `tools/list`. Both result alternatives were already objects, but that did not satisfy the client’s schema check.

## Usage

No usage change.

Closes kenn-io#1068.


Co-authored-by: Rusty Shackleford <salmonumbrella@users.noreply.github.com>
Adds opt-in Google Calendar event control through the CLI, daemon HTTP API, and MCP: create, update, delete, move, self RSVP, freebusy, and conflicts. Calendar sync remains read-only.

Writes require separate OAuth consent, exact configured calendars, delegated permissions, and current Google access. Guest changes also require invitation permission; notifications default to `none`. Delegated authorization precedes setup checks. Only actual mutations and archive writes take the daemon operation lock.

MCP previews identify the target event for owners and event-read grants. Write-only grants see requested changes with provider details redacted. Approval binds the account, plan, target, and notification mode; conditional provider writes detect later edits. Legacy stdio confirmation works through the SDK, while HTTP mutation tools require protocol `2026-07-28` or newer.

Future-series edits validate bounds, occurrence selection, recurrence compatibility, and retained-event overlaps before shortening a series. Unsupported metadata and detached future exceptions are rejected. A definite replacement failure triggers an attempt to restore the original recurrence using the returned ETag. Unknown outcomes are not retried or rolled back; the result reports restoration failures and includes archive receipts for completed writes.

The API schema is 3.1, with generated Go/browser clients and setup and recovery documentation. Write-only validation can reveal timing constraints; it does not grant event-content access. Whole-calendar exception scanning remains a bounded optimization opportunity.


Co-authored-by: Rusty Shackleford <salmonumbrella@users.noreply.github.com>
…urity] (kenn-io#1106)

This PR contains the following updates:

| Package | Change | [Age](https://docs.renovatebot.com/merge-confidence/) | [Confidence](https://docs.renovatebot.com/merge-confidence/) |
|---|---|---|---|
| [go.opentelemetry.io/otel/sdk](https://redirect.github.com/open-telemetry/opentelemetry-go) | `v1.44.0` → `v1.45.0` | ![age](https://developer.mend.io/api/mc/badges/age/go/go.opentelemetry.io%2fotel%2fsdk/v1.45.0?slim=true) | ![confidence](https://developer.mend.io/api/mc/badges/confidence/go/go.opentelemetry.io%2fotel%2fsdk/v1.44.0/v1.45.0?slim=true) |

---

### OpenTelemetry-Go: Exporter config logging may leak endpoint URLs in info logs
[CVE-2026-81870](https://nvd.nist.gov/vuln/detail/CVE-2026-81870) / [GHSA-8wmf-6v46-5gfg](https://redirect.github.com/advisories/GHSA-8wmf-6v46-5gfg)

<details>
<summary>More information</summary>

#### Details
##### Summary

OpenTelemetry Go versions 1.5.0 through 1.44.0 can include trace exporter endpoint configuration in an internal diagnostic log emitted when an SDK `TracerProvider` is created. The default OpenTelemetry logger does not emit this event. Exposure requires an application to install a logger that enables OpenTelemetry's internal Info-level diagnostics and for someone other than the intended audience to have access to those logs.

The logged configuration can disclose the address of the trace collector and whether the OTLP/HTTP connection is configured as insecure. The Zipkin exporter logs its complete collector URL, so credentials in URL userinfo or tokens in the query string are also disclosed if an application embeds them there. OTLP authentication headers, TLS key material, and exported span data are not included in this log.

Exporter `MarshalLog` implementations that caused this configuration to be included in internal logs were introduced by [`a1fff3c`](https://redirect.github.com/open-telemetry/opentelemetry-go/commit/a1fff3c2588c783d1f3f6fd2315aa2660fc6d330).

##### Details

When `sdk/trace.NewTracerProvider` constructs a provider, it records a `TracerProvider created` internal Info event containing the provider configuration. In affected versions, the configuration's `MarshalLog` methods recursively include:

1. the provider's span processors;
2. each processor's span exporter; and
3. for the OTLP trace exporter, its client configuration.

This causes the following values to be present in the event:

- OTLP trace gRPC: the configured endpoint;
- OTLP trace HTTP: the configured endpoint and the `Insecure` flag; and
- Zipkin: the complete collector URL.

OpenTelemetry Go does not emit this event with its default logger, which only emits errors. An application must explicitly configure a sufficiently verbose logger with `otel.SetLogger`. The required `logr` verbosity is version-dependent:

- versions 1.5.0 through 1.14.x use `V(1)` for this Info event; and
- versions 1.15.0 through 1.44.0 use `V(4)`.

OTLP header configuration is not part of the marshaled object, so credentials supplied with `WithHeaders` or the corresponding environment variables are not exposed. The documented OTLP `WithEndpoint` input is a collector address rather than a credential-bearing URL. The higher-risk case is therefore the Zipkin collector URL, which is retained and logged in full, or an application passing sensitive data in an OTLP endpoint outside the documented format.

##### Proof of concept

The following program demonstrates the behavior with OpenTelemetry Go 1.44.0. It deliberately places credentials and a token in the Zipkin collector URL and enables internal Info logging:

```go
package main

import (
	"bytes"
	"context"
	"fmt"

	"github.com/go-logr/logr/funcr"
	"go.opentelemetry.io/otel"
	"go.opentelemetry.io/otel/exporters/zipkin"
	sdktrace "go.opentelemetry.io/otel/sdk/trace"
)

func main() {
	var logs bytes.Buffer
	otel.SetLogger(funcr.New(func(_, args string) {
		_, _ = logs.WriteString(args)
	}, funcr.Options{Verbosity: 4}))

	exporter, err := zipkin.New(
		"http://user:pass@zipkin.internal:9411/api/v2/spans?token=secret",
	)
	if err != nil {
		panic(err)
	}

	tp := sdktrace.NewTracerProvider(sdktrace.WithBatcher(exporter))
	_ = tp.Shutdown(context.Background())

	fmt.Println(logs.String())
}
```

The `TracerProvider created` event contains:

```text
http://user:pass@zipkin.internal:9411/api/v2/spans?token=secret
```

For versions before 1.15.0, set `funcr.Options{Verbosity: 1}` instead.

##### Impact

This is a conditional disclosure through application logs. Affected applications must enable verbose OpenTelemetry internal diagnostics and configure a trace exporter containing information they do not intend to expose to readers of those logs. In that configuration, a person or system with log access can learn the trace collector address and internal network topology. If credentials or tokens are embedded directly in a Zipkin collector URL, those values can also be recovered from the logs.

There is no exposure with the default OpenTelemetry logger, and the vulnerable log is generated from local application configuration rather than remotely supplied span data. OTLP authentication headers, certificate or private-key contents, and telemetry payloads are not logged by this path.

##### Remediation

Upgrade the affected OpenTelemetry Go modules to version 1.45.0 or later. The fix in [`3a1412d`](https://redirect.github.com/open-telemetry/opentelemetry-go/commit/3a1412d2b3bc4e4231fbeac2ed42117ae541bb38) stops recursively marshaling exporter and client configuration and records their types instead.

If an immediate upgrade is not possible:

- keep OpenTelemetry internal logging below the Info verbosity described above;
- do not embed credentials or tokens in exporter endpoint URLs; use authentication headers or another supported credential mechanism; and
- restrict access to existing logs and rotate any credentials that may already have been recorded.

#### Severity
- CVSS Score: 2.0 / 10 (Low)
- Vector String: `CVSS:4.0/AV:L/AC:L/AT:P/PR:L/UI:N/VC:L/VI:N/VA:N/SC:N/SI:N/SA:N`

#### References
- [https://github.com/open-telemetry/opentelemetry-go/security/advisories/GHSA-8wmf-6v46-5gfg](https://redirect.github.com/open-telemetry/opentelemetry-go/security/advisories/GHSA-8wmf-6v46-5gfg)
- [https://nvd.nist.gov/vuln/detail/CVE-2026-81870](https://nvd.nist.gov/vuln/detail/CVE-2026-81870)
- [https://github.com/open-telemetry/opentelemetry-go/pull/8438](https://redirect.github.com/open-telemetry/opentelemetry-go/pull/8438)
- [https://github.com/open-telemetry/opentelemetry-go/commit/3a1412d2b3bc4e4231fbeac2ed42117ae541bb38](https://redirect.github.com/open-telemetry/opentelemetry-go/commit/3a1412d2b3bc4e4231fbeac2ed42117ae541bb38)
- [https://github.com/open-telemetry/opentelemetry-go/releases/tag/exporters/zipkin/v1.45.0](https://redirect.github.com/open-telemetry/opentelemetry-go/releases/tag/exporters/zipkin/v1.45.0)
- [https://github.com/open-telemetry/opentelemetry-go/releases/tag/sdk/v1.45.0](https://redirect.github.com/open-telemetry/opentelemetry-go/releases/tag/sdk/v1.45.0)
- [https://github.com/advisories/GHSA-8wmf-6v46-5gfg](https://redirect.github.com/advisories/GHSA-8wmf-6v46-5gfg)

This data is provided by the [GitHub Advisory Database](https://redirect.github.com/advisories/GHSA-8wmf-6v46-5gfg) ([CC-BY 4.0](https://redirect.github.com/github/advisory-database/blob/main/LICENSE.md)).
</details>

---

### OpenTelemetry-Go: Exporter config logging may leak endpoint URLs in info logs
[CVE-2026-81870](https://nvd.nist.gov/vuln/detail/CVE-2026-81870) / [GHSA-8wmf-6v46-5gfg](https://redirect.github.com/advisories/GHSA-8wmf-6v46-5gfg) / [GO-2026-6505](https://pkg.go.dev/vuln/GO-2026-6505)

<details>
<summary>More information</summary>

#### Details
##### Summary

OpenTelemetry Go versions 1.5.0 through 1.44.0 can include trace exporter endpoint configuration in an internal diagnostic log emitted when an SDK `TracerProvider` is created. The default OpenTelemetry logger does not emit this event. Exposure requires an application to install a logger that enables OpenTelemetry's internal Info-level diagnostics and for someone other than the intended audience to have access to those logs.

The logged configuration can disclose the address of the trace collector and whether the OTLP/HTTP connection is configured as insecure. The Zipkin exporter logs its complete collector URL, so credentials in URL userinfo or tokens in the query string are also disclosed if an application embeds them there. OTLP authentication headers, TLS key material, and exported span data are not included in this log.

Exporter `MarshalLog` implementations that caused this configuration to be included in internal logs were introduced by [`a1fff3c`](https://redirect.github.com/open-telemetry/opentelemetry-go/commit/a1fff3c2588c783d1f3f6fd2315aa2660fc6d330).

##### Details

When `sdk/trace.NewTracerProvider` constructs a provider, it records a `TracerProvider created` internal Info event containing the provider configuration. In affected versions, the configuration's `MarshalLog` methods recursively include:

1. the provider's span processors;
2. each processor's span exporter; and
3. for the OTLP trace exporter, its client configuration.

This causes the following values to be present in the event:

- OTLP trace gRPC: the configured endpoint;
- OTLP trace HTTP: the configured endpoint and the `Insecure` flag; and
- Zipkin: the complete collector URL.

OpenTelemetry Go does not emit this event with its default logger, which only emits errors. An application must explicitly configure a sufficiently verbose logger with `otel.SetLogger`. The required `logr` verbosity is version-dependent:

- versions 1.5.0 through 1.14.x use `V(1)` for this Info event; and
- versions 1.15.0 through 1.44.0 use `V(4)`.

OTLP header configuration is not part of the marshaled object, so credentials supplied with `WithHeaders` or the corresponding environment variables are not exposed. The documented OTLP `WithEndpoint` input is a collector address rather than a credential-bearing URL. The higher-risk case is therefore the Zipkin collector URL, which is retained and logged in full, or an application passing sensitive data in an OTLP endpoint outside the documented format.

##### Proof of concept

The following program demonstrates the behavior with OpenTelemetry Go 1.44.0. It deliberately places credentials and a token in the Zipkin collector URL and enables internal Info logging:

```go
package main

import (
	"bytes"
	"context"
	"fmt"

	"github.com/go-logr/logr/funcr"
	"go.opentelemetry.io/otel"
	"go.opentelemetry.io/otel/exporters/zipkin"
	sdktrace "go.opentelemetry.io/otel/sdk/trace"
)

func main() {
	var logs bytes.Buffer
	otel.SetLogger(funcr.New(func(_, args string) {
		_, _ = logs.WriteString(args)
	}, funcr.Options{Verbosity: 4}))

	exporter, err := zipkin.New(
		"http://user:pass@zipkin.internal:9411/api/v2/spans?token=secret",
	)
	if err != nil {
		panic(err)
	}

	tp := sdktrace.NewTracerProvider(sdktrace.WithBatcher(exporter))
	_ = tp.Shutdown(context.Background())

	fmt.Println(logs.String())
}
```

The `TracerProvider created` event contains:

```text
http://user:pass@zipkin.internal:9411/api/v2/spans?token=secret
```

For versions before 1.15.0, set `funcr.Options{Verbosity: 1}` instead.

##### Impact

This is a conditional disclosure through application logs. Affected applications must enable verbose OpenTelemetry internal diagnostics and configure a trace exporter containing information they do not intend to expose to readers of those logs. In that configuration, a person or system with log access can learn the trace collector address and internal network topology. If credentials or tokens are embedded directly in a Zipkin collector URL, those values can also be recovered from the logs.

There is no exposure with the default OpenTelemetry logger, and the vulnerable log is generated from local application configuration rather than remotely supplied span data. OTLP authentication headers, certificate or private-key contents, and telemetry payloads are not logged by this path.

##### Remediation

Upgrade the affected OpenTelemetry Go modules to version 1.45.0 or later. The fix in [`3a1412d`](https://redirect.github.com/open-telemetry/opentelemetry-go/commit/3a1412d2b3bc4e4231fbeac2ed42117ae541bb38) stops recursively marshaling exporter and client configuration and records their types instead.

If an immediate upgrade is not possible:

- keep OpenTelemetry internal logging below the Info verbosity described above;
- do not embed credentials or tokens in exporter endpoint URLs; use authentication headers or another supported credential mechanism; and
- restrict access to existing logs and rotate any credentials that may already have been recorded.

#### Severity
- CVSS Score: 2.0 / 10 (Low)
- Vector String: `CVSS:4.0/AV:L/AC:L/AT:P/PR:L/UI:N/VC:L/VI:N/VA:N/SC:N/SI:N/SA:N`

#### References
- [https://github.com/open-telemetry/opentelemetry-go/security/advisories/GHSA-8wmf-6v46-5gfg](https://redirect.github.com/open-telemetry/opentelemetry-go/security/advisories/GHSA-8wmf-6v46-5gfg)
- [https://nvd.nist.gov/vuln/detail/CVE-2026-81870](https://nvd.nist.gov/vuln/detail/CVE-2026-81870)
- [https://github.com/open-telemetry/opentelemetry-go/pull/8438](https://redirect.github.com/open-telemetry/opentelemetry-go/pull/8438)
- [https://github.com/open-telemetry/opentelemetry-go/commit/3a1412d2b3bc4e4231fbeac2ed42117ae541bb38](https://redirect.github.com/open-telemetry/opentelemetry-go/commit/3a1412d2b3bc4e4231fbeac2ed42117ae541bb38)
- [https://github.com/open-telemetry/opentelemetry-go](https://redirect.github.com/open-telemetry/opentelemetry-go)
- [https://github.com/open-telemetry/opentelemetry-go/releases/tag/exporters/zipkin/v1.45.0](https://redirect.github.com/open-telemetry/opentelemetry-go/releases/tag/exporters/zipkin/v1.45.0)
- [https://github.com/open-telemetry/opentelemetry-go/releases/tag/sdk/v1.45.0](https://redirect.github.com/open-telemetry/opentelemetry-go/releases/tag/sdk/v1.45.0)

This data is provided by [OSV](https://osv.dev/vulnerability/GHSA-8wmf-6v46-5gfg) and the [GitHub Advisory Database](https://redirect.github.com/github/advisory-database) ([CC-BY 4.0](https://redirect.github.com/github/advisory-database/blob/main/LICENSE.md)).
</details>

---

### OpenTelemetry-Go: Exporter config logging may leak endpoint URLs in info logs in go.opentelemetry.io/otel/exporters/otlp/otlptrace
[CVE-2026-81870](https://nvd.nist.gov/vuln/detail/CVE-2026-81870) / [GHSA-8wmf-6v46-5gfg](https://redirect.github.com/advisories/GHSA-8wmf-6v46-5gfg) / [GO-2026-6505](https://pkg.go.dev/vuln/GO-2026-6505)

<details>
<summary>More information</summary>

#### Details
OpenTelemetry-Go: Exporter config logging may leak endpoint URLs in info logs in go.opentelemetry.io/otel/exporters/otlp/otlptrace

#### Severity
Unknown

#### References
- [https://github.com/open-telemetry/opentelemetry-go/security/advisories/GHSA-8wmf-6v46-5gfg](https://redirect.github.com/open-telemetry/opentelemetry-go/security/advisories/GHSA-8wmf-6v46-5gfg)
- [https://nvd.nist.gov/vuln/detail/CVE-2026-81870](https://nvd.nist.gov/vuln/detail/CVE-2026-81870)
- [https://github.com/open-telemetry/opentelemetry-go/commit/3a1412d2b3bc4e4231fbeac2ed42117ae541bb38](https://redirect.github.com/open-telemetry/opentelemetry-go/commit/3a1412d2b3bc4e4231fbeac2ed42117ae541bb38)
- [https://github.com/open-telemetry/opentelemetry-go/pull/8438](https://redirect.github.com/open-telemetry/opentelemetry-go/pull/8438)
- [https://github.com/open-telemetry/opentelemetry-go/releases/tag/exporters/zipkin/v1.45.0](https://redirect.github.com/open-telemetry/opentelemetry-go/releases/tag/exporters/zipkin/v1.45.0)
- [https://github.com/open-telemetry/opentelemetry-go/releases/tag/sdk/v1.45.0](https://redirect.github.com/open-telemetry/opentelemetry-go/releases/tag/sdk/v1.45.0)

This data is provided by [OSV](https://osv.dev/vulnerability/GO-2026-6505) and the [Go Vulnerability Database](https://redirect.github.com/golang/vulndb) ([CC-BY 4.0](https://redirect.github.com/golang/vulndb#license)).
</details>

---

### Release Notes

<details>
<summary>open-telemetry/opentelemetry-go (go.opentelemetry.io/otel/sdk)</summary>

### [`v1.45.0`](https://redirect.github.com/open-telemetry/opentelemetry-go/compare/v1.44.0...v1.45.0)

[Compare Source](https://redirect.github.com/open-telemetry/opentelemetry-go/compare/v1.44.0...v1.45.0)

</details>

---

### Configuration

📅 **Schedule**: (UTC)

- Branch creation
  - At any time (no schedule defined)
- Automerge
  - At any time (no schedule defined)

🚦 **Automerge**: Disabled by config. Please merge this manually once you are satisfied.

♻ **Rebasing**: Whenever PR becomes conflicted, or you tick the rebase/retry checkbox.

🔕 **Ignore**: Close this PR and you won't be reminded about this update again.

---

 - [ ] <!-- rebase-check -->If you want to rebase/retry this PR, check this box

---

This PR was generated by [Mend Renovate](https://mend.io/renovate/). View the [repository job log](https://developer.mend.io/github/kenn-io/msgvault).
<!--renovate-debug:eyJjcmVhdGVkSW5WZXIiOiI0NC4xMjUuMSIsInVwZGF0ZWRJblZlciI6IjQ0LjEyNS4xIiwidGFyZ2V0QnJhbmNoIjoibWFpbiIsImxhYmVscyI6W119-->


Co-authored-by: renovate[bot] <renovate[bot]@users.noreply.github.com>
…io#1036)

## What changed

Containers and supervised installs can configure the daemon from environment variables and flags, mount server/remote/Docbank keys, and retain an automatically generated daemon key across restarts. This removes the need for startup scripts that rewrite configuration, generate keys, or seed an MCP config file.

`iface:NAME` binds to a usable address on that interface and fails when none is available. Config edits do not require live interfaces or server credential files; startup checks those resources. Existing home-directory permissions remain unchanged.

Provider keys can be installed with `credentials set` or imported from their configured environment variables. HTTP MCP can use an independent inbound bearer key while connecting to a backend configured entirely from the environment. Runtime overrides and mounted keys stay out of saved TOML. Explicit export-token and setup choices persist even when they match an override.

## Usage

```sh
MSGVAULT_HOME=/data MSGVAULT_BIND_ADDR=0.0.0.0 MSGVAULT_API_PORT=8080 msgvault serve
msgvault credentials set vector.embeddings --from-file /run/secrets/embedding-key
MSGVAULT_REMOTE_URL=https://archive.example.com MSGVAULT_REMOTE_API_KEY_FILE=/run/secrets/daemon-key msgvault mcp --http 0.0.0.0:8081 --http-token-file /run/secrets/mcp-key
```

Persist `data_dir` (the home directory by default) to retain the archive and generated key. Once minted, `tokens/server-api-key` also applies to loopback starts, including Web UI login. Local CLI clients discover it automatically.

Mounted key files must be regular, non-symlink files owned by the process user, with mode `0400` or `0600` on Unix or an owner-only ACL on Windows. Swarm mounts need explicit ownership and mode; default Swarm mounts and Kubernetes Secret volume symlinks are rejected. The configuration guide describes alternatives. A selected credential source that is missing, empty, or unsafe fails without falling back.

## Upgrade note

A stored person-enrichment suppression key now takes precedence over a custom `suppression_key_env`, matching other provider keys. Installations with both must ensure the keys agree before upgrading; using a different key for existing suppression records causes `ErrSuppressionKeyMismatch`.

Related to kenn-io#488. Store submission and Fastmail token-file support remain separate work.


Co-authored-by: Rusty Shackleford <salmonumbrella@users.noreply.github.com>
…enn-io#1115)

Slack archives can now use a token restricted to public channels. Previously,
registration required `search:read`, which can expose private conversations
even when archive filters exclude them. Restricted tokens list all public
channels, including unjoined ones, and find late thread replies through
resumable history walks. This takes more API requests than search. File
downloads are deferred without `files:read`.

Users who want private conversations can grant the corresponding scopes and
choose what to archive. The new `[slack].private_channels` setting and
`--private-channels` override work independently of `dms` and `group_dms`.
All three retain their enabled defaults. Turning a type off preserves its
archived messages and sync progress. Broader tokens continue to archive the
user's memberships.

The setup guide distinguishes archive selection from token permissions and
provides both configurations. It recommends a fresh app when reducing access
because Slack OAuth grants are additive.

Generated with Codex
Co-authored-by: Codex <noreply@openai.com>
…n-io#1005)

This removes msgvault's 425-line copy of Docbank's reranking client, so `msgvault eval` uses Docbank's directly and a fix to it reaches both projects.

The local spend stop goes away with the copy, along with `--rerank-cost-stop-usd`, because Docbank's client reports tokens only after a whole ranking finishes. `--rerank-max-requests` stays and still refuses, before the archive opens, a run whose worst-case request count is over the limit.

Each ranking gets 40 seconds for all of its calls, which keeps 10 seconds per call across the four waves of eight that a 30-candidate per-candidate ranking needs. A failed ranking reports no requests or tokens, so usage counts completed rankings only, and a response without token usage fails the run.

Refs kenn-io#877 (comment), slice 1b (eval arm uses Docbank's Jev client)


Co-authored-by: Rod Boev <rodboev@users.noreply.github.com>
…nn-io#1018)

This removes 733 lines of copied helper code. The whole diff is -125 because of 608 lines of tests that check the shared copies give the same answers as the old ones: "is this address mine" matches the same messages everywhere, full and incremental analytics exports write the same cache, and file search pages the same way.

Small jobs each lived in two to five places, so a fix to one copy missed the others. Reading a server's "retry after" header, saving Codex and Microsoft sign-in tokens, and matching "is this address mine" now have one copy each. So do analytics cache writes, Ctrl-C handling, chat previews, text search and file paging.

Refs kenn-io#1002, slice 5 (smaller copies)


Co-authored-by: Rod Boev <rodboev@users.noreply.github.com>
…enn-io#1064)

Refs kenn-io#534

`carddav.Service` now reaches the server through a `Remote` interface: discover, pull, get, put, delete, and the href for a new card. The CardDAV and Google request code moves to `davRemote` without changes. The service keeps the retry gate, the publication ledger and conflict review, so another backend can reuse them through `NewRemoteService`.

This is the first step of the Microsoft contacts work from kenn-io#534. The Graph backend follows in a separate PR.

No behavior or configuration changes. Against Google Contacts, `main` and this branch found the same book and stored the same 13 cards, with the same first, incremental and full sync results.


Co-authored-by: Mike Campbell <exactmike@users.noreply.github.com>
@arcaputo3
arcaputo3 merged commit 3a65fd8 into main Oct 5, 2026
3 of 30 checks passed
@arcaputo3
arcaputo3 deleted the sync/upstream-2026-10-05 branch October 5, 2026 18:54
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.