Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions cmd/shithubd/worker.go
Original file line number Diff line number Diff line change
Expand Up @@ -33,6 +33,7 @@ import (
"github.com/tenseleyFlow/shithub/internal/infra/storage"
"github.com/tenseleyFlow/shithub/internal/notifications"
"github.com/tenseleyFlow/shithub/internal/orgs"
repotraffic "github.com/tenseleyFlow/shithub/internal/repos/traffic"
"github.com/tenseleyFlow/shithub/internal/secretscan"
"github.com/tenseleyFlow/shithub/internal/webhook"
"github.com/tenseleyFlow/shithub/internal/webhookrelay"
Expand Down Expand Up @@ -161,6 +162,9 @@ var workerCmd = &cobra.Command{
p.Register(worker.KindJobsPurge, jobs.JobsPurge(jobs.JobsPurgeDeps{
Pool: pool, Logger: logger,
}))
p.Register(repotraffic.KindTrafficPurge, jobs.TrafficPurge(jobs.TrafficPurgeDeps{
Pool: pool, Logger: logger,
}))
p.Register(worker.KindLifecycleSweep, jobs.LifecycleSweep(jobs.LifecycleSweepDeps{
Pool: pool, RepoFS: rfs, Audit: auditRecorder(), Logger: logger,
}))
Expand Down
1 change: 1 addition & 0 deletions deploy/systemd/shithubd-cron.service
Original file line number Diff line number Diff line change
Expand Up @@ -14,3 +14,4 @@ ExecStart=/usr/local/bin/shithubd admin run-job lifecycle:sweep
ExecStart=/usr/local/bin/shithubd admin run-job jobs:purge_completed
ExecStart=/usr/local/bin/shithubd admin run-job webhook:purge_old
ExecStart=/usr/local/bin/shithubd admin run-job workflow:cleanup
ExecStart=/usr/local/bin/shithubd admin run-job traffic:purge
6 changes: 6 additions & 0 deletions docs/internal/actions-runner-api.md
Original file line number Diff line number Diff line change
Expand Up @@ -303,3 +303,9 @@ runner posts terminal job status `cancelled`.
- `shithub_actions_step_timeouts_total`
- `shithub_actions_storage_objects{kind="artifacts|step_logs|hot_log_chunks"}`
- `shithub_actions_storage_bytes{kind="artifacts|step_logs|hot_log_chunks"}`

The gauge observer refreshes queue, runner and object-count gauges every
15 s. `shithub_actions_storage_bytes{kind="hot_log_chunks"}` is refreshed
every 5 min instead: summing `octet_length(chunk)` detoasts every row of
`workflow_step_log_chunks`, so it is scraped on a slower cadence and can
lag the other gauges by up to five minutes.
20 changes: 20 additions & 0 deletions docs/internal/db-indexes.md
Original file line number Diff line number Diff line change
Expand Up @@ -94,6 +94,26 @@ rows for a 1M-row table; "medium" 1–10k; "low" most-of-the-table.
| `signup_ip_throttle` PK `(cidr)` | per-/24 lookup | high | UPSERT |
| `signup_ip_throttle_window_started_idx (window_started_at)` | periodic prune | low | scan-friendly |

## Repo traffic (S38 + retention, availability campaign)

| Index | Covers query | Selectivity | Cost notes |
|---|---|---|---|
| `repo_traffic_daily` PK `(repo_id, day)` | Traffic chart for one repo | high | UPSERT hot path |
| `repo_traffic_daily_day_idx (day DESC)` | 400-day retention purge | low | scan-friendly |
| `repo_traffic_paths` PK `(repo_id, day, path)` | popular-content rollup | high | UPSERT hot path |
| `repo_traffic_paths_day_idx (day)` | 30-day retention purge | low | 0129; without it every purge batch is a seq scan |
| `repo_traffic_referrers` PK `(repo_id, day, referrer)` | referrer rollup | high | UPSERT hot path |
| `repo_traffic_referrers_day_idx (day)` | 30-day retention purge | low | 0129 |
| `repo_traffic_uniques` PK `(repo_id, day, metric, key, visitor_hash)` | dedupe a visitor within a day | high | INSERT ... ON CONFLICT DO NOTHING, one per pageview |
| `repo_traffic_uniques_created_idx (created_at)` | 30-day retention purge | low | 0112; the purge filters this table on `created_at` rather than `day` so it can reuse this index instead of adding a second one to a table that takes an insert per pageview |

The purge (`traffic:purge`) deletes with
`WHERE ctid IN (SELECT ctid FROM <table> WHERE <cutoff column> < $1 LIMIT $2)`.
The `day`/`created_at` index drives the subselect; the outer delete is a
tid scan. Without the index the plan is a sequential scan **per batch**,
which on the production table sizes (1.26 M paths, 1.34 M uniques as of
2026-09-02) is a few hundred full scans per run.

## Future considerations (deferred)

- **`pg_stat_statements` extension.** S37's deploy doc owns the
Expand Down
30 changes: 30 additions & 0 deletions docs/internal/repository-insights.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,6 +25,36 @@ addresses, user agents, authenticated user IDs, and full referrer URLs are not
persisted. Referrers are reduced to an external host and same-site referrers are
dropped.

## Traffic retention

The traffic tables are pruned by the `traffic:purge` worker job, enqueued
nightly from `deploy/systemd/shithubd-cron.service`. Retention windows live in
`internal/repos/traffic/purge.go`:

| Table | Window | Cutoff column | Why |
|---|---|---|---|
| `repo_traffic_uniques` | 30 days | `created_at` | one row per visitor digest per repo/day/metric |
| `repo_traffic_paths` | 30 days | `day` | one row per distinct path per repo/day; crawlers inflate this hardest |
| `repo_traffic_referrers` | 30 days | `day` | one row per external host per repo/day |
| `repo_traffic_daily` | 400 days | `day` | one row per repo/day; the only long-term history, and small enough to keep |

Thirty days is deliberately more than double the fourteen the Traffic UI reads
(`traffic.DefaultWindowDays`), so a purge can never eat a bar the chart would
draw. `repo_traffic_uniques` is filtered on `created_at` rather than `day`
because that is the column it already has an index on, and the request path
stamps both from the same instant.

The job deletes in batches of 5,000 rows, each its own statement, up to 2,000
batches per table per run; the payload (`retention_days`, `daily_retention_days`,
`batch_size`, `max_batches`) overrides any of that for an ad-hoc
`shithubd admin run-job traffic:purge`. Nothing is done in one big transaction:
the tables were left unpruned from 2026-05-18 until the 2026-09-02 availability
sitrep, by which point they held 881 MB of a 988 MB database, and a single
DELETE over that backlog would have locked the write path for minutes. A run
that stops on the batch cap re-enqueues itself so the backlog drains without
waiting for the next cron beat. Re-running is always safe — the cutoff is
recomputed from the clock and rows inside the window are never touched.

## Refresh Flow

`push:process` enqueues `repo:insights_recalc` whenever the repository default
Expand Down
15 changes: 12 additions & 3 deletions docs/internal/retro/2026-09-02-availability-sitrep.md
Original file line number Diff line number Diff line change
Expand Up @@ -153,12 +153,21 @@ The verification items below are still the operator's.
- [x] Key the anonymous HTML tier by `/24` for repo history/blob/raw
routes (Meta rotates within `57.141.2.0/24`) — applied to the
whole anonymous tier, not just those routes
- [ ] Retention job for `repo_traffic_paths` / `repo_traffic_uniques`
(14-day window, matches the Traffic UI) + one-off prune migration
- [x] Retention job for `repo_traffic_paths` / `repo_traffic_uniques`
(plus `_referrers`) — `traffic:purge`, nightly from
`shithubd-cron.service`, 30-day window rather than the UI's 14 so a
purge can never truncate the chart; `repo_traffic_daily` keeps 400
days. Deletes run in 5k-row batches, capped per run, so the first
pass over the 2.5 M-row backlog never holds a long transaction —
which is also why the backfill is the job's first run and not a
bulk DELETE in a migration. 0129 adds the `day` indexes the purge
needs (`repo_traffic_uniques` reuses its existing `created_at`
index). See `docs/internal/repository-insights.md`.
- [ ] Cache per-entry last-commit for the code tab (single
`git log --name-only` walk, or an LRU keyed by tree OID) and
cache `rev-list --count` / recursive `ls-tree` per head OID
- [ ] `actionsobserver`: drop the `octet_length` sum or run it every 5 min
- [x] `actionsobserver`: the `octet_length` sum now runs every 5 min on its
own cadence; the count and queue-depth gauges stay at 15 s

### Phase 4 — observability and docs

Expand Down
6 changes: 6 additions & 0 deletions docs/internal/runbooks/actions.md
Original file line number Diff line number Diff line change
Expand Up @@ -187,6 +187,12 @@ Important metrics:
- `shithub_actions_log_chunk_bytes_total{location="server"}`
- `shithub_actions_storage_objects{kind="artifacts|step_logs|hot_log_chunks"}`
- `shithub_actions_storage_bytes{kind="artifacts|step_logs|hot_log_chunks"}`

The gauge observer refreshes queue, runner and object-count gauges every
15 s. `shithub_actions_storage_bytes{kind="hot_log_chunks"}` is refreshed
every 5 min instead: summing `octet_length(chunk)` detoasts every row of
`workflow_step_log_chunks`, so it is scraped on a slower cadence and can
lag the other gauges by up to five minutes.
- `shithub_actions_step_timeouts_total`

The committed dashboard JSON lives at:
Expand Down
1 change: 1 addition & 0 deletions docs/internal/worker.md
Original file line number Diff line number Diff line change
Expand Up @@ -41,6 +41,7 @@ backstop poll (every 5s by default) covers dropped notifications.
| `workflow:cleanup` | cron / manual ad-hoc | retention cutoff + idempotent deletes |
| `trending:compute` | recurring self-scheduled S42 job | append-only snapshots |
| `org:scheduled_reminder_sweep` | cron / manual ad-hoc | reminder delivery rows |
| `traffic:purge` | cron / manual ad-hoc | retention cutoff recomputed per run; batched deletes |

Adding a new kind: write the handler in `internal/worker/jobs/<kind>.go`,
add the `Kind` constant to `internal/worker/types.go`, register it in
Expand Down
110 changes: 99 additions & 11 deletions internal/infra/metrics/actionsobserver.go
Original file line number Diff line number Diff line change
Expand Up @@ -11,31 +11,95 @@ import (

const actionsRunnerStaleAfter = 60 * time.Second

// defaultActionsInterval is the cadence for the cheap gauges when the caller
// does not pick one.
const defaultActionsInterval = 15 * time.Second

// actionsStorageBytesInterval bounds how often the hot log-chunk byte sum
// runs. `sum(octet_length(chunk))` has to scan and detoast every row of
// workflow_step_log_chunks, which costs the same whether or not anything is
// running; at the 15s cadence of the other gauges it was a standing load on
// the database. Chunk volume moves slowly enough that a 5 minute gauge is
// still useful.
const actionsStorageBytesInterval = 5 * time.Minute

// ObserveActions starts a goroutine that periodically refreshes DB-backed
// Actions gauges. The goroutine exits when ctx is canceled.
//
// interval drives the queue, runner and object-count gauges. The hot
// log-chunk byte sum is refreshed on the slower actionsStorageBytesInterval
// cadence; see refreshActionLogChunkBytes.
func ObserveActions(ctx context.Context, pool *pgxpool.Pool, interval time.Duration) {
if pool == nil {
return
}
if interval <= 0 {
interval = 15 * time.Second
interval = defaultActionsInterval
}
slowEvery := ticksBetween(interval, actionsStorageBytesInterval)
t := time.NewTicker(interval)
go func() {
refreshActions(ctx, pool)
t := time.NewTicker(interval)
defer t.Stop()
for {
select {
case <-ctx.Done():
observeActionsLoop(ctx, t.C, slowEvery,
func(ctx context.Context) { refreshActionsFast(ctx, pool) },
func(ctx context.Context) { refreshActionLogChunkBytes(ctx, pool) },
)
}()
}

// ticksBetween returns how many ticks of length tick must elapse between two
// runs of a task that should run at most once per every. It rounds up, so the
// task never runs more often than requested, and never returns less than 1.
func ticksBetween(tick, every time.Duration) int {
if tick <= 0 || every <= tick {
return 1
}
n := int((every + tick - 1) / tick)
if n < 1 {
return 1
}
return n
}

// observeActionsLoop runs fast on every tick and slow once every slowEvery
// ticks. Both run once up front so the gauges are populated before the first
// tick. It returns when ctx is canceled or ticks is closed.
func observeActionsLoop(ctx context.Context, ticks <-chan time.Time, slowEvery int, fast, slow func(context.Context)) {
if slowEvery < 1 {
slowEvery = 1
}
fast(ctx)
slow(ctx)
sinceSlow := 0
for {
select {
case <-ctx.Done():
return
case _, ok := <-ticks:
if !ok {
return
case <-t.C:
refreshActions(ctx, pool)
}
fast(ctx)
sinceSlow++
if sinceSlow >= slowEvery {
sinceSlow = 0
slow(ctx)
}
}
}()
}
}

// refreshActions refreshes every Actions gauge, cheap and expensive alike.
// The observer loop splits the two cadences apart; this is the one-shot form.
func refreshActions(ctx context.Context, pool *pgxpool.Pool) {
if pool == nil {
return
}
refreshActionsFast(ctx, pool)
refreshActionLogChunkBytes(ctx, pool)
}

func refreshActionsFast(ctx context.Context, pool *pgxpool.Pool) {
if pool == nil {
return
}
Expand Down Expand Up @@ -155,13 +219,16 @@ GROUP BY r.id, r.name, r.status, r.capacity, r.last_heartbeat_at, r.draining_at,
ActionsRunnerStaleTotal.Set(stale)
}

// refreshActionStorageGauges publishes the object counts for all three storage
// kinds plus the two byte sums that read a plain integer column. The
// hot_log_chunks byte sum is deliberately absent: it is the only one that has
// to detoast, so refreshActionLogChunkBytes owns that gauge.
func refreshActionStorageGauges(ctx context.Context, pool *pgxpool.Pool) {
ActionsStorageObjects.WithLabelValues("artifacts").Set(0)
ActionsStorageObjects.WithLabelValues("step_logs").Set(0)
ActionsStorageObjects.WithLabelValues("hot_log_chunks").Set(0)
ActionsStorageBytes.WithLabelValues("artifacts").Set(0)
ActionsStorageBytes.WithLabelValues("step_logs").Set(0)
ActionsStorageBytes.WithLabelValues("hot_log_chunks").Set(0)

rows, err := pool.Query(ctx, `
SELECT 'artifacts'::text AS kind, count(*)::double precision, COALESCE(sum(byte_count), 0)::double precision
Expand All @@ -171,7 +238,7 @@ SELECT 'step_logs'::text AS kind, count(*)::double precision, COALESCE(sum(log_b
FROM workflow_steps
WHERE log_object_key IS NOT NULL
UNION ALL
SELECT 'hot_log_chunks'::text AS kind, count(*)::double precision, COALESCE(sum(octet_length(chunk)), 0)::double precision
SELECT 'hot_log_chunks'::text AS kind, count(*)::double precision, 0::double precision
FROM workflow_step_log_chunks`)
if err != nil {
return
Expand All @@ -184,6 +251,27 @@ FROM workflow_step_log_chunks`)
return
}
ActionsStorageObjects.WithLabelValues(kind).Set(objects)
if kind == "hot_log_chunks" {
continue
}
ActionsStorageBytes.WithLabelValues(kind).Set(bytes)
}
}

// refreshActionLogChunkBytes publishes shithub_actions_storage_bytes for the
// hot chunk table. Every row is a bytea that Postgres has to fetch out of the
// TOAST heap to measure, so this runs on actionsStorageBytesInterval rather
// than with the cheap gauges.
func refreshActionLogChunkBytes(ctx context.Context, pool *pgxpool.Pool) {
if pool == nil {
return
}
var bytes float64
err := pool.QueryRow(ctx, `
SELECT COALESCE(sum(octet_length(chunk)), 0)::double precision
FROM workflow_step_log_chunks`).Scan(&bytes)
if err != nil {
return
}
ActionsStorageBytes.WithLabelValues("hot_log_chunks").Set(bytes)
}
Loading
Loading