Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
8 changes: 4 additions & 4 deletions .agents/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,17 +6,17 @@ Everything an AI coding agent needs to operate inside this repo. The `.claude

| Path | Role |
|---|---|
| [`settings.json`](settings.json) | **Tracked** allow-list of safe, read-only Bash / git / WebFetch / pipeline commands. A fresh-clone agent gets these pre-approved so it doesn't burn turns on permission prompts. Side-effecting stages (build / fetch / extract / convert / tighten_variant) are intentionally *not* pre-approved here. |
| [`settings.json`](settings.json) | **Tracked** allow-list of safe, read-only Bash / git / WebFetch / pipeline commands. A fresh-clone agent gets these pre-approved so it doesn't burn turns on permission prompts. Side-effecting stages (build / fetch / extract / convert) are intentionally *not* pre-approved here. |
| `settings.local.json` | **Gitignored** per-machine override — additional permissions specific to the local agent's session. Don't commit. |
| [`skills/`](skills/) | 21 invokable skills following the [Agent Skills](https://agentskills.io) standard — wrappers around pipeline entrypoints (`/raincloud-build`, `/raincloud-fetch`, `/raincloud-status`, `/raincloud-validate-manifest`, `/raincloud-list-datasets`, `/raincloud-load`, `/raincloud-publish`, …) and procedural playbooks (`/raincloud-add-dataset`, `/raincloud-add-handler`, `/raincloud-debug-build`, …). See [`skills/README.md`](skills/README.md). |
| [`skills/`](skills/) | 20 invokable skills following the [Agent Skills](https://agentskills.io) standard — wrappers around pipeline entrypoints (`/raincloud-build`, `/raincloud-fetch`, `/raincloud-status`, `/raincloud-validate-manifest`, `/raincloud-list-datasets`, `/raincloud-load`, `/raincloud-publish`, …) and procedural playbooks (`/raincloud-add-dataset`, `/raincloud-add-handler`, `/raincloud-debug-build`, …). See [`skills/README.md`](skills/README.md). |
| [`context/`](context/) | Symlinks back to the repo-root canonical docs (`AGENTS.md`, `SKILLS.md`, `README.md`, `sources.schema.md`) so each `SKILL.md` can pull authoritative guidance via a stable relative path without copying. |
| `scheduled_tasks.lock` | Gitignored — agent-runtime state. |

## Where to start (fresh-clone agent)

1. Read [`../AGENTS.md`](../AGENTS.md) (auto-loaded via `../CLAUDE.md → AGENTS.md` symlink) for the invariants.
2. Run `python -m scripts.pipeline.status --fast --missing-only` to verify the env.
3. Run `python -m scripts.pipeline.validate_manifest` to confirm the manifest is well-formed.
2. Run `python -m raincloud.pipeline.status --fast --missing-only` to verify the env.
3. Run `python -m raincloud.pipeline.validate_manifest` to confirm the manifest is well-formed.
4. Browse [`skills/README.md`](skills/README.md) for the catalog of invokable commands.

## Adding or editing a skill
Expand Down
44 changes: 23 additions & 21 deletions .agents/settings.json
Original file line number Diff line number Diff line change
@@ -1,31 +1,33 @@
{
"$comment": "Tracked agent settings — safe, read-only defaults shared across fresh clones. settings.local.json sits next to this file (gitignored) and overrides per-machine. Side-effecting pipeline stages (build, fetch, extract, convert, tighten_variant) are intentionally NOT pre-approved here — those skills carry disable-model-invocation: true and confirmation lives in AGENTS.md.",
"$comment": "Tracked agent settings shared across fresh clones: read-only queries over the repo and catalog, plus `uv sync --inexact` for environment setup (the one entry here that changes anything, and only the virtualenv; a sync without --inexact silently drops the other extras, so it is not pre-approved). Deliberately NOT here: arbitrary code execution (`python -c`), commands that regenerate derived files (`raincloud.pipeline.docs` rewrites the docs/ scratch copies in a checkout, or observations under the data directory, and holds the store lock; the /raincloud-docs skill pre-approves it while that skill is active), git commands that can delete or rewire (`git branch -D`, `git remote add/set-url/remove`), and the side-effecting pipeline stages (build, fetch, extract, export, convert, hydrate, publish) — those skills carry disable-model-invocation: true and confirmation lives in AGENTS.md. settings.local.json sits next to this file (gitignored) and overrides per-machine; that is where a maintainer adds anything above.",
"permissions": {
"allow": [
"Bash(python -m scripts.pipeline.status *)",
"Bash(python -m scripts.pipeline.status)",
"Bash(python -m scripts.pipeline.validate_manifest *)",
"Bash(python -m scripts.pipeline.validate_manifest)",
"Bash(python -m scripts.pipeline.list_datasets *)",
"Bash(python -m scripts.pipeline.list_datasets)",
"Bash(python -m scripts.pipeline.docs *)",
"Bash(python -m scripts.pipeline.docs)",
"Bash(python -c *)",
"Bash(python3 -c *)",
"Bash(.venv/bin/python -c *)",
"Bash(.venv/bin/python -m scripts.pipeline.status *)",
"Bash(.venv/bin/python -m scripts.pipeline.validate_manifest *)",
"Bash(.venv/bin/python -m scripts.pipeline.docs *)",
"Bash(uv sync*)",
"Bash(uv run python -m scripts.pipeline.status *)",
"Bash(uv run python -m scripts.pipeline.validate_manifest *)",
"Bash(uv run python -m scripts.pipeline.docs *)",
"Bash(python -m raincloud.pipeline.status *)",
"Bash(python -m raincloud.pipeline.status)",
"Bash(python -m raincloud.pipeline.validate_manifest *)",
"Bash(python -m raincloud.pipeline.validate_manifest)",
"Bash(python -m raincloud.pipeline.list_datasets *)",
"Bash(python -m raincloud.pipeline.list_datasets)",
"Bash(.venv/bin/python -m raincloud.pipeline.status *)",
"Bash(.venv/bin/python -m raincloud.pipeline.validate_manifest *)",
"Bash(raincloud list*)",
"Bash(raincloud describe *)",
"Bash(raincloud config show)",
"Bash(raincloud capabilities)",
"Bash(uv sync *--inexact*)",
"Bash(uv run python -m raincloud.pipeline.status *)",
"Bash(uv run python -m raincloud.pipeline.validate_manifest *)",
"Bash(git status*)",
"Bash(git diff*)",
"Bash(git log*)",
"Bash(git show*)",
"Bash(git branch*)",
"Bash(git remote*)",
"Bash(git branch)",
"Bash(git branch -a)",
"Bash(git branch -vv)",
"Bash(git branch --show-current)",
"Bash(git branch --list*)",
"Bash(git remote -v)",
"Bash(git remote show*)",
"Bash(git ls-files*)",
"Bash(git rev-parse*)",
"Bash(git blame*)",
Expand Down
36 changes: 18 additions & 18 deletions .agents/skills/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,24 +6,24 @@ Project-local skills for AI coding agents (Claude Code, plus other tools that fo

## Script wrappers

Wrappers around `python -m scripts.pipeline.<module>`. Side-effecting ones set `disable-model-invocation: true` so Claude won't auto-trigger destructive work; read-only ones are model-invocable.
Wrappers around `python -m raincloud.pipeline.<module>`. Side-effecting ones set `disable-model-invocation: true` so Claude won't auto-trigger destructive work; read-only ones are model-invocable.

| Skill | Wraps | Purpose |
|---|---|---|
| `/raincloud-build` | `scripts.pipeline.build` | Full pipeline (fetch → … → convert) for one or more slugs. |
| `/raincloud-fetch` | `scripts.pipeline.fetch` | Download raw bytes only. |
| `/raincloud-extract` | `scripts.pipeline.extract` | Unpack archives into `_workdir/`. |
| `/raincloud-convert` | `scripts.pipeline.convert` | Stage 7 — emit sibling `.vortex` per spec opt-in. |
| `/raincloud-hydrate` | `scripts.pipeline.hydrate` | Stage 8 (optional, opt-in) — dereference a slug's URL column into a sibling parquet under `parquet-hydrated/`. Side-effecting (outbound HTTP); safety-filter-gated; `disable-model-invocation: true`. |
| `/raincloud-docs` | `scripts.pipeline.docs` | Regenerate derived docs. *(model-invocable — regen is mostly idempotent.)* |
| `/raincloud-tighten-variant` | `scripts.pipeline.tighten_variant` | In-place JSON → VARIANT promotion. |
| `/raincloud-status` | `scripts.pipeline.status` | Per-slug filesystem state (raw / workdir / parquet / vortex / variant-pending). *(read-only, model-invocable.)* |
| `/raincloud-validate-manifest` | `scripts.pipeline.validate_manifest` | Static checks for `sources.json` — JSON Schema + handler-registry / slug-uniqueness / fetch-auth cross-checks. *(read-only, model-invocable.)* |
| `/raincloud-list-datasets` | `scripts.pipeline.list_datasets` | Filter/list slugs by handler / license / fetch-type / reader / vortex / tag / showcase / size / regex. *(read-only, model-invocable.)* |
| `/raincloud-discover` | `scripts.pipeline.list_datasets` | Find "interesting" datasets via the discoverability flags — tag / showcase / size / trait / view. *(read-only, model-invocable.)* |
| `/raincloud-profile` | `scripts.pipeline.profile` | Compute per-column statistics → `outputs/v1/<slug>/profile.json` (opt-in; feeds the TUI detail pane + `list_datasets --inspect`). *(writes `profile.json`; model-invocable.)* |
| `/raincloud-load` | `raincloud.load` (loader API) | Load a prepared dataset (cache → mirror → local build) as a lazy `Dataset`; inspect metadata or materialize. *(`disable-model-invocation: true`.)* |
| `/raincloud-publish` | `scripts.pipeline.publish` | Sync built `outputs/v1/...` artefacts to a mirror, gated on the snapshot sha256. *(side-effecting — `disable-model-invocation: true`.)* |
| `/raincloud-build` | `raincloud.pipeline.build` | Full pipeline (fetch → … → write_canonical → validate → run_exporters) for one or more slugs. |
| `/raincloud-fetch` | `raincloud.pipeline.fetch` | Download raw bytes only. |
| `/raincloud-extract` | `raincloud.pipeline.extract` | Unpack archives into the recipe's scratch directory. |
| `/raincloud-export` | `raincloud.pipeline.export` | Re-derive Parquet/Vortex from the canonical Arrow already on disk, without refetching; the refresh path for one format. |
| `/raincloud-convert` | `raincloud.pipeline.convert` | Re-encode Vortex with the Python writer (v1 catalogs: from Parquet). For v2, prefer `/raincloud-export --format vortex`. |
| `/raincloud-hydrate` | `raincloud.pipeline.hydrate` | Build a `<parent>-hydrated` dataset (URL columns fetched from the open web) with the safe defaults, or write a scratch sample with non-default options (`--limit`/`--block`/`--urlhaus`/`--max-bytes`/`--timeout`/bypass), never published or served. Side-effecting (outbound HTTP); safety-filter-gated; `disable-model-invocation: true`. |
| `/raincloud-docs` | `raincloud.pipeline.docs` | Regenerate derived docs. *(model-invocable — regen is mostly idempotent.)* |
| `/raincloud-status` | `raincloud.pipeline.status` | Per-slug filesystem state (raw / work / arrow / parquet / vortex). *(read-only, model-invocable.)* |
| `/raincloud-validate-manifest` | `raincloud.pipeline.validate_manifest` | Static checks for `sources.json` — JSON Schema + handler-registry / slug-uniqueness / fetch-auth cross-checks. *(read-only, model-invocable.)* |
| `/raincloud-list-datasets` | `raincloud.pipeline.list_datasets` | Filter/list slugs by handler / license / fetch-type / reader / vortex / tag / showcase / size / regex. *(read-only, model-invocable.)* |
| `/raincloud-discover` | `raincloud.pipeline.list_datasets` | Find "interesting" datasets via the discoverability flags — tag / showcase / size / trait / view. *(read-only, model-invocable.)* |
| `/raincloud-profile` | `raincloud.pipeline.profile` | Compute per-column statistics → `outputs/v{n}/<slug>/profile.json` (opt-in; feeds the TUI detail pane + `list_datasets --inspect`). *(writes `profile.json`; model-invocable.)* |
| `/raincloud-load` | `raincloud.load` (loader API) | Load a prepared dataset (cache → mirror; never builds implicitly) as a lazy `Dataset`; inspect metadata or materialize. *(`disable-model-invocation: true`.)* |
| `/raincloud-publish` | `raincloud.pipeline.publish` | Place built artefacts in this machine's store and release the catalog (`--store`), or upload them to a mirror (`--mirror`), gated on the snapshot sha256. *(side-effecting — `disable-model-invocation: true`.)* |

## Procedural playbooks (model-invocable)

Expand All @@ -32,9 +32,9 @@ These guide multi-step procedures from [`SKILLS.md`](../context/SKILLS.md). Defa
| Skill | When to use |
|---|---|
| `/raincloud-add-dataset` | Adding a new dataset to `sources.json` and producing its first build. |
| `/raincloud-add-handler` | Writing a new transform handler under `scripts/pipeline/handlers/`. |
| `/raincloud-add-handler` | Writing a new transform handler under `raincloud/pipeline/handlers/`. |
| `/raincloud-add-kaggle-tos` | Adding a Kaggle dataset gated behind a one-time ToS click-through. |
| `/raincloud-promote-variant` | Picks the right path (in-place vs new-build) for JSON → VARIANT. |
| `/raincloud-promote-variant` | JSON → VARIANT via the transform recipe and a rebuild. |
| `/raincloud-debug-build` | Diagnostic checklist for a failing build — isolate which stage broke. |
| `/raincloud-large-build` | Run a memory- or runtime-heavy build safely (caps, nohup, logging). *(side-effecting — `disable-model-invocation: true`.)* |
| `/raincloud-remove-dataset` | Remove a dataset from the manifest and clean up its outputs. *(destructive — `disable-model-invocation: true`.)* |
Expand Down Expand Up @@ -66,4 +66,4 @@ Reference: <https://code.claude.com/docs/en/skills>. Frontmatter fields used her
- `description` — front-load the key use case (truncated at 1,536 chars in the listing).
- `argument-hint` — autocomplete hint for `/<skill> <args>`.
- `disable-model-invocation` — `true` for side-effecting skills so Claude won't auto-trigger them.
- `allowed-tools` — pre-approve specific tool patterns when the skill is active (e.g. `Bash(python -m scripts.pipeline.build *)`).
- `allowed-tools` — pre-approve specific tool patterns when the skill is active (e.g. `Bash(python -m raincloud.pipeline.build *)`).
26 changes: 5 additions & 21 deletions .agents/skills/raincloud-add-dataset/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,33 +10,17 @@ Steps:

1. **Identify the upstream.** Confirm with the user:
- A stable public URL (prefer the publisher's canonical endpoint over a mirror).
- The license — must permit redistribution-of-derivatives. Check SPDX ID and `source_url`.
- The license, recorded accurately: SPDX ID, `source_url`, `redistribution_permitted`, and `scrape_advisory` for broad-web crawls. A license that forbids redistribution does not keep a dataset out of the catalog; the redistribution gates apply only to `publish --mirror`.
- Approximate row count (used for `expect.rows`; can be `null` on first build).

2. **Append a `DatasetSpec` to `sources.json`** using the Python load-edit-dump pattern from [AGENTS.md](../../context/AGENTS.md#safe-ways-to-edit-sourcesjson) — never `sed`. Start from [`templates/minimal_spec.json`](../../../templates/minimal_spec.json) (every field present with placeholder values) rather than typing one from scratch. Minimal direct-HTTP shape:

```jsonc
{
"slug": "my-dataset",
"short_name": "My Dataset",
"full_name": "My Dataset (publisher attribution)",
"description": "One-line summary.",
"license": { "spdx": "CC0-1.0", "source_url": "...", "redistribution_permitted": true, "attribution_required": false },
"fetch": { "type": "http", "urls": ["https://..."], "auth": null },
"extract": { "type": "passthrough" },
"parse": { "reader": "csv", "options": { "delimiter": "," } },
"transform": { "handler": "tighten_types", "params": {} },
"write": { "output": "my-dataset.parquet", "compression": "zstd", "row_group_size_rows": 1048576, "statistics": true },
"expect": { "rows": 123456 }
}
```
2. **Append a `DatasetSpec` to `sources.json`** using the Python load-edit-dump pattern from [AGENTS.md](../../context/AGENTS.md#editing-sourcesjson) — never `sed`. Copy [`templates/minimal_spec.json`](../../../templates/minimal_spec.json), which has every required field and validates as-is, and edit its placeholders rather than typing a spec from scratch. Keep its `write.*` values: they match the rest of the catalog, and `write.row_group_size_rows` is a row cap on top of the byte-sized row groups, and it wins over `RAINCLOUD_ROW_GROUP_MAX_ROWS` in every Parquet writer (the sidecars receive it through that variable), so a smaller value makes every Parquet export's groups smaller. [sources.schema.md](../../context/sources.schema.md) documents each field.

3. **Validate the manifest.** Invoke `/raincloud-validate-manifest` — sub-second check that the new entry has the right shape, the handler resolves, the slug is unique, and `fetch.type`/`fetch.auth` agree. Catches typos before paying for a fetch.

4. **Run the first build.** Invoke `/raincloud-build <slug> --loose` (the `--loose` is essential when the row count is a guess). If `expect.rows` was wrong, update the manifest with the actual count once the build succeeds.
4. **Run the first build.** Invoke `/raincloud-build <slug>` (validation drift is a warning by default; add `--strict` for a hard gate). If `expect.rows` was wrong, update the manifest with the actual count once the build succeeds.

5. **Regenerate docs.** Invoke `/raincloud-docs` (or just `/raincloud-docs datasets columns_parquet coverage_parquet`).
5. **Regenerate docs.** Invoke `/raincloud-docs` (or just `python -m raincloud.pipeline.docs`).

6. **Optionally opt into Vortex.** If the dataset's types are vortex-compatible, add `"convert": { "vortex": true }` to the spec and run `/raincloud-convert <slug>`.
6. **Choose exports.** Without an `export` block a dataset exports Parquet and Vortex. `export.formats` is the only declaration of which formats it wants: `["parquet"]` leaves Vortex out, `[]` keeps only the canonical Arrow file. Keep a format listed even if its writer fails on the data: the build records the failure as the format's measured unavailability and carries on. `export.priority` prefers a writer, either for every format (`["rs", "py"]`, which must then name a writer for each exported format) or per format (`{"parquet": ["rs", "py"]}`); each format is one file whichever writer makes it. `convert.vortex` is v1-only and rejected in v2. Use `/raincloud-export <slug> --format vortex` to refresh one format without refetching.

If the upstream needs unpacking, set `extract.type` accordingly and pick a parser — see existing specs in `sources.json` for shapes. If the source has nested JSON or row-level processing, you'll likely need a custom handler — see `/raincloud-add-handler`. If it's a Kaggle dataset behind ToS acceptance, see `/raincloud-add-kaggle-tos`.
Loading
Loading