Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .github/copilot-instructions.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,7 @@ can be judged from the diff text alone.

- **Localization is all-or-nothing.** A diff that adds a key to
`birdnet_analyzer/lang/en.json`, or a `loc.localize("...")` call with a new key,
must add that key to **all** files in `birdnet_analyzer/lang/` (currently 10) with a
must add that key to **all** files in `birdnet_analyzer/lang/` with a
real translation — English text copied into `de.json` is a defect, not a
placeholder. The files stay sorted with 4-space indent (a json load/dump round-trip
with `ensure_ascii=False, indent=4, sort_keys=True`), and every translation keeps
Expand Down
8 changes: 7 additions & 1 deletion .github/workflows/documentation.yml
Original file line number Diff line number Diff line change
Expand Up @@ -25,6 +25,11 @@ on:
required: false
type: boolean
default: false
python_version:
description: "Python to build with (3.11 for tags pinning tensorflow 2.15)"
required: false
type: string
default: "3.13"

permissions:
contents: write
Expand Down Expand Up @@ -95,7 +100,8 @@ jobs:
ref: ${{ steps.target.outputs.ref }}
- uses: actions/setup-python@v7
with:
python-version: "3.13"
# Old tags pin dependencies that have no wheels for current Python.
python-version: ${{ inputs.python_version || '3.13' }}
- name: Set up uv
uses: astral-sh/setup-uv@v10.0.1
with:
Expand Down
208 changes: 70 additions & 138 deletions AGENTS.md
Original file line number Diff line number Diff line change
@@ -1,161 +1,93 @@
# AGENTS.md

Guidance for coding agents working in this repository. It is the single source of
truth for agent instructions — `CLAUDE.md` just points here.
truth for agent instructions — `CLAUDE.md` just points here. It holds only what the
codebase does not already state: anything you can read off `pyproject.toml`, a
workflow file or a directory listing is deliberately not repeated here.

Personal or machine-specific preferences (which virtualenv to use, whether an agent
may commit) do **not** belong in this file. Put those in an untracked
`CLAUDE.local.md` / `AGENTS.local.md`, which `.gitignore` already covers.
may commit) belong in an untracked `CLAUDE.local.md` / `AGENTS.local.md`, which
`.gitignore` already covers.

## Setup

Python >= 3.11 (CI tests 3.11–3.13 on Ubuntu, macOS and Windows).

```bash
pip install -e .[dev] # ruff + pytest + docs
pip install -e .[dev,embeddings,train,gui-tests] # what CI installs
git submodule update --init --recursive # test audio fixtures
```

Extras: `gui`, `train`, `embeddings`, `docs`, `tests`, `gui-tests`, `dev`, `all`.

The test audio lives in the `tests/data` submodule; without it most tests cannot run.
`CONTRIBUTING.md` covers updating the submodule ref.

### Models are downloaded on first use

Inference comes from the `birdnet` dependency, which downloads model files on first
use — roughly 120 MB (acoustic), 35 MB (geo) and 380 MB (perch-v2). Two consequences
worth knowing before running the suite:

- `birdnet` reads `BIRDNET_APP_DATA` **at import time** and puts the downloads there,
skipping any file that already exists. CI sets it to a workspace path and caches
that directory; `Dockerfile` bakes the models into the image the same way.
- Every test has a 120s timeout, and a cold download can exceed it on a slow link.
CI therefore pre-downloads outside pytest (`ci.yml`, "Pre-download birdnet models").
Do the same locally if a first run times out:
`python -c "import birdnet; birdnet.load('acoustic', '2.4', 'tf', lang='en_us')"`

## Commands

- **Tests**: `python -m pytest`. Single test:
`python -m pytest tests/analyze/test_analyze.py::test_name`. Config lives in
`[tool.pytest.ini_options]`.
- **Lint**: `ruff check` and `ruff format`. Line length 88. Ruff is pinned to the same
version in the `[dev]` extra and in `.github/workflows/lint.yml` — bump both
together.
- **Docs**: `pip install -e .[docs]` then `sphinx-build -E docs _build`. The CLI
reference is generated from `birdnet_analyzer/cli.py` via sphinx-argparse, which is
why the docs workflow also triggers on that file.
- **Run the GUI**: `python -m birdnet_analyzer.gui`
- **Run a CLI**: `python -m birdnet_analyzer.analyze` (likewise `.species`,
`.segments`, `.embeddings`, `.search`, `.train`); installed as `birdnet-analyze`
etc., plus `birdnet-evaluate` and the `birdnet-gui` GUI script.

## CI

Six workflows in `.github/workflows/`: `lint.yml` and `ci.yml` (the test matrix) run
on every PR to `main`, `documentation.yml` on docs changes and on release,
`docker-build.yml` on Dockerfile/`pyproject.toml` changes and on release, and
`publish.yml` / `test-publish.yml` on release. All of them use `paths:` filters, so a
PR touching only docs will not run the test matrix.

The published docs are **versioned**: `documentation.yml` deploys each release to
`/vX.Y.Z/` on `gh-pages` (mirrored at `/stable/`, the default landing spot), pushes
to `main` deploy to `/dev/`, and old releases can be backfilled via the workflow's
manual trigger (`tag`, plus `set_stable` when that tag is the newest release).
Prereleases get their own directory but never become `/stable/`. The site-root
redirect, 404 handler and version switcher live in `docs/_site/`, which is deployed
to the `gh-pages` root and excluded from the Sphinx build.

Two things worth knowing before touching this:

- **`/stable/` only exists once a release has been deployed to the versioned site.**
Until then the root redirect and the 404 handler fall back to the newest version
directory and then to `/dev/`, so the site still works, but it is serving
unreleased docs. Backfill the current release (manual trigger, `set_stable`
ticked) right after the versioning workflow first lands on `main`.
- **Backfilling only reaches `v2.1.0`–`v2.4.0`.** `v2.0.0` and `v2.0.0-rc` have a
`docs/conf.py` but no `docs` extra, so no Sphinx gets installed; every `v1.x` tag
predates `pyproject.toml` entirely. (The bare `1.4.0` tag is also rejected by the
workflow's `vX.Y.Z` pattern.)

Run `ruff check` and `python -m pytest` before handing work back.
`pip install -e .[dev]` gets ruff, pytest and the docs toolchain; CI additionally
installs `embeddings,train,gui-tests`.

**The test audio is a git submodule.** Without
`git submodule update --init --recursive` most tests cannot run. `CONTRIBUTING.md`
covers updating the ref.

**Models download on first use.** Inference comes from the `birdnet` dependency,
which fetches roughly 120 MB (acoustic), 35 MB (geo) and 380 MB (perch-v2) into the
directory named by `BIRDNET_APP_DATA` — **read at import time** — skipping files
that already exist. CI points it at a cached workspace path; `Dockerfile` bakes the
models into the image. Tests have a 120 s timeout that a cold download can exceed,
which is why CI pre-downloads outside pytest. Do the same locally if a first run
times out:
`python -c "import birdnet; birdnet.load('acoustic', '2.4', 'tf', lang='en_us')"`

## Before handing work back

Run `ruff check` and `python -m pytest`. **The ruff version is pinned in two
places** — the `[dev]` extra and `lint.yml` — and has to be bumped in both.

## Architecture

Python package (`birdnet_analyzer`) built on top of the `birdnet` library, which
provides the actual model inference.

- **Feature subpackages** (`analyze/`, `embeddings/`, `search/`, `species/`,
`segments/`, `train/`): each follows the pattern `core.py` (public API, re-exported
in `__init__.py`), `cli.py` (argparse entry point), `__main__.py`, plus an optional
`utils.py`. Shared argparse helpers live in the top-level `cli.py`.
- **`evaluation/` is the exception**: it has no `core.py`/`cli.py` pair. Its `main()`
lives in `__init__.py` and the logic sits under `assessment/` and `preprocessing/`.
Don't assume the other subpackages' shape when working there.
- **Shared top-level modules**: `config.py` (global constants/defaults), `audio.py`,
`model.py`/`model_utils.py`, `utils.py`, `settings.py` (paths, persisted settings),
`logs.py` (logging setup — the package itself only attaches a `NullHandler` and
leaves configuration to the CLI/GUI entry points), `params.py` (reads the
`*-params.csv` files a run writes back into keyword arguments; shared by the CLI,
where they become argument defaults, and the GUI).
- **GUI** (`gui/`): Gradio app wrapped in a pywebview window. One module per tab
exposing a `build_*_tab` function, assembled in `gui/__init__.py:main` via
`gui.utils.open_window`.
- **Tests** (`tests/`): mirror the package layout, e.g. `tests/analyze/`,
`tests/gui/`. There is no `conftest.py`.
`birdnet_analyzer` wraps the `birdnet` library, which does the actual inference.

Feature subpackages (`analyze/`, `embeddings/`, `search/`, `species/`, `segments/`,
`train/`) follow `core.py` (public API, re-exported in `__init__.py`) + `cli.py` +
`__main__.py`. **`evaluation/` does not**: its `main()` lives in `__init__.py` and
the logic sits under `assessment/` and `preprocessing/`.

`params.py` reads the `*-params.csv` files a run writes back into keyword arguments,
shared by the CLI (where they become argument defaults) and the GUI. `logs.py`
leaves logging configuration to the CLI/GUI entry points — the package itself only
attaches a `NullHandler`. The GUI is a Gradio app in a pywebview window, one module
per tab exposing `build_*_tab`, assembled in `gui/__init__.py:main`.

### State that lives outside the repo

`settings.APPDIR` resolves to a per-OS user data directory named
`BirdNET-Analyzer-GUI` (`%APPDATA%` on Windows, `~/.local/share` on Linux,
`~/Library/Application Support` on macOS). GUI settings (`gui-settings.json`),
tab state (`state.json`) and the log files live there, **not** in the checkout. So GUI
behaviour can differ between machines because of leftover local state, and deleting
that directory is the way to test a first-run experience.
`settings.APPDIR` is a per-OS user data directory (`BirdNET-Analyzer-GUI`) holding
GUI settings, tab state (`state.json`) and the logs — **not** the checkout. GUI
behaviour can therefore differ between machines because of leftover local state, and
deleting that directory is how you test a first-run experience.

## Conventions that are easy to get wrong

- **The GUI test environment has no pywebview and no plotly.** CI's `gui-tests` extra
installs gradio but neither pywebview nor plotly (both are `gui`-only), so no GUI
module may import `webview` or `plotly` at module import time — `gui/utils.py`
imports `webview` inside the dialog functions and `open_window` only, and its
`_WINDOW` annotation is a string under `TYPE_CHECKING`. Keep it that way; the
tests import `gui.utils` unstubbed on purpose. A test that builds species-list
controls must stub `gu.plot_map_scatter_mapbox` (it draws a plotly map). A local
venv with the `gui` extra hides both problems, so before pushing GUI tests run them
once with `webview`/`plotly` made unimportable.
- **Localization is all-or-nothing.** `lang/*.json` holds one file per language
(currently 10), located via `settings.LANG_DIR`. A new UI string has to be added to
*every* language file with a real translation — not the English text copied over.
Write the files with a json load/dump round-trip
(`ensure_ascii=False, indent=4, sort_keys=True`). `tests/gui/test_language.py`
- **The GUI test environment has no pywebview and no plotly** — `gui-tests` installs
gradio, both others are `gui`-only. So no GUI module may import `webview` or
`plotly` at module import time: `gui/utils.py` imports `webview` inside the dialog
functions and `open_window` only, and its `_WINDOW` annotation is a string under
`TYPE_CHECKING`. The tests import `gui.utils` unstubbed on purpose. A test that
builds species-list controls must stub `gu.plot_map_scatter_mapbox` (it draws a
plotly map). A local venv with the `gui` extra hides both problems, so run GUI
tests once with `webview`/`plotly` made unimportable before pushing.
- **Localization is all-or-nothing.** `lang/*.json` holds one file per language. A
new UI string has to be added to *every* file with a real translation — not the
English text copied over. Write them with a json load/dump round-trip
(`ensure_ascii=False, indent=4, sort_keys=True`); `tests/gui/test_language.py`
enforces that formatting, key parity across languages, and that each translation
keeps the `str.format` fields of its `en.json` source.
- **A GUI setting has three ways to get its value, not one.** Every persisted control
(`TabState.persist`) is set (1) at build time from `state.json`, (2) by the user, and
(3) programmatically by presets and loaded `*-params.csv` files
(`TabState.updates_for`, which sets `value` only). So when one control's state
depends on another (the sensitivity slider is disabled for BirdNET 3.0/Perch, the
locale dropdown offers only the model's languages, …), cover all three: pass the
dependency into the builder for the initial state, handle it in a `.change` handler
(which gradio fires for programmatic updates too — `.input` does not), and make the
handler also *reset* the dependent value if the loaded one is no longer valid, since
the batch update landed before the handler ran. Doing only (2) is the classic gap.
- **A GUI setting has three ways to get its value, not one.** Every persisted
control (`TabState.persist`) is set (1) at build time from `state.json`, (2) by the
user, and (3) programmatically by presets and loaded `*-params.csv` files
(`TabState.updates_for`, which sets `value` only). So when one control depends on
another (the sensitivity slider disabled for BirdNET 3.0/Perch, the locale dropdown
offering only the model's languages, …), cover all three: pass the dependency into
the builder for the initial state, handle it in a `.change` handler (gradio fires
it for programmatic updates too — `.input` does not), and have the handler *reset*
the dependent value when the loaded one is no longer valid, since the batch update
landed before the handler ran. Doing only (2) is the classic gap.
- **Code comments: only when the code cannot say it.** Default to none. A comment
earns its place only for a constraint, a non-obvious *why*, or a measured value
that justifies a bound — something a reader could not recover from the code and
its names. If the code already explains itself, delete the comment rather than
keep it. No history ("once was", "used to fail", "previously") and no narration
of what the next line does or how the code was arrived at — that belongs in
commit messages and the changelog. The same bar applies to test comments: the
earns its place for a constraint, a non-obvious *why*, or a measured value that
justifies a bound. No history ("used to", "previously"), no narration of the next
line — that belongs in commit messages and the changelog. Same bar for tests: the
test name and its assertions are the explanation.
- **Don't widen the ruff config to make a fix pass.** The `select`/`ignore` lists in
`pyproject.toml` are deliberate.
- **Packaging is an allow-list, not a deny-list.** What ships is decided by the
explicit `packages` and `package-data` lists in `[tool.setuptools]`; a **new
subpackage or data file has to be added there or it silently won't be installed**.
Arbitrary top-level files are not distributed, so they need no `MANIFEST.in` entry
(`prune tests` is the one line there that actually does work). `.dockerignore` is
separate and does need updating for new top-level cruft.
- **Packaging is an allow-list.** `[tool.setuptools]`'s `packages` / `package-data`
decide what ships, so a new subpackage or data file that is not listed there is
silently missing from the install. `.dockerignore` is separate and needs its own
update for new top-level cruft.