Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,7 @@ and this project adheres to [Semantic Versioning](http://semver.org/spec/v2.0.0.
## [Unreleased]

### Added
- `fiboa create-stac-collection` describes the fiboa properties in `table:columns`.
- Added `pixi run check-hcat`, which compares the HCAT mapping tables the converters read online with the taxonomy.
- Added `FiboaDuckDBBaseConverter` for SQL-based conversion of large Parquet sources.
- Added `PerFileBaseConverter` to process multi-file sources incrementally.
Expand Down Expand Up @@ -80,6 +81,7 @@ and this project adheres to [Semantic Versioning](http://semver.org/spec/v2.0.0.
- US-CSB: Editions now cover 2017-2024.

### Removed
- Removed `fiboa publish`. Datasets are published as a Portolan catalog; the README describes the steps.
- CH: Removed the national `ch` converter, whose single licence could not cover the cantons' differing terms; the canton converters replace it, and a Swiss file is `fiboa merge` of their outputs.

### Fixed
Expand Down
91 changes: 48 additions & 43 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -62,7 +62,7 @@ fiboa CLI supports various commands to work with the files:
- [Improve a fiboa Parquet file](#improve-a-fiboa-parquet-file)
- [Update an extension template with new names](#update-an-extension-template-with-new-names)
- [Converter for existing datasets](#converter-for-existing-datasets)
- [Publish datasets to source coop or your own s3 repository](#publish-datasets-to-source-coop-or-your-own-s3-repository)
- [Publishing with Portolan](#publishing-with-portolan)
- [Development](#development)
- [Implement a converter](#implement-a-converter)
- [Run in Docker](#run-in-docker)
Expand Down Expand Up @@ -193,48 +193,53 @@ Use any of the IDs from the list to convert an existing dataset to fiboa:

See [Implement a converter](#implement-a-converter) for details about how to

### Publish datasets to source coop or your own s3 repository

`fiboa publish <dataset> -o <target>`

The publish converts and publishes a fiboa dataset to source coop or your own s3 repository. The target directory
will be filled with the following files:

```
<target>/
<dataset>.parquet
<dataset>.pmtiles # requires working ogr2ogr and tippecanoe
stac/collection.json
README.md # generated if --generate-meta/-gm flag is present
LICENSE.txt # generated if --generate-meta/-gm flag is present
```

This directory is synchronized to the s3 repository (default source.coop/fiboa/data).

**Requirements**: Requires the [aws CLI](https://docs.aws.amazon.com/cli/latest/userguide/getting-started-install.html) to be installed,
and `AWS_ACCESS_KEY_ID` with `AWS_SECRET_ACCESS_KEY` environment variables. Also, for generating the pmtiles file,
it requires [ogr2ogr](https://gdal.org/programs/ogr2ogr.html) and [tippecanoe](https://github.com/mapbox/tippecanoe).

The command executes the following steps:

- `fiboa convert` to generate a fiboa parquet dataset. All convert parameters are passed to the converter.
- `fiboa validate` to validate the fiboa dataset
- creates a <dataset>.pmtiles from the parquet file. Uses ogr2ogr and tippecanoe
- `fiboa create-stac-collection` to create a STAC collection
- `fiboa publish` to publish the fiboa dataset to a source coop or your own s3 repository

Examples:

- `fiboa publish at_crop -o data/at_crop`
- `fiboa publish -c /tmp/cache -gm br_conab -o data/br_conab`

Relevant parameters:

- `--generate-meta/-gm` Generatse the README.md and LICENSE.txt files if absent, based on data-survey and converter properties.
- `--data-url` The URL to the data repository, used when generating the README
- `--s3-upload-path` The `aws s3 sync` target. Defaults to `s3://source.coop/fiboa/data` . Uploading requires the `aws` CLI, and `AWS_ACCESS_KEY_ID` with `AWS_SECRET_ACCESS_KEY` environment variables.

Check `fiboa publish --help` for more details.
### Publishing with Portolan

fiboa datasets are published as a [Portolan](https://github.com/portolan-sdi/portolan-cli) catalog. The fiboa CLI
converts and validates; Portolan writes the STAC metadata, the PMTiles, the checksums and the README, and uploads.
The steps below are written so that an agent can follow them; the
[Portolan skills](https://github.com/portolan-sdi/portolan-skills) cover the Portolan side in more depth.

Inside a Portolan catalog (`portolan init`), for dataset `<id>` and edition `<variant>`:

1. Convert and validate. Name the file after the edition, so each edition is its own asset:
```bash
fiboa convert <id> --variant <variant> -c <cache> -o <id>/<id>-<variant>.parquet
fiboa validate <id>/<id>-<variant>.parquet
```
2. Seed the collection metadata from the converter: title, description, providers, license, fiboa version, and the
columns with a description for every fiboa property. The columns move to the collection; the assets go, as
Portolan adds its own. Only for a new collection; Portolan keeps these fields afterwards.
```bash
fiboa create-stac-collection <id>/<id>-<variant>.parquet -o <id>/stac.json
jq '."table:columns" = .assets.data."table:columns" | del(.assets)' <id>/stac.json > <id>/collection.json
rm <id>/stac.json
```
3. Add the file with its campaign date, and generate the PMTiles (requires
[tippecanoe](https://github.com/felt/tippecanoe)):
```bash
portolan add <id> --datetime <variant>-01-01 --pmtiles
```
4. Describe the dataset from its [data survey](https://github.com/fiboa/data-survey/tree/main/data): usually
`<ID>.md` with the id upper-cased and `_` as `-` (`de_nrw` is `DE-NRW.md`), otherwise the country's file
(`nl_block` is in `NL.md`).
- Run `portolan metadata init <id>` and fill `<id>/.portolan/metadata.yaml`:

| data survey | metadata.yaml |
|---|---|
| Data Provider (Legal Entity) | `providers`, role `producer` (and `licensor`) |
| Homepage, Data URL | `source_url` |
| License | `license`, `license_url` |
| Overview and the dataset's section | `description` |
| Caveats in the text (coverage, preliminary editions) | `known_issues` |
| the fiboa project, publishing this copy | `contact`, and `providers` with role `host` |

- Describe the dataset's own columns in `<id>/collection.json` (`table:columns[].description`) from the
survey's Properties table; the converter's `columns` show which source column each one comes from.
- Run `portolan readme <id>`.
5. Complete what else Portolan asks for until `portolan check` passes, such as a thumbnail (skill
`portolan-thumbnails`).
6. Upload: `portolan push <remote> --collection <id>` (skill `sourcecoop` for Source Cooperative).

## Development

Expand Down
26 changes: 26 additions & 0 deletions fiboa_cli/create_stac.py
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,25 @@

from fiboa_cli.fiboa.version import get_versions

# The fiboa properties mean the same in every dataset; a converter's own columns are described per dataset
DESCRIPTIONS = {
"id": "Unique identifier",
"collection": "The collection identifier",
"inspire:id": "The INSPIRE identifier",
"determination:datetime": "Timestamp of the determination of the field boundary",
Comment on lines +11 to +15
"metrics:area": "Field area in square meters",
"metrics:perimeter": "Field perimeter in meters",
"crop:code_list": "A link to the code list",
"crop:code": "The crop code",
"crop:name": "Crop name in the original language",
"crop:name_en": "Crop name in English",
"hcat:name": "The machine-readable HCAT name of the crop",
"hcat:code": "The 10-digit HCAT code indicating the hierarchy of the crop",
"hcat:name_en": "The HCAT crop name translated into English",
"admin:country_code": "ISO 3166-1 alpha-2 country code",
"admin:subdivision_code": "ISO 3166-2 principal subdivision code (e.g. province or state)",
}


class CreateStacCollection(Base):
temporal_property = "determination:datetime"
Expand Down Expand Up @@ -45,3 +64,10 @@ def create(self, collection: Collection, gdf: GeoDataFrame, *args, **kwargs) ->
data.setdefault("vecorel_extensions", {k: list(v) for k, v in schemas.items()})

return data

def create_from_file(self, *args, **kwargs) -> dict:
stac = super().create_from_file(*args, **kwargs)
for column in stac["assets"]["data"].get("table:columns", []):
if column["name"] in DESCRIPTIONS:
column.setdefault("description", DESCRIPTIONS[column["name"]])
return stac
Loading
Loading