diff --git a/CHANGELOG.md b/CHANGELOG.md index ed5131f0..fe6b53c7 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -8,6 +8,7 @@ and this project adheres to [Semantic Versioning](http://semver.org/spec/v2.0.0. ## [Unreleased] ### Added +- `fiboa create-stac-collection` describes the fiboa properties in `table:columns`. - Added `pixi run check-hcat`, which compares the HCAT mapping tables the converters read online with the taxonomy. - Added `FiboaDuckDBBaseConverter` for SQL-based conversion of large Parquet sources. - Added `PerFileBaseConverter` to process multi-file sources incrementally. @@ -80,6 +81,7 @@ and this project adheres to [Semantic Versioning](http://semver.org/spec/v2.0.0. - US-CSB: Editions now cover 2017-2024. ### Removed +- Removed `fiboa publish`. Datasets are published as a Portolan catalog; the README describes the steps. - CH: Removed the national `ch` converter, whose single licence could not cover the cantons' differing terms; the canton converters replace it, and a Swiss file is `fiboa merge` of their outputs. ### Fixed diff --git a/README.md b/README.md index 57982654..6d90e008 100644 --- a/README.md +++ b/README.md @@ -62,7 +62,7 @@ fiboa CLI supports various commands to work with the files: - [Improve a fiboa Parquet file](#improve-a-fiboa-parquet-file) - [Update an extension template with new names](#update-an-extension-template-with-new-names) - [Converter for existing datasets](#converter-for-existing-datasets) - - [Publish datasets to source coop or your own s3 repository](#publish-datasets-to-source-coop-or-your-own-s3-repository) + - [Publishing with Portolan](#publishing-with-portolan) - [Development](#development) - [Implement a converter](#implement-a-converter) - [Run in Docker](#run-in-docker) @@ -193,48 +193,53 @@ Use any of the IDs from the list to convert an existing dataset to fiboa: See [Implement a converter](#implement-a-converter) for details about how to -### Publish datasets to source coop or your own s3 repository - -`fiboa publish -o ` - -The publish converts and publishes a fiboa dataset to source coop or your own s3 repository. The target directory -will be filled with the following files: - -``` -/ - .parquet - .pmtiles # requires working ogr2ogr and tippecanoe - stac/collection.json - README.md # generated if --generate-meta/-gm flag is present - LICENSE.txt # generated if --generate-meta/-gm flag is present -``` - -This directory is synchronized to the s3 repository (default source.coop/fiboa/data). - -**Requirements**: Requires the [aws CLI](https://docs.aws.amazon.com/cli/latest/userguide/getting-started-install.html) to be installed, -and `AWS_ACCESS_KEY_ID` with `AWS_SECRET_ACCESS_KEY` environment variables. Also, for generating the pmtiles file, -it requires [ogr2ogr](https://gdal.org/programs/ogr2ogr.html) and [tippecanoe](https://github.com/mapbox/tippecanoe). - -The command executes the following steps: - -- `fiboa convert` to generate a fiboa parquet dataset. All convert parameters are passed to the converter. -- `fiboa validate` to validate the fiboa dataset -- creates a .pmtiles from the parquet file. Uses ogr2ogr and tippecanoe -- `fiboa create-stac-collection` to create a STAC collection -- `fiboa publish` to publish the fiboa dataset to a source coop or your own s3 repository - -Examples: - -- `fiboa publish at_crop -o data/at_crop` -- `fiboa publish -c /tmp/cache -gm br_conab -o data/br_conab` - -Relevant parameters: - -- `--generate-meta/-gm` Generatse the README.md and LICENSE.txt files if absent, based on data-survey and converter properties. -- `--data-url` The URL to the data repository, used when generating the README -- `--s3-upload-path` The `aws s3 sync` target. Defaults to `s3://source.coop/fiboa/data` . Uploading requires the `aws` CLI, and `AWS_ACCESS_KEY_ID` with `AWS_SECRET_ACCESS_KEY` environment variables. - -Check `fiboa publish --help` for more details. +### Publishing with Portolan + +fiboa datasets are published as a [Portolan](https://github.com/portolan-sdi/portolan-cli) catalog. The fiboa CLI +converts and validates; Portolan writes the STAC metadata, the PMTiles, the checksums and the README, and uploads. +The steps below are written so that an agent can follow them; the +[Portolan skills](https://github.com/portolan-sdi/portolan-skills) cover the Portolan side in more depth. + +Inside a Portolan catalog (`portolan init`), for dataset `` and edition ``: + +1. Convert and validate. Name the file after the edition, so each edition is its own asset: + ```bash + fiboa convert --variant -c -o /-.parquet + fiboa validate /-.parquet + ``` +2. Seed the collection metadata from the converter: title, description, providers, license, fiboa version, and the + columns with a description for every fiboa property. The columns move to the collection; the assets go, as + Portolan adds its own. Only for a new collection; Portolan keeps these fields afterwards. + ```bash + fiboa create-stac-collection /-.parquet -o /stac.json + jq '."table:columns" = .assets.data."table:columns" | del(.assets)' /stac.json > /collection.json + rm /stac.json + ``` +3. Add the file with its campaign date, and generate the PMTiles (requires + [tippecanoe](https://github.com/felt/tippecanoe)): + ```bash + portolan add --datetime -01-01 --pmtiles + ``` +4. Describe the dataset from its [data survey](https://github.com/fiboa/data-survey/tree/main/data): usually + `.md` with the id upper-cased and `_` as `-` (`de_nrw` is `DE-NRW.md`), otherwise the country's file + (`nl_block` is in `NL.md`). + - Run `portolan metadata init ` and fill `/.portolan/metadata.yaml`: + + | data survey | metadata.yaml | + |---|---| + | Data Provider (Legal Entity) | `providers`, role `producer` (and `licensor`) | + | Homepage, Data URL | `source_url` | + | License | `license`, `license_url` | + | Overview and the dataset's section | `description` | + | Caveats in the text (coverage, preliminary editions) | `known_issues` | + | the fiboa project, publishing this copy | `contact`, and `providers` with role `host` | + + - Describe the dataset's own columns in `/collection.json` (`table:columns[].description`) from the + survey's Properties table; the converter's `columns` show which source column each one comes from. + - Run `portolan readme `. +5. Complete what else Portolan asks for until `portolan check` passes, such as a thumbnail (skill + `portolan-thumbnails`). +6. Upload: `portolan push --collection ` (skill `sourcecoop` for Source Cooperative). ## Development diff --git a/fiboa_cli/create_stac.py b/fiboa_cli/create_stac.py index ae291d01..72bfc38f 100644 --- a/fiboa_cli/create_stac.py +++ b/fiboa_cli/create_stac.py @@ -7,6 +7,25 @@ from fiboa_cli.fiboa.version import get_versions +# The fiboa properties mean the same in every dataset; a converter's own columns are described per dataset +DESCRIPTIONS = { + "id": "Unique identifier", + "collection": "The collection identifier", + "inspire:id": "The INSPIRE identifier", + "determination:datetime": "Timestamp of the determination of the field boundary", + "metrics:area": "Field area in square meters", + "metrics:perimeter": "Field perimeter in meters", + "crop:code_list": "A link to the code list", + "crop:code": "The crop code", + "crop:name": "Crop name in the original language", + "crop:name_en": "Crop name in English", + "hcat:name": "The machine-readable HCAT name of the crop", + "hcat:code": "The 10-digit HCAT code indicating the hierarchy of the crop", + "hcat:name_en": "The HCAT crop name translated into English", + "admin:country_code": "ISO 3166-1 alpha-2 country code", + "admin:subdivision_code": "ISO 3166-2 principal subdivision code (e.g. province or state)", +} + class CreateStacCollection(Base): temporal_property = "determination:datetime" @@ -45,3 +64,10 @@ def create(self, collection: Collection, gdf: GeoDataFrame, *args, **kwargs) -> data.setdefault("vecorel_extensions", {k: list(v) for k, v in schemas.items()}) return data + + def create_from_file(self, *args, **kwargs) -> dict: + stac = super().create_from_file(*args, **kwargs) + for column in stac["assets"]["data"].get("table:columns", []): + if column["name"] in DESCRIPTIONS: + column.setdefault("description", DESCRIPTIONS[column["name"]]) + return stac diff --git a/fiboa_cli/publish.py b/fiboa_cli/publish.py deleted file mode 100644 index 6ddcfa6e..00000000 --- a/fiboa_cli/publish.py +++ /dev/null @@ -1,445 +0,0 @@ -import json -import os -import re -import sys -from datetime import date -from functools import cache -from pathlib import Path - -import click -import requests -import spdx_license_list -from vecorel_cli.basecommand import BaseCommand, runnable -from vecorel_cli.cli.options import VECOREL_TARGET -from vecorel_cli.encoding.auto import create_encoding - -from .convert import ConvertData -from .converters import Converters -from .create_stac import CreateStacCollection -from .registry import Registry -from .validate import ValidateData - -STAC_EXTENSION = "https://stac-extensions.github.io/web-map-links/v1.2.0/schema.json" -DESCRIPTIONS = { - "id": "Unique identifier", - "collection": "The collection identifier", - "inspire:id": "The INSPIRE identifier", - "determination:datetime": "Timestamp of the determination of the field boundary", - "metrics:area": "Field area in square meters", - "metrics:perimeter": "Field perimeter in square meters", - "crop:code_list": "A link to the code list", - "crop:code": "The crop code", - "crop:name": "Crop name in the original language", - "crop:name_en": "Crop name in English", - "hcat:name": "The machine-readable HCAT name of the crop", - "hcat:code": "The 10-digit HCAT code indicating the hierarchy of the crop", - "hcat:name_en": "The HCAT crop name translated into English", - "admin:country_code": "ISO 3166-1 alpha-2 country code.", - "admin:subdivision_code": "ISO 3166-2 principal subdivision code (e.g. province or state)", -} - -is_windows = os.name == "nt" - - -class Publish(BaseCommand): - cmd_name = "publish" - cmd_help = f"Convert and publish a {Registry.project} dataset to source coop." - url_base = "https://data.source.coop/fiboa/data" - - @staticmethod - def get_cli_args(): - return { - **ConvertData.get_cli_args(), - "target": VECOREL_TARGET(folder=True), - "generate_meta": click.option( - "--generate-meta", - "-gm", - is_flag=True, - type=click.BOOL, - help="Generate README.txt and LICENSE.txt for the dataset if not present.", - default=False, - ), - "data_url": click.option( - "--data-url", - type=click.STRING, - help="When generating documentation, this is the link to the data.", - ), - "s3_upload_path": click.option( - "--s3-upload-path", - type=click.STRING, - help="Upload to this path on S3. By default it's the source coop fiboa data repository.", - ), - "yes": click.option( - "--yes", - "-y", - is_flag=True, - type=click.BOOL, - help="Answer yes to all questions.", - default=False, - show_default=True, - ), - "data_survey_url": click.option( - "--data-survey-url", - type=click.STRING, - help="URL to the data survey markdown file.", - default=os.getenv("FIBOA_DATA_SURVEY"), - show_default=True, - ), - "editor": click.option( - "--editor", - type=click.STRING, - help="Editor to use when editing generated files.", - default=os.getenv("EDITOR", "edit" if is_windows else "nano"), - show_default=True, - ), - "converted_by": click.option( - "--converted-by", - type=click.STRING, - help="Name of the person or organization that converted the data.", - default=os.getenv("FIBOA_CONVERTED_BY"), - show_default=True, - ), - } - - @staticmethod - def get_cli_callback(cmd): - def callback(dataset, *args, **kwargs): - return Publish(dataset).run(*args, **kwargs) - - return callback - - def __init__(self, dataset: str, data_url=None, s3_upload_path=None): - super().__init__() - self.cmd_title = f"Publish {dataset}" - self.dataset = dataset - self.data_url = data_url or f"{self.url_base}/{self.dataset}" - self.s3_upload_path = ( - s3_upload_path or f"s3://us-west-2.opendata.source.coop/fiboa/data/{self.dataset}/" - ) - - try: - self.converter = Converters().load(self.dataset) - except (ImportError, NameError, OSError, RuntimeError, SyntaxError) as e: - raise Exception(f"Converter for '{self.dataset}' not available or faulty: {e}") from e - - def exc(self, cmd): - assert os.system(cmd) == 0 - - def check_command(self, cmd, name=None): - if os.system(f"{cmd} --version") != 0: - self.error(f"Missing command {cmd}. Please install {name or cmd}") - sys.exit(1) - - def download_data_survey(self, base, **kwargs): - data_survey = ( - kwargs.get("data_survey_url") - or f"https://raw.githubusercontent.com/fiboa/data-survey/refs/heads/main/data/{base}.md" - ) - response = requests.get(data_survey) - if not response.ok: - self.warning( - f"Missing data survey {base}.md at {data_survey}. Falling back to converter declared properties." - ) - else: - return response.text - - @cache - def collect_meta_data(self, parquet_file, **kwargs): - base = self.dataset.replace("_", "-").upper() - data = { - "provider": self.converter.provider, - "license": self.converter.license, - "projection": "", - "homepage": "", - "submitter": "Fiboa project", - "header": "", - } - text = self.download_data_survey(base, **kwargs) - mapping = { - "data provider (legal entity)": "provider", - "submitter (affiliation)": "submitter", - } - properties = {} - if text: - data["header"] = ( - f"\n- **Data Survey:** https://github.com/fiboa/data-survey/blob/main/data/{base}.md" - ) - data.update( - { - mapping.get(a.lower(), a.lower()): b - for a, b in re.findall(r"- \*\*(.+?):\*\* (.+?)\n", text) - } - ) - properties = { - a.lower(): b.strip() - for a, b in re.findall(r"\n\|\s*(\w+)[^|]*\|[^|]*\|[^|]*\|([^|]*)\|", text) - } - try: - # Try read projection from parquet metadata - meta = create_encoding(parquet_file).get_geoparquet_metadata() - crs = meta["columns"]["geometry"]["crs"] - data["projection"] = f"{crs['id']['authority']}:{crs['id']['code']} ({crs['name']})" - except Exception: - pass - converted_by = kwargs.get("converted_by") - if converted_by: - data["submitter"] = converted_by - - assert data["provider"], "Cannot determine data provider from converter or data survey." - return data, properties - - def readme_attribute_table(self, stac_data, properties): - def description(name): - m = self.converter.columns - reverse = dict(zip(m.values(), m.keys())) - return ( - properties.get(reverse.get(name)) - or properties.get(name) - or DESCRIPTIONS.get(name, "") - ) - - cols = [["Property", "**Data Type**", "Description"]] + [ - [ - s["name"], - re.search(r"\w+", s["type"])[0], - description(s["name"]), - ] - for s in stac_data["assets"]["data"]["table:columns"] - if s["name"] not in ("geometry", "bbox", "collection") - ] - widths = [max(len(c[i]) for c in cols) for i in range(3)] - aligned_cols = [[f" {c:<{w}} " for c, w in zip(row, widths)] for row in cols] - aligned_cols.insert(1, ["-" * (w + 2) for w in widths]) - return "\n".join(["|" + "|".join(cols) + "|" for cols in aligned_cols]) - - def make_license(self, parquet_file, **kwargs): - text = "" - try: - data, properties = self.collect_meta_data(parquet_file, **kwargs) - text = data["license"] - if getattr(self.converter, "license") not in (None, "", data["license"]): - text += "\n" + self.converter.license + "\n" - - found = False - for _license in (data["license"], self.converter.license): - if not _license or "<(https://" in _license: - continue - - # Include full-license text - _license = _license.upper() - if _license in spdx_license_list.LICENSES: - response = requests.get( - f"https://raw.githubusercontent.com/spdx/license-list-data/refs/heads/main/text/{_license}.txt" - ) - if response.ok: - found = True - text += f"\n\n{response.text}\n" - break - if not found: - self.warning(f"License {text} could not be found in SPDX license list") - - except Exception as e: - self.exception(e) - return text - - def make_readme(self, parquet_file, file_name, stac, **kwargs): - version = Registry.get_version() - converter = self.converter - with open(stac) as f: - stac_data = json.load(f) - count = stac_data["assets"]["data"]["table:row_count"] - data, properties = self.collect_meta_data(parquet_file, **kwargs) - columns = self.readme_attribute_table(stac_data, properties) - urls = converter.get_urls() or "manually downloaded file" - urls = urls.keys() if isinstance(urls, dict) else [urls] - downloaded_urls = "\n".join([(" - " + url) for url in urls]) - - return f"""# Field boundaries for {converter.short_name} - -Provides {count} official field boundaries from {converter.short_name}. -It has been converted to a fiboa GeoParquet file from data obtained from {data["provider"]}. - -- **Source Data Provider:** [{data["provider"]}]({data["homepage"]}) -- **Converted by:** {data["submitter"]} -- **License:** {data["license"]} -- **Projection:** {data["projection"]}{data["header"]} - ---- - -- [Download the data as fiboa GeoParquet]({self.data_url}/{file_name}.parquet) -- [STAC Browser](https://radiantearth.github.io/stac-browser/#/external/data.source.coop/fiboa/data/{self.dataset}/stac/collection.json) -- [STAC Collection]({self.data_url}/stac/collection.json) -- [PMTiles]({self.data_url}/{file_name}.pmtiles) - -## Columns - -{columns} - -## Lineage - -- Data downloaded on {date.today()} from: -{downloaded_urls} -- Converted to GeoParquet using [fiboa-cli](https://github.com/fiboa/cli), version {version} -""" - - @runnable - def publish( - self, - target, - generate_meta=False, - yes=False, - data_survey_url=None, - editor=None, - converted_by=None, - **kwargs, - ): - """ - You need GDAL 3.8 or later (for ogr2ogr) with libgdal-arrow-parquet, tippecanoe, and AWS CLI - - https://gdal.org/ - - https://github.com/felt/tippecanoe - - https://aws.amazon.com/cli/ - """ - Path(target).mkdir(parents=True, exist_ok=True) - - file_name = self.dataset - # the choice convert() makes, so the README lists the source files of this edition - self.converter.select_variant(kwargs["variant"]) - kwargs["variant"] = self.converter.variant - if kwargs["variant"]: - file_name += f"-{kwargs['variant']}" - parquet_file = Path(target) / f"{file_name}.parquet" - - has_write_access = bool( - os.getenv("AWS_ACCESS_KEY_ID") and os.getenv("AWS_SECRET_ACCESS_KEY") - ) - - stac_file = Path(target) / "stac" / "collection.json" - - ## Create parquet file - if not parquet_file.exists(): - self.info(f"Converting file for {self.dataset} to {parquet_file}") - ConvertData(self.dataset).run(parquet_file, **kwargs) - self.success(f"Converted file for {self.dataset} to {parquet_file}") - else: - self.success(f"Using existing file {parquet_file} for {self.dataset}") - - ## Validate parquet file, we only want to publish valid files - self.info(f"Validating {parquet_file}") - ValidateData().validate(parquet_file, num=-1) - self.log("\n => VALID\n", "success") - - ## Create STAC collection.json - self.create_stac_collection(target, file_name, parquet_file, stac_file) - - if generate_meta: - self.generate_meta( - target, - file_name, - stac_file, - data_survey_url=data_survey_url, - converted_by=converted_by, - yes=yes, - editor=editor, - ) - - self.generate_pmtiles(target, file_name, parquet_file) - if not has_write_access: - self.info("Get your credentials through the source coop organization.") - self.info("Login to AWS Console and generate an access key:") - self.info( - " - In AWS console, click on account (right top) press 'Security credentials'," - ) - self.info(" - Go to 'Access keys' and press 'Create access key'") - self.info( - " - Run `export AWS_ACCESS_KEY_ID=<> AWS_SECRET_ACCESS_KEY=<>`\n" - " (Linux/Mac only) where you copy-paste the access key and secret to <>.", - ) - self.error("Please set AWS_ environment variables for uploading") - return - self.upload_to_aws(target) - - def create_stac_collection(self, target, file_name, parquet_file, stac_file): - p_stac = Path(stac_file) - if p_stac.exists() and p_stac.stat().st_mtime >= Path(parquet_file).stat().st_mtime: - return - - self.success(f"Creating STAC collection.json for {parquet_file}") - p_stac.parent.mkdir(exist_ok=True) - CreateStacCollection().create_cli(parquet_file, stac_file) - - Path(target, "stac").mkdir(parents=True, exist_ok=True) - data = json.load(open(stac_file, "r")) - assert data["id"] == self.dataset, ( - f"Wrong collection dataset id: {data['id']} != {self.dataset}, for {stac_file}" - ) - - data["assets"]["data"]["href"] = f"{self.data_url}/{file_name}.parquet" - - if STAC_EXTENSION not in data["stac_extensions"]: - data["stac_extensions"].append(STAC_EXTENSION) - - if not any(d.get("rel") == "pmtiles" for d in data["links"]): - data["links"].append( - { - "href": f"{self.data_url}/{file_name}.pmtiles", - "type": "application/vnd.pmtiles", - "rel": "pmtiles", - } - ) - - with open(stac_file, "w", encoding="utf-8") as f: - json.dump(data, f, indent=2) - - def generate_meta(self, target, file_name, stac_file, **kwargs): - parquet_file = Path(target) / f"{file_name}.parquet" - for required in ("README.md", "LICENSE.txt"): - path = Path(target) / required - if not path.exists(): - self.warning(f"Missing {required}. Generating at {path}") - if required == "README.md": - text = self.make_readme( - parquet_file, - file_name=file_name, - stac=stac_file, - **kwargs, - ) - else: - text = self.make_license(parquet_file, **kwargs) - self.info( - f"\nGenerated the following file {required}:\n{'-' * 80}\n\n{text}\n{'-' * 80}\n" - ) - action = ( - "C" - if kwargs.get("yes") - else input("Do you want to Continue (C), Edit (E) or Abort (A)?") - ) - if action.lower() not in "ce": - self.warning("Bailing out") - sys.exit(1) - with open(path, "w") as f: - f.write(text) - editor = kwargs.get("editor") - if action.lower() == "e" and editor: - os.system(f"{editor} {path}") - - def generate_pmtiles(self, target, file_name, parquet_file): - if is_windows: - self.warning( - "PMTiles generation through tippecanoe is not supported on Windows, skipping." - ) - return - - pm_file = Path(target) / f"{file_name}.pmtiles" - if not pm_file.exists(): - self.info("Running ogr2ogr | tippecanoe") - self.check_command("tippecanoe") - self.check_command("ogr2ogr", name="GDAL") - self.exc( - f"ogr2ogr -t_srs EPSG:4326 -f geojson /vsistdout/ {str(parquet_file)} | tippecanoe -zg --projection=EPSG:4326 -o {str(pm_file)} -l {self.dataset} --drop-densest-as-needed" - ) - - def upload_to_aws(self, target): - self.info("Uploading to aws") - - self.check_command("aws") - self.exc(f"aws s3 sync --exclude '.*' {target} {self.s3_upload_path}") diff --git a/fiboa_cli/registry.py b/fiboa_cli/registry.py index 5b61137e..02da9a85 100644 --- a/fiboa_cli/registry.py +++ b/fiboa_cli/registry.py @@ -34,7 +34,6 @@ def register_commands(self): from .describe import DescribeFile from .improve import ImproveData from .merge import MergeDatasets - from .publish import Publish from .rename_extension import RenameExtension from .validate import ValidateData from .validate_schema import ValidateSchema @@ -49,7 +48,6 @@ def register_commands(self): DescribeFile, ImproveData, MergeDatasets, - Publish, RenameExtension, ValidateData, ValidateSchema, diff --git a/tests/data-files/publish/BE-VLG-survey.md b/tests/data-files/publish/BE-VLG-survey.md deleted file mode 100644 index 9ac80ea3..00000000 --- a/tests/data-files/publish/BE-VLG-survey.md +++ /dev/null @@ -1,78 +0,0 @@ -# Vlaanderen, Belgium - -## Submission Details - -- **Submitter (Affiliation):** Matthias Mohr -- **Data Provider (Legal Entity):** Agriculture and Marine Fisheries Agency of the Flemish government (Government) -- **Homepage:** https://landbouwcijfers.vlaanderen.be/open-geodata-landbouwgebruikspercelen -- **Alternative URL:** https://www.vlaanderen.be/datavindplaats/catalogus/landbouwgebruikspercelen-lv-2022 - -## Overview - -Since 2020, the Department of Agriculture and Fisheries has been publishing a more extensive set of data related to agricultural use plots (from the 2008 campaign). - -From 2023, the downloadable dataset of agricultural use plots will also include the specialization given by the company (= company typology) and that is given to the plots of the company. Based on the typology, the companies are divided into 4 major specializations: arable farming, horticulture, livestock farming and mixed farms. The specialization of each company is calculated annually according to a European method and is based on the standard output of the various agricultural productions on the company. It is therefore an economic specialization and not a reflection of all agricultural production on the company. - -## Data & Metadata - -- **URL:** https://landbouwcijfers.vlaanderen.be/open-geodata-landbouwgebruikspercelen -- **Documentation:** contained in the ZIP packages -- **File Format:** GeoPackage / Shapefile -- **Projection:** EPSG:31370 (Belgian Lambert 72) -- **License:** CC-0 (described as "Publiek" and "Toegang zonder voorwaarden") - -### Properties - -Some of the documented fields are missing in the GeoPackage. These are marked with "(missing)". - -| Property | **Data Type** | Constraints | Description | -|-----------------------|---------------|----------------------------|----------------------------------------------------------------------------------------------| -| fid | integer | | Identifier | -| BT_OMSCH | string | 200 chars | Business type (economic specialization) | -| BT_BRON | string | 50 chars | Source of the business type (year of calculation or specialization indicated) | -| GRAF_OPP | number | | Area (ha, accurate to 1m²) | -| REF_ID | integer | | Unique identification number for the field. | -| GWSCOD_V | string | 5 chars (digits), nullable | Pre-cultivation code | -| GWSNAM_V | string | 90 chars, nullable | Pre-cultivation name | -| GWSCOD_H | string | 5 chars (digits), nullable | Main cultivation/crop code | -| GWSNAM_H | string | 90 chars, nullable | Main cultivation/crop name | -| GWSGRPH_LB | string | 150 chars, nullable | Main cultivation/crop group name | -| GWSCOD_N | string | 5 chars (digits), nullable | First cultivation/crop code | -| CWSNAM_N | string | 90 chars, nullable | First cultivation/crop name | -| GWSCOD_N2 | string | 5 chars (digits), nullable | Second cultivation/crop code | -| GWSNAM_N2 | string | 90 chars, nullable | Second cultivation/crop name | -| AMKM (missing) | string | | Agri-environment code | -| AMKM_LB (missing) | string | | Agri-environment name | -| ECOREGELING (missing) | string | | Eco-regulation code | -| ECOR_LB (missing) | string | | Eco-regulation name | -| BLS (missing) | string | | Planting subsidy code (forest farming systems) | -| BLS_LB (missing) | string | | Planting subsidy name (forest farming systems) | -| GESP_PM | string | 11 chars, nullable | Specialized production method | -| GESP_PM_LB | string | 150 chars, nullable | Description of specialized production method | -| BIOCERT (missing) | string | `J` or `N` | Plot under bio-control with a bio-control body. | -| ERO_NAM | string | 20 chars, | Erosion color code for the field | -| STAT_BGV | string | 2 chars, nullable | Status Permanent Grassland under greening (BG) | -| MEERJARIG_GRASLAND | string | | Status Perennial Grassland (MG6 or higher). Example: MG16 = 16th year grassland | -| LANDBSTR | string | 2 chars, nullable | Agricultural region in which the center of the field is located | -| STAT_AAR | string | 10 chars, nullable | Status Potatoes, follow up rotation duty | -| PCT_EKBG | string | 10 chars, nullable | Percentage range of field that is ecologically sensitive permanent pasture. Example: `0-10%` | -| PCT_WETVEEN | string | 10 chars, nullable | Percentage range of field that is wetland and/or peatland. Example: `0-10%` | -| PRC_GEM | string | 30 chars | Municipality in which the center of the field is located | -| PRC_NIS | string | 5 chars (digits) | NIS code of the municipality in which the center of the field is located | -| X_REF | number | | X coordinate of the center of the field (Lambert) | -| Y_REF | number | | Y coordinate of the center of the field (Lambert) | -| WGS84_LG | string | 11 chars | Longitude of the center of the field (WGS84). Example: `3°21'44"` | -| WGS84_BG | string | 11 chars | Latitude of the center of the field (WGS84). Example: `51°11'39"` | - -Note: Many integer-like numbers are encoded as strings. - -## API - -The open data viewer https://geopunt.be/ shows the data in a viewer (search term: landbouwgebruikspercelen) -See https://www.vlaanderen.be/datavindplaats/catalogus/landbouwgebruikspercelen-lv-2022 for more info - -| Standard | URL | Documentation | -|--------------|-------------------------------------------------------------|------------------------------------------------------------------------------------------------------| -| OGC WFS | https://geo.api.vlaanderen.be/Landbgebrperc/wfs | https://www.vlaanderen.be/datavindplaats/catalogus/wfs-landbouwgebruikspercelen | -| OGC Features | https://geo.api.vlaanderen.be/Landbgebrperc/ogc/features/v1 | https://metadata.vlaanderen.be/srv/dut/catalog.search#/metadata/01f408db-df8a-49a2-8ce4-0f66b8efe17b | -| OGC WMS | https://geo.api.vlaanderen.be/ALV/wms | https://www.vlaanderen.be/datavindplaats/catalogus/wms-departement-landbouw-en-visserij | diff --git a/tests/test_command_modules.py b/tests/test_command_modules.py index 0db55672..880163c2 100644 --- a/tests/test_command_modules.py +++ b/tests/test_command_modules.py @@ -29,4 +29,5 @@ def test_registry_registers_fiboa_commands(): R.instance.register_commands() names = {getattr(c, "__name__", str(c)) for c in R.instance.commands} - assert {"publish", "improve", "create-stac-collection", "merge"} <= names + assert {"improve", "create-stac-collection", "merge"} <= names + assert "publish" not in names # replaced by Portolan, see the README diff --git a/tests/test_create_stac.py b/tests/test_create_stac.py index a1debaf4..9c55356c 100644 --- a/tests/test_create_stac.py +++ b/tests/test_create_stac.py @@ -2,7 +2,7 @@ from vecorel_cli.vecorel.util import load_file -from fiboa_cli.create_stac import CreateStacCollection +from fiboa_cli.create_stac import DESCRIPTIONS, CreateStacCollection from fiboa_cli.registry import Registry @@ -44,3 +44,16 @@ def pop_path(*args, value=None): expected["vecorel_extensions"]["de_nrw"].sort() assert created_file == expected + + +def test_fiboa_columns_are_described(tmp_folder: Path): + from fiboa_cli.create_geoparquet import CreateGeoParquet + + parquet = tmp_folder / "fiboa.parquet" + CreateGeoParquet().create([Path("tests/data-files/fiboa-example.json")], parquet) + out_file = tmp_folder / "collection.json" + CreateStacCollection().create_cli(parquet, out_file) + + columns = {c["name"]: c for c in load_file(out_file)["assets"]["data"]["table:columns"]} + assert columns["metrics:area"]["description"] == "Field area in square meters" + assert all("description" not in c for n, c in columns.items() if n not in DESCRIPTIONS) diff --git a/tests/test_publish.py b/tests/test_publish.py deleted file mode 100644 index cf960760..00000000 --- a/tests/test_publish.py +++ /dev/null @@ -1,30 +0,0 @@ -import responses - -from fiboa_cli.publish import Publish - - -class PublishTest(Publish): - def generate_pmtiles(self, target, file_name, parquet_file): - pass - - def upload_to_aws(self, target): - pass - - -@responses.activate -def test_publish(tmp_folder): - converter = "be_vlg" - base = "BE-VLG" - path = f"tests/data-files/convert/{converter}" - rsp1 = responses.Response( - method="GET", - url=f"https://raw.githubusercontent.com/fiboa/data-survey/refs/heads/main/data/{base}.md", - body=open(f"tests/data-files/publish/{base}-survey.md").read(), - ) - responses.add(rsp1) - PublishTest(converter).run( - variant="2023", target=tmp_folder, cache=path, generate_meta=True, yes=True - ) - files = [f.name for f in tmp_folder.iterdir() if f.is_file()] - for f in ("README.md", "LICENSE.txt", "be_vlg-2023.parquet"): - assert f in files, f"Missing file {f}"