ckanext-gztr

Enable Geoconnex integration

Learn how to sync your CKAN instance with the Geoconnex water data knowledge graph.

To allow for minting persistent identifiers for your CKAN datasets and properly indexing your CKAN dataset based on its geospatial metadata (e.g. to identify the Geoconnex reference features a CKAN dataset is about or related to), you'll need to enable syncing Geoconnex with your CKAN instance.

We enable this CKAN --> Geoconnex sync using github.com/dathere/ckan_geoconnex_bulk_runner which is a software system built with Rust.

CKAN to Geoconnex data flow diagram

Enabling sync with GitHub

To enable the sync, you'll need to have a GitHub account and follow the steps below.

Fork the internetofwater/geoconnex.us repository and clone the fork to your local device

Fork the github.com/internetofwater/geoconnex.us repository and then clone your fork to your local device.

git clone https://github.com/YOUR_GITHUB_USERNAME/geoconnex.us.git
cd geoconnex.us

Add a namespace for your CKAN instance in namespaces/bulk/ckan

Add a namespace directory to the namespaces/bulk/ckan directory based on your CKAN instance's name. For example for the New Mexico Water Data Catalog, we'd add a directory with no spaces named New_Mexico_Water_Data_Catalog.

mkdir namespaces/bulk/ckan/New_Mexico_Water_Data_Catalog

Add a metadata.json file

Add a metadata.json file and modify the entries described in the table below. We use the New Mexico Water Data Catalog as an example.

cd namespaces/bulk/ckan/New_Mexico_Water_Data_Catalog
touch metadata.json

Different semantics for 'datasets'

In the following metadata.json example, a "dataset" is not the same as a CKAN dataset. In this JSON example a "dataset" represents a collection of all of your web resources that are being synced with Geoconnex.
namespaces/bulk/ckan/New_Mexico_Water_Data_Catalog/metadata.json
{
    "contact_email": "support@dathere.com",
    "dataset_description": "New Mexico Water Data Catalog web resources",
    "dataset_has_html_landing_pages": true,
    "add_associated_mainstems": false,
    "source_code_link": "https://github.com/dathere/ckan_geoconnex_bulk_runner",
    "dataset_documentation_link": "https://catalog.newmexicowaterdata.org",
    "bulk_container_image": "ghcr.io/dathere/ckan_geoconnex_bulk_runner:new_mexico_water_data_catalog"
}
KeyValue
contact_emailThe email that can be reached out to in case there are issues when attempting to sync Geoconnex with your CKAN instance's data.
dataset_descriptionA description about your CKAN instance.
dataset_documentation_linkYour CKAN instance URL.
bulk_container_imageKeep ghcr.io/dathere/ckan_geoconnex_bulk_runner: and then add the lowercase version of your namespace such as new_mexico_water_data_catalog.

You should keep the example's values for dataset_has_html_landing_pages, add_associated_mainstems, and source_code_link.

Add a web_resources.csv file

You'll also need a CSV file that maps the newly minted Geoconnex persistent identifier URLs to your CKAN web resources URLs.

touch web_resources.csv

For now there is support for dataset-oriented web resources, so you can modify the following example:

namespaces/bulk/ckan/New_Mexico_Water_Data_Catalog/web_resources.csv
id,target
https://geoconnex.us//ckan/New_Mexico_Water_Data_Catalog/([-a-zA-Z0-9_]+).*$,https://catalog.newmexicowaterdata.org/dataset/$1
  • For the id value replace New_Mexico_Water_Data_Catalog with your namespace.
  • For the target value replace https://catalog.newmexicowaterdata.org with your CKAN instance site URL.

Commit changes and submit your pull request

By now you should have the following:

  • New namespace directory in namespaces/bulk/ckan based on your CKAN instance name
  • metadata.json file describing your CKAN instance as needed
  • CSV file (e.g. web_resources.csv) to map Geoconnex PIDs to your CKAN web resource URLs

Now you'll need to stage, commit, and push your changes to your fork.

git add -A
git commit -m "feat: add NMWDC CKAN bulk loader namespace"
git push

Go to your fork on GitHub and create a new pull request. The geoconnex.us administrators should review your pull request and merge it if everything looks correct.

On this page