Enable Geoconnex integration
Learn how to sync your CKAN instance with the Geoconnex water data knowledge graph.
To allow for minting persistent identifiers for your CKAN datasets and properly indexing your CKAN dataset based on its geospatial metadata (e.g. to identify the Geoconnex reference features a CKAN dataset is about or related to), you'll need to enable syncing Geoconnex with your CKAN instance.
We enable this CKAN --> Geoconnex sync using github.com/dathere/ckan_geoconnex_bulk_runner which is a software system built with Rust.
Enabling sync with GitHub
To enable the sync, you'll need to have a GitHub account and follow the steps below.
Fork the internetofwater/geoconnex.us repository and clone the fork to your local device
Fork the github.com/internetofwater/geoconnex.us repository and then clone your fork to your local device.
git clone https://github.com/YOUR_GITHUB_USERNAME/geoconnex.us.git
cd geoconnex.usAdd a namespace for your CKAN instance in namespaces/bulk/ckan
Add a namespace directory to the namespaces/bulk/ckan directory based on your CKAN instance's name. For example for the New Mexico Water Data Catalog, we'd add a directory with no spaces named New_Mexico_Water_Data_Catalog.
mkdir namespaces/bulk/ckan/New_Mexico_Water_Data_CatalogAdd a metadata.json file
Add a metadata.json file and modify the entries described in the table below. We use the New Mexico Water Data Catalog as an example.
cd namespaces/bulk/ckan/New_Mexico_Water_Data_Catalog
touch metadata.jsonDifferent semantics for 'datasets'
metadata.json example, a "dataset" is not the same as a CKAN dataset. In this JSON example a "dataset" represents a collection of all of your web resources that are being synced with Geoconnex.{
"contact_email": "support@dathere.com",
"dataset_description": "New Mexico Water Data Catalog web resources",
"dataset_has_html_landing_pages": true,
"add_associated_mainstems": false,
"source_code_link": "https://github.com/dathere/ckan_geoconnex_bulk_runner",
"dataset_documentation_link": "https://catalog.newmexicowaterdata.org",
"bulk_container_image": "ghcr.io/dathere/ckan_geoconnex_bulk_runner:new_mexico_water_data_catalog"
}| Key | Value |
|---|---|
contact_email | The email that can be reached out to in case there are issues when attempting to sync Geoconnex with your CKAN instance's data. |
dataset_description | A description about your CKAN instance. |
dataset_documentation_link | Your CKAN instance URL. |
bulk_container_image | Keep ghcr.io/dathere/ckan_geoconnex_bulk_runner: and then add the lowercase version of your namespace such as new_mexico_water_data_catalog. |
You should keep the example's values for dataset_has_html_landing_pages, add_associated_mainstems, and source_code_link.
Add a web_resources.csv file
You'll also need a CSV file that maps the newly minted Geoconnex persistent identifier URLs to your CKAN web resources URLs.
touch web_resources.csvFor now there is support for dataset-oriented web resources, so you can modify the following example:
id,target
https://geoconnex.us//ckan/New_Mexico_Water_Data_Catalog/([-a-zA-Z0-9_]+).*$,https://catalog.newmexicowaterdata.org/dataset/$1- For the
idvalue replaceNew_Mexico_Water_Data_Catalogwith your namespace. - For the
targetvalue replacehttps://catalog.newmexicowaterdata.orgwith your CKAN instance site URL.
Commit changes and submit your pull request
By now you should have the following:
- New namespace directory in
namespaces/bulk/ckanbased on your CKAN instance name metadata.jsonfile describing your CKAN instance as needed- CSV file (e.g.
web_resources.csv) to map Geoconnex PIDs to your CKAN web resource URLs
Now you'll need to stage, commit, and push your changes to your fork.
git add -A
git commit -m "feat: add NMWDC CKAN bulk loader namespace"
git pushGo to your fork on GitHub and create a new pull request. The geoconnex.us administrators should review your pull request and merge it if everything looks correct.