Plazi taxon treatments no longer on GBIF?

Looking at the new portal and the move to using Catalogue of Life as the taxonomic backbone, I’ve realised that the Plazi treatments are no longer didplayed(!) For example, the treatment for Ceratogyrus attonitifer has a one sentence phrase and a couple of links, whereas before there was text extracted from the paper publishing the name, figures from that paper, etc. (you can see this on the Wayback Machine for https://www.gbif.org/species/10614692).

I stumbled across this change because there is a RDF version of GBIF on QLever that has triples linking images to taxa, and these images no seem to be on the GBIF site.

Am I correct in thinking that the treatments are not longer being displayed on GBIF? The display in ChecklistBank for the Plazi record captures some of what GBIf originally displayed, and obviously the treatment is also avauilable on Plazi, just curious as to the rationale for this change.

Hi @rdmpage yes the PLAZI-provided HTML files are no longer displayed and have been replaced by links. This is to encourage users to check the original publications as well as avoid more messy aggregations (like in this example: https://old.gbif.org/species/4822633/treatments).

The images are still displayed on the dataset-specific taxon pages on GBIF (https://www.gbif.org/dataset/c8635406-888f-49ed-b0a8-b8463677eae5/taxon/426E06A02C11DFE26260EE205AD40568.taxon) and checklistbank (https://www.checklistbank.org/dataset/6731/taxon/426E06A02C11DFE26260EE205AD40568.taxon) as well as aggregated for the taxon on checklisbank (ChecklistBank).

1 Like

Thanks for the explanation @mgrosjean, I agree that for some taxa such as T. rex things can get messy. Perhaps what we need is a way to summarise taxonomic information concisely so that users aren’t swamped with detail, but still get a useful overview. I suspect GBIF might think this is out of scope for the moment.

1 Like

I have finally exposed the html treatments on checklistbank: https://www.checklistbank.org/dataset/6731/taxon/426E06A02C11DFE26260EE205AD40568.taxon#treatment

This is for individual datasets only, so no aggregation for the same name. That is what GBIF does with the links out - which I think is the best option. It gives you the overview what exists and direct links to look them up

2 Likes

Thanks @markus, was just curious given that the old GBIF species integer ids were linked to taxon images in @Andrawaag GBIF RDF dump, and now both integers and integers have gone. Might pose some interesting challenges if that RDF gets updated to the new Catalogue of Life-based GBIF.

Please note that if you use the aggregated pages in ChecklistBanks Names Index - the names index identifiers are not stable and not designed to ever be. E.g. we just rebuilt that index and the link Marie posted about resolves now to a different name. This is by design, do not store those ints.

I am not following. What you are hinting at is our RDF rendering of the full occurrence dataset and one backbone as RDF, which opens IMO a treasure trove of possibilities,

This is a scaling effort of an approach I have been using for a while now, where I have been using Ontop to render GBIF subsets (particularly small ones) to integrate with other RDF stores. For example, to link out to UniProt to align with protein annotations for species surfacing in the occurrence data.
We recently were able to do this on the full-scale GBIF version, which really opens up a treasure trove of aligning GBIF data with other RDF resources (Wikidata, OpenStreetMap, UniProt, Pubchem, etc). Which identifiers are not persistent? A quick comparison with Wikidata, shows that most taxa identifiers linked from Wikidata to GBIF, are resolving. If these identifiers are not persistent, does that imply that the Wikidata links to GBIF are deprecated or soon going to be?

While I still not follow what the real issues are, I don’t expect this to be a difficult issue to fix. Remodelling the full RDF graph, takes approx 18 hours. So I guess it might “just” be a remodelling of the RDF transformation.The “old” id’s still resolve, but indeed under a new identifer format.
Where can I find information about what changed.

We are building from the parquet files from the open data cloud and I am not seeing anything in there. Am I missing something?

Either way, this is exciting :slight_smile:

@Andrawaag I’m still trying to make sense of this myself, but with the switch to CoL old-style GBIF URLs such as https://www.gbif.org/species/10614692 are now HTTP 302 redirects to /taxon URLs (in this case https://www.gbif.org/taxon/SN8G ). So in that sense there is little change in user experience (e.g., the old URLs in Wikidata will likely resolve).

I have not looked at the Parquet files, but the switch to CoL has also meant that the direct links between a species URL and descriptions and images have changed. For example, if I do a

DESCRIBE <https://www.gbif.org/species/10614692>

on https://qlever.dev/api/gbif I get:

{
  "@context": {
    "rdf": "http://www.w3.org/1999/02/22-rdf-syntax-ns#",
    "rdfs": "http://www.w3.org/2000/01/rdf-schema#",
    "owl": "http://www.w3.org/2002/07/owl#",
    "xsd": "http://www.w3.org/2001/XMLSchema#",
    "schema": "https://schema.org/",
    "foaf": "http://xmlns.com/foaf/0.1/",
    "dcterms": "http://purl.org/dc/terms/",
    "dc": "http://purl.org/dc/elements/1.1/",
    "skos": "http://www.w3.org/2004/02/skos/core#",
    "wd": "http://www.wikidata.org/entity/",
    "wdt": "http://www.wikidata.org/prop/direct/",
    "prov": "http://www.w3.org/ns/prov#",
    "geo": "http://www.opengis.net/ont/geosparql#",
    "dcat": "http://www.w3.org/ns/dcat#",
    "dwc": "http://rs.tdwg.org/dwc/terms/",
    "dwciri": "http://rs.tdwg.org/dwc/iri/",
    "gbif": "https://rs.gbif.org/terms/",
    "gbif1": "http://rs.gbif.org/terms/1.0/"
  },
  "@graph": [
    {
      "@id": "https://www.gbif.org/species/10614692",
      "dwc:class": "Arachnida",
      "dwc:family": "Theraphosidae",
      "dwc:genericName": "Ceratogyrus",
      "dwc:genus": "Ceratogyrus",
      "dwc:kingdom": "Animalia",
      "dwc:namePublishedIn": "Midgley, J. M., & Engelbrecht, I. (2019). New collection records for Theraphosidae (Araneae, Mygalomorphae) in Angola, with the description of a remarkable new species of Ceratogyrus. African Invertebrates 60(1): 1-, 13.",
      "dwc:order": "Araneae",
      "dwc:phylum": "Arthropoda",
      "dwc:scientificName": "Ceratogyrus attonitifer Engelbrecht, 2019",
      "dwc:scientificNameAuthorship": "Engelbrecht, 2019",
      "dwc:specificEpithet": "attonitifer",
      "dwc:taxonID": "10614692",
      "dwc:taxonRank": "SPECIES",
      "rdf:type": {
        "@id": "http://rs.tdwg.org/dwc/terms/Taxon"
      },
      "rdfs:label": "Ceratogyrus attonitifer",
      "skos:altLabel": {
        "@language": "en",
        "@value": "Ceratogyrus attonitifer"
      },
      "skos:broader": {
        "@id": "https://www.gbif.org/species/2153849"
      },
      "foaf:depiction": [
        {
          "@id": "https://binary.pensoft.net/fig/262540"
        },
        {
          "@id": "https://binary.pensoft.net/fig/262541"
        },
        {
          "@id": "https://binary.pensoft.net/fig/262542"
        }
      ],
      "gbif:hasDescription": [
        {
          "@id": "https://www.gbif.org/species/10614692/description/441571"
        },
        {
          "@id": "https://www.gbif.org/species/10614692/description/441572"
        },
        {
          "@id": "https://www.gbif.org/species/10614692/description/441573"
        },
        {
          "@id": "https://www.gbif.org/species/10614692/description/441574"
        }
      ],
      "gbif:rank": {
        "@id": "https://rs.gbif.org/terms/rank/species"
      },
      "gbif:sourceDataset": {
        "@id": "https://www.gbif.org/dataset/7ddf754f-d193-4cc9-b351-99906754a03b"
      },
      "gbif:taxonomicStatus": {
        "@id": "https://rs.gbif.org/terms/status/accepted"
      }
    }
  ]
}

As far as I can work out, the links between taxon and image are gone, and the description links do not resolve (404 on GBIF). So my expectation is that as the Parquet files are updated to include CoL, any RDF you generate will lose a lot of the the taxonomic information that is in the current RDF version.

P.S. I’m working on a SPARQL browser as part of another project, and when I pointed it at the https://qlever.dev/api/gbif one of the things I discovered is that it has URIs with namespaces https://rs.gbif.org/terms/ and https://rs.gbif.org/terms/1.0/, is that intentional? Seems likely that they are the same thing and could be merged.

Are there mappingfiles available to feed the 302 header?

Seems like a question for @markus

1 Like

@Markus_B Thank you, this is really helpful!. It raizes two questions though.

  1. How should Wikidata mappings be treated? Should we do a major curation effort by deleting all the mappings that currently point to GBIF from Wikidata items?
  2. As far as I have seen this update has not been updated yet on the parquet files being shared on the AWS open data cloud.

It is a bit unfortunate that our efforts in aligning the GBIF with the linked-open-data cloud. The RDF @rdmpage mentions is completely based on the parquet files being shared.

I guess that with this mapping file we should be able to update this easily, but if this change will soon be reflected in the parquet files we might simply wait.
For now, most if not all, SPARQL queries still run as expected.

@markus When looking at the mapping file shared here, I noticed that there are quite some gbif:IDs that lack a col:ID, even when the status is accepted.

e.g.

gbif:ID col:ID

0       

1       N

2       

3       

4       C

5       F

6       P

7       Z

8       

9       9J9G3

11      

12      KT6SY

13      9JHQ8

14      L2QNW

17      9XBJJ

18      4M

19      5P

22      B8V3M

25      

26      

28      

29      

30      

31      

32      5K

How should we treat those species ID’s that lack a col:ID?

@Andrawaag I’d argue against deleting any GBIF ids in Wikidata! Obsolete identifiers are still valuable. I’d recommend deprecating them, not deleting them. That would let people with links to old ids still have a bridge to Wikidata. Alternatively, one could imagine keeping the GBIF integer ids as a legacy identifier, and the CoL identifier could then serve as a key to CoL and GBIF.

Depends what you see as GBIF keys, but I would also think that the keys for the GBIF-ID property P846 should not change. That backbone taxonomy is still in existence and use, even by GBIF in its v1 API. The new keys are not really GBIF keys. We use the keys from Catalogue of Life and I don’t see a point in duplicating every single identifier with the new ‎GBIF taxon ID - Wikidata . The existing Catalogue of Life ID - Wikidata should be sufficient.

Btw, I am curating a list of taxonomy relevant identifier “scopes” with wikidata properties where known: https://api.checklistbank.org/vocab/identifier-scope

Regarding missing COL-GBIF mappings: we clearly miss some mappings like gbif:2 (Archaea) being col:CRLT8. I will try to update those, but cannot promise to get that done soon.

But there must be quite a few names that indeed do not have any mapping. Not sure how to handle that on your side, but there won’t be a 100% coverage for various reasons.

In case it helps @Andrawaag I’ve submitted a SQL download to create a file with the occurrence ID, taxon key and accepted taxon key (all COL identifiers). It’s in the queue for processing and will be tab delimited:

https://www.gbif.org/occurrence/download/0033503-260623161305970