Georeference.it - a new open platform for crowdsourced georeferencing of GBIF records

Hi all,

I’d like to share a platform I’ve had in my head for five years and finally built: georeference.it: an open-source web tool for collaboratively georeferencing GBIF occurrence records.

The platform works around three contributions:

  • Filling gaps - volunteers georeference locality groups with no coordinates following Georeferencing Best Practices; where GBIF coordinates already exist within the group, the system generates an automatic candidate suggestion as a starting point
  • Community validation - a consensus model where suggestions are validated through peer agreement, so no single user’s opinion determines the outcome
  • Error detection - automatic flagging of spatially inconsistent records for the same location, presenting them as validation tasks for the community

A browser extension for Chrome and Firefox is also available, surfacing georeferencing status directly on gbif.org occurrence pages.

Platform: https://georeference.it
Source code: GitHub - joaquimsantos1978/georeference-it · GitHub

Any feedback, bug reports or contributions are very welcome!
*
Disclaimer: I should mention that I’m submitting this to the 2026 GBIF Ebbe Nielsen Challenge, but I think the project is needed regardless, so I intend to maintain it for as long as possible, and hopefully build a community around it.*

4 Likes

@joaquimsantos , this looks like a great project but I’m having trouble understanding who will benefit from it. On your GitHub page you state that

Only about 55% of natural history specimen records accessible via biodiversity aggregators such as GBIF are georeferenced

How did you calculate that? What base (denominator) did you use - does “natural history specimen records” refer only to specimens in museums and herbaria? It’s certainly not true for the ca 90% of GBIF-mediated records that come from human observations, overwhelmingly from citizen-science platforms that require locations.

Assuming your project targets missing or wrong georeferencing of museum and herbarium specimens, how will the curators and data managers at those institutions access community-sourced corrections? Do you expect them to do so any more frequently than they currently do with GBIF spatial-issue flags, and do you expect them to do corrections or add annotations in their CMSes?

How do georeference.it annotations link to the GBIF-mediated datasets? If they don’t, is georeference.it sitting to one side of CMS and the GBIF platform, requiring end users to visit georeference.it to see if the GBIF-mediated records they’re working with have been annotated there? If so, how do you imagine end users will find those annotations when working with large datasets?

Hi @datafixer , thanks for the feedback.

The percentages refer to natural history specimens (museums and herbaria), and are a citation from this paper: Marcer A, Haston E, Groom Q, et al. Quality issues in georeferencing: From physical collections to digital data repositories for ecological research. Divers Distrib.2021;27:564–567. https://doi.org/10.1111/ddi.13208

The platform provides an API that can be queried by dataset identifier, so yes, the idea is that they can programmatically query the API and ingest the values into their CMSs. In a more intrusive approach, the platform could even send to the dataset publishers the list of records with the approved values - I’m not planning to do so, but when I was managing data from the herbarium of Coimbra I received a list of things to correct from Darwin Core checker: Introduction and I loved it!

The platform has also an API that users can use to see individual records georeferencing. For bulk analysis they could use the API quering by occurrence IDs. I recognise this is not very easy for most users though.

Ideally, if the quality of community georeferencing gets recognised, GBIF could use georeference.it as a preferred source for georeferencing and attach the values to the occurrence records, tagging them as suggestions/corrections.

1 Like

@joaquimsantos Nice! I have a less sophisticated tool that simply treats GBIF as a geocoding tool GBIF geocoder. You paste in a locality, e.g. “Cambodia: Ratanakiri Province, Virachey National Park” and it finds localities for existing records that look like that GBIF geocoder.

I naively pasted that same string into https://georeference.it and the results were unexpected (it took a while to load then I got results from France(!). Does the tool accept raw strings or do I need to have a more structured query?

Hi @rdmpage , thanks for your message.

I knew about GBIF Geocoder :slight_smile: . I find the idea very interesting and looked into it to figure out if it would be a good way to infer automatic suggestions for ungeoreferenced specimens. I chose not to because there are some situations where a biased result may occur - that would be a very interesting discussion :slight_smile:

About your question, the Focus Area accepts text input and does a Fulltext query to the locations table. What probably happened with your query is that it didn’t find any matches, so returned the first location in line. The idea is that people select a broad area they would like to work on during the session, like Ratanakiri, and the platform will show matching locations to georeference - I realise queries are taking a long time; I have indexes for the table, but need to implement some optimisation for sure.

Acually I see now that a condition I have for not logged users is delaying the results. Apparently all ungeoreferenced occurrences in that area have already pending suggestions created by the system based on existing georreferenced occurrences, so it keeps searching but finds none. For logged in users the loading was quick. I just removed the condition, so you will now see results for Ratanakiri or other places.

I’ll try to improve the searches for Focus area anyway.

That sounds like something you could also easily do in DiSSCo for all relevant specimens, basically by duplicating the MAS (Machine Annotation Service) that already exists to georeference with the Geopick API, and make it using the DiSSCo batch functionality to repeat for the same string in other specimens. May need to be a bit more sophisticated though in a production environment to cater for red list species that need an obfuscated georeference.

Don’t “can” and “could” get to the heart of the matter? Collections already have GBIF flagging their datasets with annotations for the more obvious problems, like swapped lat/lons, lats and lons with wrong signs, country-coordinate mismatch. Collections “can” take note of these flags and correct their CMSes, but overwhelmingly they don’t.

georeference.it would add even more flags (and propose corrections for them), like stateProvince/county/municipality/placename-mismatch. If collections are largely unwilling to do the GBIF correction work, why do you think they might welcome doing the additional georeference.it work?

Yes, it’s great that end-users would be able to check their datasets of interest on the georeference.it platform (if “the community” had already done something with those records), but please note that occurrenceIDs in GBIF-mediated datasets are often not globally unique, so an end-user querying by occID would first need to be sure of a record match.

Ideally, if the quality of community georeferencing gets recognised, GBIF could use georeference.it as a preferred source for georeferencing and attach the values to the occurrence records, tagging them as suggestions/corrections.

GBIF already has an annotation system, where user-provided corrections and remarks can be added to records after GBIF processing. I haven’t seen a review of its effectiveness - perhaps someone at GBIF could comment?

I realise I sound pessimistic about this, but mechanisms for effective annotations of shared biological records have been discussed at least since 2008, when Donald Hobern argued for one in Atlas of Living Australia. The mechanisms have become increasingly sophisticated, and your community-based platform for georeferencing is excellent.

But the key problem in annotating collection records isn’t a technical one that can be solved with a sophisticated annotation system. The key is a data management by humans problem at the contributing institutions (time, skills, resources…).

At the user end, my reading of the SDM literature suggests that consumers of records typically use GBIF spatial-data flags to discard (“clean” [sic]) records, because that saves work. I’m not sure how georeference.it would reduce the work required.

@joaquimsantos, thanks for the comment about the Pensoft data check, and I’m glad you found it useful.

In the past couple of months I’ve done two similar museum collection audits, of 908119 and 570361 records, respectively, outside the Pensoft data-paper frame. Both datasets are shared with GBIF. One of the data managers said they were “making improvements where possible”, the other “doesn’t have time to action what [I found]”.

In both cases I found far more errors and inconsistencies than GBIF did with its automated checks, which makes my audits add to collections workload in the same way georeference.it might do.

Like you, GBIF is optimistic that collections will pay attention to its advice on data quality:

Publishers thus play an essential role not simply in sharing datasets, but also in managing their quality, completeness and usefulness as well as ensuring their integration and value within GBIF’s global knowledge base. (https://www.gbif.org/data-quality-requirements)

In previous posts on this forum I’ve suggested why this optimism is unrealistic, and noted that non-collections records are not only vastly more abundant through GBIF than collections records (ca 90%), but are also the ones most likely to be used for conservation planning and other real-world applications for datasets. These non-collections records have vanishingly few georeferencing issues. The temptation will be there for end-users to simply ignore collections records with spatial issues, rather than use georeference.it to fix them.

Hi @datafixer ,

I understand why you are skeptical about the real advantage of such systems, and you definitely got a point.

However I would argue that even if the systems are underused in relation to their full potential, the few uses that they make possible are worthy.

Regarding collection records, they might be useless for many situations when you have much more human observations. This is the case specially for the so-called Western World, but not so much for many places in Africa for example, where most of the biodiversity records are the specimens in museums and herbaria. Millions of historical records have been digitised in the last years, and many millions are still to be (hopefully soon). Of all the tasks that digitisation implies, I would say that georeferencing is the one institutions will do in the end, if there is time and resources, mostly due to the lack of tools in CMSs to do this operation in bulks. This is the gap that georeference.it fills, which is even better if I’m able to bring the contributions of volunteers. I have run other crowdsourcing projects and volunteers contribution is amazing (and very accurate, I must say).

About the collections ingesting or correcting data, I think what georeference.it has to offer is a bit different than the cases you mentioned. The quality checks flagged in GBIF are hardly automatically ingested in a completely autonomous pipeline - I mean, you can have a specimen flagged with a country-coordinate mismatch, but all you can do is delete the coordinates; to have the right coordinates you would need extra work. The same for the other flags, you see the flag tag, not the value you should fill in. With georeference.it they could just ingest the values coming from there (assuming they will be trustworthy - and I don’t have reasons to believe they won’t). Also there are cases flagged that GBIF doesn’t do, like the same location description georeferenced to two very different points - there are cases when all might be correct, but certainly there are many cases when at least one of them should be corrected.

That being said, I believe something good can come out of it. Let me bring the case of Bionomia, developed and maintained by @dshorthouse. I started collaborating there, mostly because I had an interest in using the data to enrich the collection dataset I was managing - and I did use it! But potentially helped many other collections or researches that can use that data as well. Maybe the community benefiting from Bionomia is not so wide as it could be, but I’m sure it has enhanced other projects. Let me shamelessly mention this side project I’m maintaining that is only possible thanks to Bionomia: https://www.expeditia.info/ - it gets research expeditions from Wikidata, and for each participant in an expedition gets the list of specimens collected by them between the dates of the expedition. Maybe one day all GBIF specimen records will have person identifiers for the collectors and allow such filtering, but while not, we trust Bionomia to do that for us. This is also an example of how collection records can be used for purposes different from those of human observation records.

@joaquimsantos , I understand and appreciate what you are trying to provide. My skepticism is unimportant if (as you say) there will be valuable end uses somewhere at some time. The excellent Bionomia likewise serves some purposes, and serves them very well.

But with any new project, I like to ask “What problem does this solve, or what obstacle does it overcome?” I have no doubt at all that there are lots of volunteers ready and able to do georeferencing. That’s the supply side of your project.

But do you have a clear picture of the demand side? Are there lots of collections that aren’t doing georeferencing and correcting georeferencing errors simply because there isn’t a service like georeference.it that can do at least some of the job for them? Or is the situation more that there are lots of collections that aren’t doing georeferencing and correcting georeferencing errors because that’s so far down their list of priorities that it just won’t happen in the foreseeable future, and could only happen if the collections got grants to employ someone to do that particular job? Who do you see as your potential customers?

I’m also skeptical that a collection could simply “ingest” data from georeference.it into whatever particular CMS the collection is using, because that’s an IT job beyond the abilities of most curators. Further, one large collection I’ve dealt with told me that anything not on the CMS menus, like global edits, would have to be done by the (commercial) CMS provider, at a cost in addition to annual maintenance and upgrade charges.

Your “expeditia.info” project is also admirable. I’ve often wished that more expeditions published their itineraries - where they collected and when, not just who collected. For more on this, please see An Australian collector's authority file, 1973–2020

Thanks for your thoughts on this @datafixer.

Thanks also for the paper about your collector’s authority file. It resonates with my experience as collection manager, on the effort to don’t loose the information coming from the collecting events - and I must say the workflows for integrating specimens on a collection are very prone to loose pieces of information on the way.

You are right about collections ingesting data (on a side note, I always thought that existing CMSs, despite being structured and flexible, are still way behind what they should be). That’s why I envision a future where GBIF completes/corrects the information for them, to be centrally available for researches and people who could use it. Mechanisms such as georeference.it, Bionomia, DiSSCo, etc. could be the trusted source. I’m being totally naive, I know :slight_smile:.

Hi @waddink, nice to hear from you.

I’m not sure if your comment was addressed to me or to @rdmpage.

About your suggestion, I’m totally in favor of such automatism, specially the ones that are backed up by existing records. I know GeoPick (Arnald kind-of developed it on my request) and I doubt it could come up with something accurate for strings like “Bank Nyong River, near the new bridge, about 65 Km SSW of Eséka”, I think we still need humans in the loop for that kind of things. But yes, once you have a value, replicate it on every specimen having the same string.

@joaquimsantos I wanted to insert a screenshot or two, but seems like GBIF doesn’t let you do that :face_exhaling:

In the explore tab the list of country codes includes numbers and stray punctuation, maybe this is what is in GBIF records? I’d suggest having a clean list, and show country names alongside each code.

I did a search for “Ratanakiri” in KH (which is much faster than “Ratanakiri Province”), when I click on an record to georeference it the map shows Spain(!). Maybe it would be better to have the map automatically pan across to the country of interest, so users don’t have to do that themselves? Maybe even do the same sort of search as GBIF Geocoder just to give user a starting point (“we think what you seek is somewhere near here”).

P.S. good luck in the challenge!

Hi @rdmpage

I guess that the importing process from GBIF data broke at some point. I’ll do a new import soon and hopefully those will be fixed.

Thanks for the feedback and for the tip about positioning the map. Sorry about the map being centered on Iberia, that’s the default view, because I started the prototype with a dataset from Portugal - I will correct that ASAP.

The proxy for what you are suggesting is the searchbox bellow the locality, wich uses Nominatim to geocode input. This would be the locality to georeference, rather than the focus area. But you are right, this part of the UX/UI needs some improovement to be more user freindly and maybe also more informative on what is happening behind the scenes. I’ll give it some serious thought.

Thanks!

Wow! This is awesome! You have clearly put your heart & soul into this. And, the fact that others here have responded in such useful and constructive ways by making you contemplate other ideas or approaches (or to think about motivations) is a testament to your drive and vision. It’s in this vein that I too have a few things to offer.

How might you categorize the signal strength on localities that await georeferencing? And, given that you might be able to do that, how do you get it to scale? Generally, any website has roughly three seconds to capture users’ attention & to absorb what’s being asked of them.

I see ORCID as a login. This is good. Is there a signal in ORCID profiles that might guide new users to occurrence records that are in need of georeferencing & make the approach more personable?

If Datasets is one entry point in the top menu, how might you bubble-up categories of occurrence records? Something like, “Try these ones, they might be the least challenging to georeference.” vs “We think these one’s will be really difficult. Try them after you gain confidence in how this works.” Is there an algorithm that could be turned into a stand-alone library of code or as a publishable unit?

The core, least interesting aspect of doing things like the above is search. But, it is the most critical part. I would investigate how far you can push Elasticsearch to make your site & its filtering as fast as possible. This will increase hosting costs (and there will be some compromises), but it will make the site stickier and keep people engaged. It might force you to concentrate on whatever are the most information-rich signals (and least costly to determine or index) so that you can guide users in ways the push the percent georeferenced up to 100%.

Onward & good luck with your Ebbe Nielsen Challenge submission.

1 Like

Hi @dshorthouse. Thank you very much for your kind words and your suggestions. You know Bionomia has been such an inspiration for me to move forward with this.

All the points you rise are very important, and ideally implemented ASAP, but I fear some might be to complex without extra tools.

Categorising the localities to georeference by difficulty is in principle possible with a set of rules, at least tagging the ones that include things like “[N/NW/NE/S/SW/etc] of” as difficult, and other rules alike. But there will be for sure many tagged as simple, say, because they are a single word, but this happens to be the name of a little village in Angola that doesn’t exist since 1860 and nobody know where it was anymore. Definitely worth trying it. Georeference.it has already a ranking system for users, based on their performance on the platform: all start as beginners, and as their suggestions get approved, they are moved to the next level. If we rank the localities by difficulty, we can show only the easy ones to beginners and then progressively difficult as they move to upper levels (I had implemented this on a transcription platform I developed - easy fields, as collection dates for beginners, determinations only for experts, etc.).

The ORCID profile as basis for suggestions is also nice… if someone is ornithologist, present them localities form bird records… I’ll look it up!

For the Elasticsearch it would probably need a big change and definitely a bigger cost. I’m running it on a low cost server, and it is already a stretch to my budget… maybe if I win the challenge :smiley:.

1 Like