# Limit columns in csv from API (R)

**URL:** <https://discourse.gbif.org/t/limit-columns-in-csv-from-api-r/5659>\
**Category:** Data Use\
**Created:** [January 21, 2025, 2:38pm UTC](https://discourse.gbif.org/t/limit-columns-in-csv-from-api-r/5659 "2025-01-21T14:38:09Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![flopet](https://avatars.discourse-cdn.com/v4/letter/f/898d66/32.png) [@flopet](https://discourse.gbif.org/u/flopet)\
**Post date:** [January 21, 2025, 2:38pm UTC](https://discourse.gbif.org/t/limit-columns-in-csv-from-api-r/5659/1 "2025-01-21T14:38:09Z")

</div>

Hello !

I’m retrieving data from GBIF via the API in R and it’s working (thanks by the way =) )  
But the file I’m retrieving is too big from my taste (and the memory of my computer).  
Is it possible to limit the columns I’m receiving in my csv ?

My call looks like this :  
data ← occ\_download(pred\_in(“taxonKey”, gbif\_taxon\_keys), pred(“hasCoordinate”, TRUE), format = “SIMPLE\_CSV”,user=gbif\_user, pwd=gbif\_pwd, email=gbif\_mail)

If possible, I would like to only download the following columns “species, decimalLongitude, decimalLatitude, gbifID, countryCode”.  
(I can clean the csv myself, but it would be faster to not download it in the first place)

Thanking you in advance,

Florent.

---

<div class="post-metadata">

**Author:** ![mgrosjean](https://avatars.discourse-cdn.com/v4/letter/m/97f17d/32.png) [@mgrosjean](https://discourse.gbif.org/u/mgrosjean)\
**Post date:** [January 21, 2025, 3:49pm UTC](https://discourse.gbif.org/t/limit-columns-in-csv-from-api-r/5659/2 "2025-01-21T15:49:22Z")

</div>

Hi @flopet you could consider trying some SQL API download functions (which allows you to select specific columns). The documentation is available here: [API SQL Downloads :: Technical Documentation](https://techdocs.gbif.org/en/data-use/api-sql-downloads) and [GBIF SQL Downloads • rgbif](https://docs.ropensci.org/rgbif/articles/gbif_sql_downloads.html)

---

<div class="post-metadata">

**Author:** ![pieter](https://avatars.discourse-cdn.com/v4/letter/p/82dd89/32.png) [@pieter](https://discourse.gbif.org/u/pieter)\
**Post date:** [January 23, 2025, 8:51am UTC](https://discourse.gbif.org/t/limit-columns-in-csv-from-api-r/5659/3 "2025-01-23T08:51:15Z")

</div>

I agree with Marie that the SQL API is probably the easiest. You might also find the Parquet dump interesting.

> **[Using Apache Arrow and Parquet with GBIF-mediated occurrences](https://data-blog.gbif.org/post/apache-arrow-and-parquet/)**
>
> As written about in a previous blog post, GBIF now has database snapshots of occurrence records on AWS. This allows users to access large tables of GBIF-mediated occurrence records from Amazon s3 remote storage. This access is free of charge.
