What is passport data?
If you're preparing to publish your collection on Genesys for the first time, we'll start by setting the foundations. In this section, we explain what passport data is, how it differs from other data types you may hold, and the practical questions you should consider before deciding what to publish, including how to handle collector identities and sensitive locations.
A definition that fits in one sentence
Passport data is the basic descriptive information that identifies an accession and records where it came from. It's the "who, what, where, when" of each sample in your genebank: what species it is, your accession number, where it was collected (or where it came from if it wasn't collected), when it was acquired, who collected or bred it, and what status it has today.
Passport data makes an accession findable and citable, whether for you, researchers, other genebanks, or automated systems. Without it, an accession is just material on a shelf. With it, the accession becomes a record that can be searched, compared, linked, and reused.
The multi-crop passport descriptor standard
Passport data is not a free-form description. It follows a defined international standard: the FAO/Bioversity Multi-Crop Passport Descriptors (MCPD), currently at version 2.1, December 2015.
The MCPD was jointly developed by Bioversity International (formerly IPGRI) and FAO. It defines a list of descriptors (fields) that genebanks worldwide use to describe their accessions consistently. Genesys, WIEWS, EURISCO, and many information systems all use MCPD as their common reference.
The MCPD lists 28 main descriptors plus the Persistent Unique Identifier (PUID) along with several sub-descriptors. They include:
- Identifiers: institute code (
INSTCODE), accession number (ACCENUMB), collecting number (COLLNUMB), persistent unique identifier (PUID) - Taxonomy: genus, species, subtaxa, common crop name, accession name
- Origin and collection: country of origin, location of collecting site, latitude, longitude, elevation, collecting date, collecting source
- People and institutes: collecting institute, breeding institute, donor institute
- Status: biological status of the accession (wild, landrace, breeding material, cultivar, etc.), Multilateral System (MLS) status under the International Treaty on Plant Genetic Resources for Food and Agriculture, safety duplication location
- Remarks: a free-text field for anything that doesn't fit the structured descriptors
The section on minimum data requirements explains which of these descriptors Genesys treats as mandatory, and the preparing MCPD-compliant data section walks you through how to format each one correctly.
How passport data differs from other accession data
Your genebank likely holds, or could hold, several types of data about each accession. Only the first type is "passport data" in the MCPD sense. The other data types are covered in beyond passport data.
| Data type | What it describes |
|---|---|
| Passport data | Identity, origin and basic status of the accession |
| Characterization data | Observable, heritable traits, morphological, agronomic or quality descriptors recorded under standardized conditions (e.g. seed color, plant height, days to flowering) |
| Evaluation data | Performance traits recorded under specific environmental or stress conditions (e.g. yield under drought, resistance to a specific disease isolate) |
| Genotypic data | Molecular markers, sequence data, SNP genotypes, typically generated by laboratory analysis |
| Image data | Photographs of plants, seeds, fruits or flowers associated with accessions |
Passport data is the foundation layer. You can publish passport data alone and have a useful presence on Genesys. You can't meaningfully publish characterization, evaluation, or genotypic data without passport data, because those datasets reference accessions by their passport identifiers.
This is why Genesys advises new providers to start with passport data and add other data types incrementally. Data curation and validation elaborates on this "start small, improve over time" principle.
What passport data is for, in practice
Here are three concrete examples of why passport data matters and why getting it right is worth the effort.
Discovery. For example, a breeder looking for sorghum landraces from West Africa filters Genesys by genus, biological status (landrace), and country of origin. If your accessions have all three fields populated correctly, they appear in the search. If the genus is missing or the country of origin is wrong, your collection is invisible to this search.
Cross-collection comparison. A curator at one institute wants to know whether a duplicate of their wheat accession IT-12345 exists in another collection. Genesys' Similarity Search uses passport data (collecting number, accession name, origin, and collecting site) to find candidate duplicates. The more passport fields you populate, the better the matching works.
Citation and attribution. A published paper cites the accessions used in a study by their MCPD identifiers (institute code + accession number or by GLIS DOI). The published record is only as informative as the underlying passport data. If a researcher cannot find your collecting site or species, they cannot interpret the result.
In short, passport data is the bridge between your genebank's internal records and the global community.
Handling personally identifiable and sensitive information
Passport data inherently contains information about people and places. Some of it is sensitive, some of it is not. You, the genebank, are responsible for deciding what is appropriate to publish, not Genesys.
We recommend thinking about these three categories explicitly before uploading. This guidance reflects standard practice for biodiversity data publication and the structure of the MCPD itself; the Crop Trust's specific position on these questions is documented separately and should be cross-referenced with your institutional data-protection policy.
Collector names
MCPD COLLCODE records the institute that collected the sample. Descriptors COLLNAME and COLLINSTADDRESS record the institute's name and address, used only when a FAO WIEWS code is not available. Individual collector names are not part of the MCPD standard.
However, individual names sometimes appear in:
- REMARKS, where notes such as "collected by Dr. X and team" may have been written historically
- Internal records that you may be tempted to migrate into
REMARKSduring data preparation
Recommended approach:
- Use the institutional record (
INSTCODEorCOLLNAME) rather than individual names wherever possible. "ICRISAT collecting mission, 1984" is informative and uncontroversial. "Dr. Jane Smith, ICRISAT, 1984" carries personally identifiable information (PII) that may be subject to data-protection law. - If a collector's contribution is significant enough to warrant acknowledgement, consider whether published acknowledgement is appropriate and obtain their consent.
- For collectors who have passed away, contributions are conventionally acknowledged in published form, but the responsibility for the decision sits with your institute.
When in doubt, default to institutional attribution and omit individual names from the published record.
Sensitive collecting locations
MCPD COLLSITE, DECLATITUDE, DECLONGITUDE and others record the location and coordinates of collection. Coordinates can be highly precise (decimal degrees to multiple decimal places).
Precise coordinates are sometimes inappropriate to publish because they could reveal:
- The location of an endangered population that could be subject to over-collection or vandalism
- Sites on land controlled by indigenous communities or protected areas where access is restricted
- Specific farms or fields where farmers may not have consented to public disclosure
- Sites within zones of armed conflict or contested territory
The MCPD anticipates this. Two of its descriptors give you the tools to publish useful location data without publishing precise coordinates:
- COORDUNCERT (Coordinate uncertainty, in meters), lets you record that a coordinate is approximate. You can deliberately publish less precise coordinates and use this field to make the imprecision explicit.
- GEOREFMETH (Georeferencing method), lets you record how the coordinate was derived (e.g. "generalized to nearest 0.1 degree").
Recommended approach for generalizing sensitive coordinates:
- Truncate to one decimal place (approx. ±11 km of precision) if you want to indicate the general region while protecting the specific site.
- Truncate to zero decimal places (approx. ±111 km) if even regional precision is sensitive.
- Withhold coordinates entirely and record country of origin only (ORIGCTY) if the location is too sensitive for any precision.
- Record the uncertainty in COORDUNCERT so users understand the data has been generalized, and do not pretend a generalized coordinate is precise.
Do not randomly shift coordinates to obscure the true location while keeping the same decimal precision. This corrupts any downstream spatial analysis a researcher might do and is misleading rather than informative.
Confidential, embargoed or third-party data
The Data Provider Agreement (DPA, Key Characteristics of the DPA) commits you to publishing only data that is not subject to confidentiality restrictions. This includes:
- Data collected under research projects with publication embargoes, wait until the embargo expires before uploading
- Data shared with you by third parties under restrictive terms, confirm the original holder permits republication or omit
- Data covered by Access and Benefit-Sharing (ABS) agreements with specific source countries that limit republication
You are responsible for the decision and, under the DPA, for any liability that arises from a breach of confidentiality.
Before publishing a record, ask three questions:
- Would publishing this harm anyone? A collector's identity, a community's land, an endangered population?
- Do I have the right to publish this? Was it collected under terms that allow open publication?
- Is what I'm publishing accurate? Or am I publishing imprecise data with precise-looking values?
If all three pass, publish. If any fail, generalize, omit or wait.