Skip to main content

Passport Data Completeness Index (PDCI)

The PDCI is a single numerical score from 0 to 10 that summarizes how complete an accession's passport data is. It was originally proposed by van Hintum, Menting and van Strien in their 2011 paper Quality indicators for passport data in ex situ genebanks and has been adopted by Genesys (and EURISCO and other systems) as a standard completeness metric.

What the PDCI measures​

PDCI measures completeness, whether information is present, not whether it is correct. A high PDCI score means an accession has a lot of fields populated; a low score means it has very few. The Validator Tool can confirm a record is mechanically correct within its scope; PDCI tells you whether it is sufficiently complete. A score of 10 means an accession has all the descriptors populated that are relevant to its type.

What "completeness" means depends on the accession type​

The PDCI does not weight every descriptor equally. The weight of each descriptor depends on the biological status of the accession (SAMPSTAT). This reflects a sensible principle: not every field is relevant to every accession type.

  • A wild accession should have a well-documented collecting site, coordinates, location description, collecting date, but it should not have a variety name (because wild material is not a cultivar). So COLLSITE, DECLATITUDE and DECLONGITUDE are weighted heavily for wild accessions, while ACCENAME is not.
  • A modern cultivar should have a variety name, breeding institute and ancestral information, but it would not have collecting coordinates (because it was bred, not collected from the wild). For cultivars, ACCENAME and BREDCODE matter; COLLSITE matters much less.
  • A landrace sits in between, both name and origin information are relevant.

This is why an accession can score the maximum PDCI of 10 regardless of whether it is wild, landrace, breeding material or cultivar: the index measures completeness relative to the kind of accession it is.

Where to see your PDCI scores in Genesys​

PDCI is calculated automatically for every accession in Genesys and is visible at two levels:

  • Per-accession, on the individual accession's Genesys record
  • Per-collection, aggregated across your genebank, useful for understanding where your collection stands overall

Using PDCI in practice​

For your genebank, PDCI is a diagnostic tool, not a target. A few practical ways to use it:

  • Identify records with the largest improvement potential. Typically, accessions with intermediate PDCI scores were the ones where curation effort produced the most improvement, records with very high scores are already nearly complete and records with very low scores often cannot be improved because the information was never captured or has been lost. Look first at the middle of your PDCI distribution for the best return on curation effort.
  • Spot patterns across your collection. If the wild accessions in your collection have a much lower average PDCI than your landraces, the gap is probably in collecting-site data (coordinates, dates, descriptions). That tells you where a historical archives review might pay off.
  • Track improvement over time. Note your collection's average PDCI today. Set a modest target, for example, +0.5 in 12 months, and use it as a measure of your curation work. Frequency of Data Updates section covers update cadence in more detail.
  • Don't game the metric. Adding placeholder values like "unknown" to inflate completeness is worse than leaving fields empty. The MCPD itself is explicit that missing fields should be left blank and placeholder strings would be flagged or stripped by validation processes.
note

PDCI is one signal among several. A collection with PDCI 7 but consistent factual errors is worse than one with PDCI 5 and accurate records. Use it alongside spot-checking and user feedback.