Article

From population genomics to biomarker discovery: pangenomics for crop improvement

brown wheat field

Two studies on Ricinus communis and Eragrostis tef, presented at the SIGA Annual Congress, show how graph-based pangenomics connects genomic diversity, phenotype and functional annotation to support biomarker discovery and varietal selection.

Crop improvement starts with understanding diversity across populations.

Different varieties carry genomic differences that may influence yield, seed characteristics, resilience or nutritional properties. The challenge is not only detecting that variation, but identifying which differences matter and how they connect to traits of interest.

Graph-based pangenomics moves beyond the limitations of comparing every sample against a single linear reference by representing multiple genomes within the same structure, with each accession represented as an explicit path through it. With the GenoGra platform, genomic variation can be explored together with gene annotations and phenotypic metadata, supporting population studies, genotyping, biomarker discovery and varietal selection within the same genomic resource.


Ricinus communis: from population diversity to candidate variants

As part of a commissioned research project, GenoGra applied graph-based pangenomics to Ricinus communis, a diploid species with ten chromosomes characterised by substantial intraspecific diversity.

Using the platform, a chromosome-resolved pangenome was built as ten graphs, one per chromosome, integrating publicly available genomic resources with selected long-read sequenced accessions.

A substantial share of the species’ molecular variation occurs within populations, making a single reference insufficient to capture its full diversity. This is particularly relevant for structural variants such as insertions, deletions, translocations and copy number variations.

The resulting graphs captured more than 5 million variants, providing a broader representation of the genetic diversity within the species. 

By combining this variation with functional annotations and trait information, GenoGra made it possible to move from millions of variants towards a smaller set of biologically relevant candidates, prioritised according to genomic location, predicted impact and association with agronomic traits. 

As new accessions are added, they can be analysed within the existing pangenomic framework, progressively expanding the population resource and refining the search for candidates of interest.

View the full Ricinus communis poster


Eragrostis tef: linking structural variation to phenotype

The second study was conducted by GenoGra in collaboration with the Institute of Plant Sciences at Scuola Superiore Sant’Anna and IGA Technology Services

Eragrostis tef is an allotetraploid cereal whose two subgenomes carry duplicated and highly similar gene copies.

Using GenoGra, a pangenome graph was built from 26 chromosome-scale assemblies representing its agrobiodiversity, with one graph per chromosome and the A and B subgenomes kept distinct. The analysis then focused on the genomic regions associated with seed colour. 

In the graph, every accession is represented as an explicit path. Structural alleles at a locus of interest can therefore be read directly from the graph and sample-specific paths explored together with seed-colour metadata.

At the analysed locus, the graph resolved distinct insertions and deletions and showed which accessions carried each alternative allele. 

Importantly, the structural variant itself can become a genotypable marker, rather than relying only on nearby SNPs as indirect proxies. 

View the full Eragrostis tef poster


From diversity to selection

Across two species with different genomic challenges, the principle is the same: population-scale pangenomics provides a broader view of diversity, genotyping identifies how that diversity is represented in individual samples, and functional and phenotypic information helps connect genomic variation to traits of interest.

This progressively narrows the search towards candidate biomarkers that can support varietal selection and crop improvement.

This is where pangenomics moves beyond describing diversity and starts making it usable.