Interpretable Vector Symbolic Architectures for scalable
Genome Interpretation
We are looking for a motivated PhD student to work at the
interface of Artificial Intelligence, Machine Learning,
Bioinformatics, and Genomics. The project will investigate the use
of Vector Symbolic Architectures (VSA), also known as
Hyperdimensional Computing, as a new framework for predicting
phenotypes and disease risk directly from genomic sequencing data.
Project description
Genome Interpretation aims to understand how genetic variation
determines phenotypes, including quantitative traits and disease
risk. This remains a challenging machine-learning problem because
genomic datasets contain millions of variables but comparatively few
samples, making conventional neural networks prone to overfitting
and difficult to interpret.
The PhD project will develop a new approach based on Vector Symbolic
Architectures, in which genomic information is represented using
high-dimensional vectors and compositional algebraic operations.
The student will develop methods to encode genomic information at
multiple biological scales:
variant → gene → pathway → individual
using operations such as binding, bundling, and permutation. These
representations will then be used to build predictive models of
genotype–phenotype relationships.
The project will first be prototyped using yeast whole-genome
sequencing data and hundreds of quantitative phenotypes, allowing
rapid comparison of different VSA representations and learning
strategies. The methods will subsequently be applied to human exome
sequencing data , using publicly available case-control cohorts.
A major component of the project will focus on interpretability. The
student will develop methods based on VSA decoding and unbinding to
identify the variants, genes, and biological pathways contributing
to individual predictions and compare the discovered signals with
known disease-associated loci.
The project therefore combines methodological development in
machine learning with applications to real genomic datasets and
clinically relevant problems.
Main research objectives
The PhD student will:
• develop multi-scale VSA representations for genomic sequencing
data; investigate different hypervector representations, including
binary, bipolar, ternary, and continuous encodings;
• design supervised and potentially self-supervised learning
methods operating on genomic hypervectors;
• benchmark VSA models against conventional machine-learning and
deep-learning approaches;
• develop interpretable decoding methods to quantify variant-,
gene-, and pathway-level contributions;
• apply the developed methods to yeast genotype–phenotype
prediction and human IBD disease-risk prediction;
• analyze the biological relevance of the associations
discovered by the models;
• publish the methodological and biological results in
international journals and conferences.
Required skills
Candidates should have a Master's degree, or equivalent, in Computer
Science, Artificial Intelligence, Machine Learning, Bioinformatics,
Computational Biology, Applied Mathematics, or a related discipline.
Strong candidates should have:
• solid Python programming skills and pytorch;
• bioinformatics or computational genomics knowledge;
• good knowledge of machine learning and statistical learning;
• familiarity with linear algebra, vector representations, and
optimization;
• experience working with scientific datasets;
• ability to independently design, implement, and evaluate
computational methods;
• good written and spoken English (at least B2). The job
interview will be in english
Useful but not mandatory experience
Experience in one or more of the following would be advantageous:
• deep learning and PyTorch;
• hyperdimensional computing or Vector Symbolic Architectures;
• genomic data formats such as VCF;
• population genetics or genotype–phenotype prediction;
• dimensionality reduction and representation learning;
• interpretable or explainable machine learning;
• high-performance computing and large-scale data analysis.
Previous biological training is not required, provided the candidate
is interested in learning the necessary genomics and genetics
concepts.
Candidate profile
We are particularly interested in candidates who enjoy developing
new machine-learning methodology rather than only applying existing
models. The project requires a combination of algorithmic thinking,
mathematical reasoning, programming, and curiosity about biological
problems.
The successful candidate will work on a highly interdisciplinary
project at the frontier between AI and genome biology, with the
opportunity to develop a largely unexplored computational paradigm
for interpretable genome analysis.
Salary is 2300€ brut. The start date is the 1/11/2026
(mandatory).
Applications should be sent to daniele.raimondi@igmm.cnrs.fr
with the following object:
[PHD VSA]: name applicant
(Emails are automatically filtered).