Internship: Towards expressive and tractable surrogate models for large scale inverse problems

Inria Montbonnot Montbonnot, France
Researcher / Scientist Software Development Statistics and Probability Remote Sensing 26 days left

About the position

Context
This internship is part of an ongoing collaboration between the Statify research team at Inria and the Institut de Planétologie et d’Astrophysique de Grenoble (IPAG) at Université Grenoble Alpes (UGA).
The internship builds on GLLiM (Gaussian Locally-Linear Mapping), a statistical modelling approach developed by the Statify team for solving Bayesian inverse problems using physical forward models and simulations. The approach is implemented in the open-source xLLiM scientific library and is also used in the PlanetGLLiM application.
GLLiM provides an efficient surrogate modelling framework for inverse problems in which evaluating the physical model may be computationally expensive. However, the current implementation has several limitations. In particular, model training currently relies on a batch implementation that requires the complete training dataset to be loaded into memory, limiting its applicability to moderately large datasets. In addition, the current model parameterization is primarily designed for real-valued data that are not bounded, and provides only limited flexibility for modelling the noise component.
The objective of this internship is to address some of these limitations and extend GLLiM towards more expressive, scalable, and computationally efficient surrogate models, with a particular focus on high-dimensional remote-sensing applications.
Assignment
The internship will focus on developing and evaluating one or several extensions of the GLLiM framework. Depending on the candidate's background and interests, possible research directions include:
Constrained parameter estimation: incorporating physical constraints, such as bounded domains of variation for physical parameters.
Incremental learning: developing an online/incremental learning strategy that processes training data sequentially, reducing memory requirements and enabling the use of substantially larger datasets.
More flexible noise modelling: introducing an additional latent component to obtain a more parsimonious and expressive parameterization of the noise.
Complex-valued modelling: reformulating the model using complex-valued Gaussian distributions to support complex-valued observations.
The selected developments will be implemented efficiently in C++ with Python bindings and integrated into the existing GLLiM ecosystem, including the xLLiM toolbox.
The proposed methods will then be validated through experiments and benchmarks, with particular attention to accuracy, computational efficiency, scalability, and robustness. The ultimate goal is to improve the applicability of GLLiM to challenging, high-dimensional inverse problems arising in remote sensing.
Main activities
Depending on the chosen research direction, the internship will involve:
Formulating mathematically one or more extensions of the GLLiM methodology.
Studying and implementing the corresponding algorithms in Python and C++.
Designing and conducting experiments, tests, and performance benchmarks.
Integrating the resulting developments into the existing xLLiM codebase.
Ensuring backward compatibility and performing non-regression testing.
Evaluating the accuracy, efficiency, and scalability of the proposed approaches.
Applying the methods to representative high-dimensional remote-sensing problems.
Writing technical and user documentation.
Presenting and discussing results with the research and development team.
Skills
Required background
Currently pursuing an M2 degree or equivalent in computer science, applied mathematics, statistics, or a related field.
Good programming skills in C++ and Python.
Solid knowledge of probability and statistics, with familiarity with topics such as Gaussian mixture models, the EM algorithm, or Bayesian modelling.
Strong foundations in linear algebra and optimization.
Experience with scientific computing and statistical modelling.
Familiarity with software development practices and tools such as GitHub/GitLab, continuous integration, and Docker.
Additional qualities :
Interest in the interaction between mathematical modelling, machine learning, and inverse problems.
Ability to translate mathematical concepts into robust and efficient software implementations.
Analytical and modelling skills, including the ability to formulate specifications and document technical developments.
Curiosity and willingness to work in a research environment and explore new approaches.
Ability to work independently while collaborating effectively with researchers and software engineers.
Rigorous, well-organized approach to problem solving.
Good communication and interpersonal skills.

This listing was collected from a public source and is reproduced here for information only. Always confirm the details on the original posting before applying.
View the original posting

Similar positions