About the position
Context
Numerical simulation is a strategic tool, useful for accelerating discoveries and innovations. The performance of simulators are thus of paramount importance to enable scalable high temporal resolution studies and to consider a wider range of scientific hypotheses. High Performance Computing (HPC) systems offer parallelism in multiple ways to address this computational demand: they are composed of many nodes, each with multiple cores (which may be heterogeneous or rely on Simultaneous Multi-Threading, SMT), using Single Instruction Multiple Data (SIMD) vectorization, with Graphics Processing Unit (GPU) accelerators.
Both the software and hardware stacks present an array of diverse levers (tunable, such as process parallelism, SIMD, thread and data mapping for Non-Uniform Memory Access, burst buffers, frequency, or prefetching) to align applications (without code changes) with the unique characteristics of the hardware they run on. Interestingly, the recent code transformation capabilities of LLMs, assuming they can be verified, further offer a lever to tune applications through prompting describing code changes to apply. While tuning these levers provides gains, co-optimizing them together is key as their behaviors can synergize: thread and data mapping co-optimization is more efficient than a greedy approach, where threads and data are optimized in order (up to 1.7x speedup for userspace users) [4]. A full-scale optimization further achieves 1.8x speedup and 6x energy savings [5], demonstrating that it is of paramount importance for exascale sustainability.
Exploring such parameter optimizations is the core task of this postdoc position: it leverages statistical/AI methods to exploit patterns, transform codes, and focus the search to avoid local optima. To address this challenge, we rely on an ongoing collaboration between Inria TADaaM [1], IFPEN [2], and LIP6 [3] (Sorbonne University). LiP6 brings expertise in SIMD code transformation, Inria in LLM-based agentic workflows for HPC applications, and IFPEN in full-stack performance tuning and representative real-world simulation workloads.
Assignment
This work aims to express diverse OpenMP and SIMD code transformations and assess their correctness with respect to the original code. We will rely on LLMs (e.g., llama, deepseek, queen coder) and an agentic framework to apply such code transformations [7]. We start by defining a set of prompts (with chain of thought or prompting) to apply simple source code transformations (e.g., loop interchange, unroll, structure change). These transformations are conditioned by the used LLM: two LLMs might not produce the same transformation on the same code with the same prompt. We will explore existing LLMs and select prompts across already explored transformations: OpenMP and SIMD. Then, we will diversify the transformations by incrementally complexifying them along with the target codes. A key insight of the proposed work is the synergistic relationship with LLMs: as their capabilities grow, they enable more aggressive and valid transformations across diverse codebases. However, LLMs alone are unlikely to maximize performance efficiency, as their training data reflects average hardware rather than system-specific conditions. To overcome this, we will enhance the characterization of each transformation with its impact over the rest of the stack thanks to CORHPEX [7]. Finally, we propose to employ SOTA verification methods (e.g., checkpoint-restart CaRV, alive symbolic analysis) along with unit tests (e.g., as illustrated by the MIPP infrastructure [8]) to assess the semantic equivalence between the original and the LLM transformed program. We hope that the proposed methodology could be transformed for other programming models and architectures.
Reference
[1]
https://team.inria.fr/tadaam/
[2]
https://www.ifpenergiesnouvelles.fr/
[3]
https://www.lip6.fr/
[4] Efficient thread/page/parallelism autotuning for NUMA systems. M Popov, A Jimborean, D Black-Schaffer. ICS 2019
[5] Optimizing performance and energy across problem sizes through a search space exploration and machine learning. L Scravaglieri, M Popov, L Lima Pilla, A Guermouche, O Aumage, E Saillard. JPDC 2023
[6] Compiler, Runtime, and Hardware Parameters Design Space Exploration. L Scravaglieri, A Anciaux-Sedrakian, O Aumage, T Guignon, M Popov. IPDPS 2025
[7] Llm-vectorizer: Llm-based verified loop vectorizer. J Taneja, A Laird, C Yan, M Musuvathi, SK Lahiri. CGO 2025
[8] MIPP: A Portable C++ SIMD Wrapper and its use for Error Correction Coding in 5G Standard. A Cassagne, O Aumage, D Barthou, C Leroux, C Jégo. WPMVP 2018.
Main activities
We identify the following activities:
1. Build a basic SIMD code transformation workflow with LLMs.
2. Build an infrastructure to assess their semantic equivalence with respect to the original code. We consider checkpoint-restart CaRV or alive symbolic analysis techniques.
3. Expand the scope of codes and transformations. A potential lead is the optimization manuals of the target hardware.
4. Contextulaize the impact of source code transformation with respect to the rest of the computing stack. We can rely on CORHPEX to navigate the resulting optimization space.
Skills
The profile we are looking for this job is someone who communicates well in English and has initiative and autonomy to conduct this research project while tackling the diverse related technical challenges.
The person should have experience in AI techniques, HPC, autotuning, performance evaluation and statistical methods.
This listing was collected from a public source and is reproduced here for
information only. Always confirm the details on the original posting before applying.
View the original posting