HPC Software & Operations Intern
About Corintis
Corintis offers innovative microfluidic cooling technologies for AI chips/GPUs and CPUs used in data centres. Working with many of the world’s largest tech companies, our solutions improve compute sustainability and tackle the excessive electricity consumption associated with data centre cooling, which consumes more electricity than New York and London combined.
Ranked as the No.1 Engineering startup in Switzerland for 2025, Corintis offers a friendly and team-oriented workplace, bringing together over 85 people from over 35 nationalities to solve the most significant computing challenges of tomorrow. Based in the EPFL Campus in St. Sulpice, we are closely connected to the local ecosystem and are located a few minutes walk from Lake Geneva.
The Role
As an HPC Software & Operations Intern, you will help us operate, optimise and develop our High Performance Computing (HPC) environment. Your primary objective is to learn how our engineering software and workflows can run efficiently on multiple cores, multiple GPUs and multiple nodes.
This internship is a hands-on learning experience. You will see the full HPC execution chain, from software adaptation and parallel workloads to software modules, data management, scheduling and cluster operations. You will start with assisted work and progressively work with more technical autonomy. We do not measure you as a junior employee. We measure your progress from assisted work to autonomous work.
Key Responsibilities
Adapt existing code and engineering workflows to the HPC environment, including multi-core, multi-GPU and multi-node execution.
Write and maintain launch scripts and job scripts, and work with the scheduler and the job queues.
Run and troubleshoot MPI workloads, and analyse jobs that fail or perform poorly.
Benchmark workloads, test their scalability and validate them across multiple nodes.
Analyse CPU, GPU, memory, network and I/O utilisation to find bottlenecks and scaling limits.
Manage HPC software environments, modules and dependencies.
Support cluster users with their workloads and with data staging, data access and storage workflows.
Take part in day-to-day cluster operations and incident analysis, together with the infrastructure and software teams.
Write procedures, best practices and runbooks.
Analyse failed or poorly performing jobs
Work on MPI tuning, workload decomposition, CPU/GPU affinity, UCX, profiling, multi-node GPU execution and containerised HPC environments.
Requirements
Student in Computer Science, Software Engineering, Computational Science, Applied Mathematics, Physics or a related field
No professional HPC experience is necessary
Knowledge of Linux and of Python, Bash, C or C++
Knowledge of, or interest in, MPI and parallel or distributed computing
Knowledge of, or interest in, GPU computing (CUDA)
Strong problem-solving skills
Preferred: knowledge of software dependency and module management, storage and data workflows, or performance analysis
This Is a Great Fit If You:
Want to know how software behaves on real computing infrastructure
Like to find the root cause when a job fails or runs slowly
Want practical knowledge of HPC cluster architecture, distributed execution models and job scheduling
Want to understand how applications interact with the compute infrastructure below them
Value guidance from mentors and want to become more autonomous step by step
This Won’t Be the Right Role for You If:
You want to work only on application code and do not want operations or user support tasks.
You prefer to work alone. This role requires close work with the infrastructure and software teams.
You prefer to work fully remotely. This role requires on-site collaboration.
- Department
- Glacierware
- Location
- Lausanne
- Employment type
- Internship