Find a vacancy that works for you. Send us your CV to receive a personalized offer.
Find me a jobWe are looking for a Senior GPU Compute & AI Kernel Engineer to work at the lowest and most performance-critical layer of the AI stack — where compute kernels, accelerator architecture and machine learning workloads meet. This is a hands-on engineering role for a systems-minded C++ practitioner who is fluent in GPU and AI accelerator programming models and who measures success in achieved throughput, latency and memory efficiency rather than in features shipped.
The successful candidate will design and optimize high-performance kernels for modern deep learning workloads, collaborate closely with compiler, hardware and ML engineering teams, and contribute to the software stacks that make AI accelerators usable in production. The role suits an engineer who is comfortable moving across the full stack — from framework-level operator implementations down to instruction-level performance analysis on the target hardware.
- Design, implement and optimize high-performance compute kernels for GPUs and AI accelerators using modern C++ and one or more of CUDA, SYCL, OpenCL, Triton or Apple Metal
- Profile, analyze and tune parallel algorithms for performance and memory efficiency, identifying and removing bottlenecks across compute, bandwidth and occupancy
- Accelerate machine learning workloads in PyTorch or TensorFlow, including the implementation of custom operators and kernels for Transformer, LLM and CNN architectures
- Map deep learning operators onto target accelerator architectures, taking advantage of memory hierarchies, tensor units and vendor-specific execution models
- Collaborate with compiler engineers on kernel code generation and lowering paths, working with technologies such as LLVM and MLIR
- Partner with hardware and architecture teams to evaluate performance characteristics of new accelerators and inform hardware/software co-design decisions
- Establish and maintain performance benchmarking, regression tracking and profiling methodology for kernel and workload performance
- Debug complex correctness and performance issues spanning framework, runtime, compiler and hardware layers
- Contribute to accelerator software stacks, runtimes and framework integrations, upstream where applicable
- Apply and uphold software engineering best practices across Linux development environments — version control, code review, automated testing and CI/CD
- Document design decisions, optimization techniques and performance findings, and share knowledge across compiler, hardware and ML engineering teams
- 5+ years of software development experience with modern C++ (C++17/20 preferred)
- Strong experience developing high-performance compute kernels
- Hands-on GPU programming experience with one or more of CUDA, SYCL, OpenCL, Triton or Apple Metal
- Solid understanding of GPU and AI accelerator architectures
- Experience optimizing parallel algorithms for performance and memory efficiency
- Experience developing or optimizing machine learning workloads using PyTorch or TensorFlow
- Good understanding of modern deep learning architectures, including Transformers, LLMs and CNNs
- Familiarity with compiler technologies such as LLVM and/or MLIR
- Knowledge of High Performance Computing (HPC) concepts and performance optimization techniques
- Experience working with low-level hardware or processor architectures (RISC-V an advantage)
- Strong debugging, profiling and performance analysis skills
- Ability to work across the software stack in close collaboration with compiler, hardware and ML engineers
- Experience with Linux development environments and software engineering best practices (Git, CI/CD, code reviews)
- Experience with AI accelerator software stacks
- Experience implementing custom ML operators or kernels
- Knowledge of distributed AI training or inference
- Experience contributing to compiler, runtime or framework development
- Background in embedded systems or computer architecture
- Familiarity with tensor compilers and graph optimization frameworks
