Senior AI Engineer (Performance Optimization Lead)
About the Role
About the Role
We are looking for a Senior AI Engineer to lead research and engineering efforts focused on improving the performance and efficiency of AI workloads. This is a hands-on role where you will work on low-level optimization of modern AI systems, identifying bottlenecks and developing novel techniques that reduce latency, improve throughput, and lower inference cost.
If you enjoy understanding how AI systems work under the hood and solving complex systems problems, we'd like to hear from you.
Responsibilities
Design and implement high-performance AI inference systems.
Optimize GPU utilization, memory bandwidth, and model execution.
Research and implement quantization, KV cache, and attention optimizations.
Profile AI workloads and identify system bottlenecks.
Develop benchmarking and performance analysis frameworks.
Work on distributed inference and multi-GPU execution.
Evaluate and reproduce state-of-the-art research.
Collaborate with researchers to transform ideas into production-quality implementations.
Mentor junior engineers on performance engineering best practices.
Minimum qualifications
Bachelor's or Master's degree in Computer Science, Computer Engineering, or a related field.
5+ years of software engineering experience, esp. working on AI inference.
Strong proficiency in C++, Python, and Linux.
Experience with CUDA, PyTorch, or Triton.
Understanding of GPU architecture and parallel programming.
Experience profiling and optimizing large-scale systems.
Strong problem-solving and debugging skills.
Preferred qualifications
Experience with modern inference optimization techniques, including disaggregated serving, prefill optimization, speculative decoding, continuous batching, and KV cache optimization.
Experience with distributed inference and large-scale model execution using tensor, pipeline, and expert parallelism.
Strong debugging, profiling, and performance optimization skills for large-scale GPU-accelerated AI systems.