← All open roles

AI Inference Kernel Optimization Engineer

Engineering5+ yearsRemote · Bangalore, India

Vidman AI is a capacity business behind an API: every fraction of utilization won in the serving stack is room to cut the price of a token. As an AI Inference Kernel Optimization Engineer you work at the layer where that is decided — the compute kernels, memory paths, and serving runtimes that execute open models on our GPUs. Your work shows up in two places customers feel: latency, and the price list.

What you'll do

  • 01Write and tune custom compute kernels — attention, matrix multiplication, normalisation — in CUDA, Triton, and CUTLASS where the serving stack leaves performance on the table.
  • 02Optimise memory movement — cache reuse, KV-cache management, operator fusion — so kernels are bound by math, not by bandwidth.
  • 03Implement low-precision inference — INT4, FP8, INT8 quantization paths — matched to the hardware actually run.
  • 04Profile at the hardware level with Nsight Systems, Compute Sanitizer, and PTX/SASS inspection, and act on what you find into shipped improvements.
  • 05Work inside the deployed serving runtimes — vLLM, TensorRT-LLM — and upstream what belongs upstream.
  • 06Partner with research and platform engineering so new model architectures land on kernels that can actually serve them.

What you bring

  • At least 5 years in systems software or ML systems engineering, with real time on GPU performance work.
  • Deep proficiency in C++ and Python, plus mastery of at least one low-level parallel programming framework: CUDA, Triton, or CUTLASS.
  • Working fluency with hardware-level profiling tools and the ability to read assembly-level execution (PTX/SASS) when the profiler is not enough.
  • Familiarity with the modern serving stack — vLLM, TensorRT-LLM, or a comparable runtime — and with how transformer inference actually spends its time and memory.
  • Evidence over adjectives: kernels written or tuned, with before-and-after numbers.

Nice to have

  • Experience with deep-learning compiler stacks: MLIR, LLVM, or torch.compile internals.
  • Contributions to vLLM, TensorRT-LLM, Triton, or another inference-relevant open-source project, readable by us.
  • KV-cache and paged-attention internals, past the tutorial level.
  • Hardware-software co-design exposure — working with compiler or silicon teams on next-generation hardware should support.
  • An advanced degree in CS, CE, or EE is welcome; equivalent depth built in production counts the same.

Apply for this role

Five minutes: your resume, your links, and a paragraph about why this one. We read every application.

← All open roles