Senior Machine Learning Engineer, LLM Inference Optimization
jobgether
Ireland
Full-time
This listing is from Lever
Accountabilities:
- Own optimization initiatives for specific model families, customer endpoints, and inference serving backends.
- Evaluate inference engines and recommend practical serving configurations based on workload requirements.
- Diagnose and resolve model quality, performance, and reliability regressions during production rollouts.
- Optimize LLM and VLM endpoints for latency, throughput, memory efficiency, GPU utilization, model quality, and cost per token.
- Deploy, configure, benchmark, and extend modern inference engines such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, or equivalent technologies.
- Build and productionize model-compression workflows, including quantization, quantization-aware training, distillation, low-bit serving, and accuracy recovery.
- Implement or integrate advanced inference techniques such as speculative decoding, draft models, KV-cache optimization, prefix caching, chunked prefill, continuous batching, and disaggregated prefill/decode serving.
- Develop reproducible benchmark harnesses covering TTFT, TPOT, tokens per second per GPU, p95/p99 latency, GPU memory usage, reliability, and cost per token.
- Partner with GPU kernel and platform engineers to identify bottlenecks across model code, kernels, runtimes, schedulers, gateways, and cluster infrastructure.
- Investigate performance trade-offs quantitatively and use benchmark results to guide optimization decisions.
- Produce clear design documentation, performance reports, rollout plans, and technical explanations for internal and customer-facing stakeholders.
- Contribute to safe, measurable, and reliable production rollouts of inference improvements.
Requirements
- Strong software engineering skills in Python and PyTorch.
- Hands-on experience deploying, operating, or optimizing LLM, VLM, or high-throughput transformer inference systems.
- Practical experience with at least one modern inference stack, such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, or an equivalent internal system.
- Strong understanding of transformer inference bottlenecks, including KV cache, attention mechanisms, memory bandwidth, batching, parallelism, and long-context serving.
- Ability to reason quantitatively about latency, throughput, model quality, resource utilization, and cost trade-offs.
- Experience diagnosing complex performance problems and translating findings into production improvements.
- Strong communication skills and the ability to collaborate effectively with research, kernel, infrastructure, product, and customer teams.
- Experience with quantization-aware training, post-training quantization, FP8, INT8, INT4, NVFP4, MXFP4, AWQ, GPTQ, SmoothQuant, or related optimization techniques is a plus.
- Familiarity with distillation, speculative decoding, EAGLE, Medusa, multi-token prediction, or other inference acceleration approaches is advantageous.
- Experience supporting agentic workloads involving tool calling, structured outputs, streaming APIs, high concurrency, or multi-step orchestration is a plus.
- Familiarity with CUDA or Triton is beneficial, even if the role is not primarily focused on kernel engineering.
- Contributions to open-source projects such as vLLM, SGLang, TensorRT-LLM, FlashInfer, LMCache, PyTorch, Triton, Ray, KServe, or related technologies are advantageous.
- Ability to work independently, take ownership, and operate effectively in a fast-moving technical environment.
- Competitive compensation.
- Career growth and continuous learning opportunities.
- Flexibility and significant ownership over technical work.
- Collaborative and innovative working environment.
- Opportunity to work on impactful AI infrastructure and inference optimization projects.
- Exposure to advanced LLM and VLM serving technologies and large-scale AI workloads.
- International environment with experienced engineering and AI teams.
This listing is from Lever. View original listing ↗