via Lever · 24. September 2026 ·vor 1 Tag

Senior Machine Learning Engineer, LLM Inference Optimization

jobgether
Germany Vollzeit
Diese Anzeige stammt von Lever
Zum Original-Inserat ↗

Accountabilities:

  • Own optimization initiatives for specific model families, customer endpoints, and inference serving backends.

  • Evaluate inference engines and recommend practical serving configurations based on workload requirements.

  • Diagnose and resolve model quality, performance, and reliability regressions during production rollouts.

  • Optimize LLM and VLM endpoints for latency, throughput, memory efficiency, GPU utilization, model quality, and cost per token.

  • Deploy, configure, benchmark, and extend modern inference engines such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, or equivalent technologies.

  • Build and productionize model-compression workflows, including quantization, quantization-aware training, distillation, low-bit serving, and accuracy recovery.

  • Implement or integrate advanced inference techniques such as speculative decoding, draft models, KV-cache optimization, prefix caching, chunked prefill, continuous batching, and disaggregated prefill/decode serving.

  • Develop reproducible benchmark harnesses covering TTFT, TPOT, tokens per second per GPU, p95/p99 latency, GPU memory usage, reliability, and cost per token.

  • Partner with GPU kernel and platform engineers to identify bottlenecks across model code, kernels, runtimes, schedulers, gateways, and cluster infrastructure.

  • Investigate performance trade-offs quantitatively and use benchmark results to guide optimization decisions.

  • Produce clear design documentation, performance reports, rollout plans, and technical explanations for internal and customer-facing stakeholders.

  • Contribute to safe, measurable, and reliable production rollouts of inference improvements.

Requirements


  • Strong software engineering skills in Python and PyTorch.

  • Hands-on experience deploying, operating, or optimizing LLM, VLM, or high-throughput transformer inference systems.

  • Practical experience with at least one modern inference stack, such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, or an equivalent internal system.

  • Strong understanding of transformer inference bottlenecks, including KV cache, attention mechanisms, memory bandwidth, batching, parallelism, and long-context serving.

  • Ability to reason quantitatively about latency, throughput, model quality, resource utilization, and cost trade-offs.

  • Experience diagnosing complex performance problems and translating findings into production improvements.

  • Strong communication skills and the ability to collaborate effectively with research, kernel, infrastructure, product, and customer teams.

  • Experience with quantization-aware training, post-training quantization, FP8, INT8, INT4, NVFP4, MXFP4, AWQ, GPTQ, SmoothQuant, or related optimization techniques is a plus.

  • Familiarity with distillation, speculative decoding, EAGLE, Medusa, multi-token prediction, or other inference acceleration approaches is advantageous.

  • Experience supporting agentic workloads involving tool calling, structured outputs, streaming APIs, high concurrency, or multi-step orchestration is a plus.

  • Familiarity with CUDA or Triton is beneficial, even if the role is not primarily focused on kernel engineering.

  • Contributions to open-source projects such as vLLM, SGLang, TensorRT-LLM, FlashInfer, LMCache, PyTorch, Triton, Ray, KServe, or related technologies are advantageous.

  • Ability to work independently, take ownership, and operate effectively in a fast-moving technical environment.
Benefits:
  • Competitive compensation.

  • Career growth and continuous learning opportunities.

  • Flexibility and significant ownership over technical work.

  • Collaborative and innovative working environment.

  • Opportunity to work on impactful AI infrastructure and inference optimization projects.

  • Exposure to advanced LLM and VLM serving technologies and large-scale AI workloads.

  • International environment with experienced engineering and AI teams.

Der Markt für diese Art von Stelle

Ähnliche Angebote
79
Ingenieurwesen in Germany
Vollzeit
81%
der Ingenieurwesen-Angebote in Deutschland
Remote möglich
29%
der Ingenieurwesen-Angebote
jobgether

200 offene Stellen · Argentina, Austria, Belgium, Denmark, France +12

📊 Ingenieurwesen · Deutschland
1.826
aktive Stellen
32.1%
Remote
Ø 3d
Ø online
Gefragte Skills
ExcelERPISOPythonAWSCI/CDSQLAzureAgileLean

Häufige Fragen

Wie viele Ingenieurwesen-Jobs gibt es in Germany?
Aktuell 79 Stellen im Bereich Ingenieurwesen in Germany auf AlmostHired, bei 26 verschiedenen Unternehmen. Unsere Daten werden täglich aktualisiert.
Bieten Ingenieurwesen-Stellen Home Office an?
29% der Ingenieurwesen-Angebote in Deutschland erlauben Remote-Arbeit, teilweise oder vollständig. Um gezielt nach Remote-Stellen zu filtern, nutze AlmostHired.
Wie erfahre ich, ob ich für diese Stelle passe?
Lad deinen CV hoch — unsere KI vergleicht dein Profil mit den Stellenanforderungen und zeigt dir einen präzisen Match-Score, inklusive passender und fehlender Skills.