via Lever · 24 septembre 2026 ·il y a 1 jour

Senior Machine Learning Engineer, LLM Inference Optimization

jobgether
France Temps plein
Cette annonce provient de Lever
Voir l'annonce originale ↗

Accountabilities:

  • Own optimization initiatives for specific model families, customer endpoints, and inference serving backends.

  • Evaluate inference engines and recommend practical serving configurations based on workload requirements.

  • Diagnose and resolve model quality, performance, and reliability regressions during production rollouts.

  • Optimize LLM and VLM endpoints for latency, throughput, memory efficiency, GPU utilization, model quality, and cost per token.

  • Deploy, configure, benchmark, and extend modern inference engines such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, or equivalent technologies.

  • Build and productionize model-compression workflows, including quantization, quantization-aware training, distillation, low-bit serving, and accuracy recovery.

  • Implement or integrate advanced inference techniques such as speculative decoding, draft models, KV-cache optimization, prefix caching, chunked prefill, continuous batching, and disaggregated prefill/decode serving.

  • Develop reproducible benchmark harnesses covering TTFT, TPOT, tokens per second per GPU, p95/p99 latency, GPU memory usage, reliability, and cost per token.

  • Partner with GPU kernel and platform engineers to identify bottlenecks across model code, kernels, runtimes, schedulers, gateways, and cluster infrastructure.

  • Investigate performance trade-offs quantitatively and use benchmark results to guide optimization decisions.

  • Produce clear design documentation, performance reports, rollout plans, and technical explanations for internal and customer-facing stakeholders.

  • Contribute to safe, measurable, and reliable production rollouts of inference improvements.

Requirements


  • Strong software engineering skills in Python and PyTorch.

  • Hands-on experience deploying, operating, or optimizing LLM, VLM, or high-throughput transformer inference systems.

  • Practical experience with at least one modern inference stack, such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, or an equivalent internal system.

  • Strong understanding of transformer inference bottlenecks, including KV cache, attention mechanisms, memory bandwidth, batching, parallelism, and long-context serving.

  • Ability to reason quantitatively about latency, throughput, model quality, resource utilization, and cost trade-offs.

  • Experience diagnosing complex performance problems and translating findings into production improvements.

  • Strong communication skills and the ability to collaborate effectively with research, kernel, infrastructure, product, and customer teams.

  • Experience with quantization-aware training, post-training quantization, FP8, INT8, INT4, NVFP4, MXFP4, AWQ, GPTQ, SmoothQuant, or related optimization techniques is a plus.

  • Familiarity with distillation, speculative decoding, EAGLE, Medusa, multi-token prediction, or other inference acceleration approaches is advantageous.

  • Experience supporting agentic workloads involving tool calling, structured outputs, streaming APIs, high concurrency, or multi-step orchestration is a plus.

  • Familiarity with CUDA or Triton is beneficial, even if the role is not primarily focused on kernel engineering.

  • Contributions to open-source projects such as vLLM, SGLang, TensorRT-LLM, FlashInfer, LMCache, PyTorch, Triton, Ray, KServe, or related technologies are advantageous.

  • Ability to work independently, take ownership, and operate effectively in a fast-moving technical environment.
Benefits:
  • Competitive compensation.

  • Career growth and continuous learning opportunities.

  • Flexibility and significant ownership over technical work.

  • Collaborative and innovative working environment.

  • Opportunity to work on impactful AI infrastructure and inference optimization projects.

  • Exposure to advanced LLM and VLM serving technologies and large-scale AI workloads.

  • International environment with experienced engineering and AI teams.

Le marché pour ce type de poste

Offres similaires
62
postes Ingénierie à France
Temps plein
83%
des offres Ingénierie en France
Télétravail possible
4%
des offres Ingénierie
jobgether

200 postes ouverts · Argentina, Austria, Belgium, Denmark, France +12

📊 Ingénierie · France
27 935
offres actives
3.5%
Remote
Ø 1d
Ø en ligne
Compétences les plus demandées
ExcelERPISOPythonAWSCI/CDSQLAzureAgileLean

Questions fréquentes

Combien d'offres Ingénierie sont disponibles à France ?
Actuellement 62 postes en Ingénierie à France sur AlmostHired, dans 20 entreprises différentes. Nos données sont mises à jour quotidiennement.
Est-ce que les postes Ingénierie offrent du télétravail ?
4% des offres Ingénierie en France permettent le télétravail, partiel ou total. Pour filtrer spécifiquement les postes en remote, utilisez AlmostHired.
Comment savoir si je corresponds à cette offre ?
Déposez votre CV — notre IA compare votre profil aux exigences du poste et vous donne un score de compatibilité précis, avec les compétences qui correspondent et celles qui manquent.