via Lever · 18. September 2026 ·vor 1 Tag

Senior Site Reliability Engineer — Token Factory (Inference Platform)

jobgether
Germany Vollzeit
Diese Anzeige stammt von Lever
Zum Original-Inserat ↗

Accountabilities

  • Own the reliability, performance, and observability of the inference platform and its supporting infrastructure.

  • Design, implement, and continuously improve telemetry pipelines covering metrics, logs, and traces.

  • Build monitoring and observability solutions capable of processing large volumes of production signals and converting them into actionable insights.

  • Configure and optimize Kubernetes infrastructure for high availability, scalability, and efficient resource utilization.

  • Tune Kubernetes autoscaling mechanisms to improve the efficiency and utilization of GPU resources.

  • Develop and maintain Terraform modules and infrastructure-as-code patterns that embed resilience and reliability into new clusters and services.

  • Design and improve request-routing, retry, and failure-handling mechanisms to minimize the impact of transient infrastructure or service failures.

  • Develop automation and operational tooling to detect, isolate, and remediate incidents quickly.

  • Create, maintain, and improve runbooks for incident response and operational procedures.

  • Participate in production incident management, troubleshooting issues and restoring services within demanding reliability objectives.

  • Lead or contribute to post-mortem processes and implement corrective actions to prevent recurring incidents.

  • Define and improve reliability practices for high-throughput APIs, including alerting strategies and Service Level Objectives (SLOs).

  • Investigate distributed-system failures and performance issues across infrastructure and application layers.

  • Optimize systems from the kernel and infrastructure layer through to the application layer.

  • Support and improve the operation of GPU-intensive inference workloads and accelerator-based infrastructure.

  • Contribute to scaling the inference platform while balancing performance, reliability, and infrastructure costs.

  • Collaborate closely with software engineers to incorporate reliability and operational excellence into product and platform development.

  • Promote automation, self-healing capabilities, and engineering practices that reduce operational overhead and improve system resilience.

Requirements


  • Significant experience in Site Reliability Engineering, Production Engineering, DevOps, or a closely related infrastructure discipline.

  • Deep practical knowledge of Kubernetes in production environments.

  • Strong experience with Prometheus and Grafana for monitoring, metrics, dashboards, and observability.

  • Advanced experience with Terraform and infrastructure-as-code practices.

  • Strong scripting and automation skills using Python and/or Bash.

  • Solid understanding of distributed systems and the ways production backends can fail under real-world conditions.

  • Experience designing effective alerts, monitoring strategies, and SLOs for high-throughput services or APIs.

  • Strong troubleshooting and debugging skills across infrastructure, networking, operating systems, and application layers.

  • Experience designing systems for high availability, resilience, scalability, and graceful failure recovery.

  • Hands-on experience with GPU-heavy workloads or accelerator-based infrastructure is highly valuable.

  • Familiarity with GPU inference technologies such as vLLM, Triton, Ray, or comparable accelerator and model-serving stacks.

  • Experience with MLOps, model hosting, AI infrastructure, or machine-learning platforms is advantageous.

  • Strong understanding of infrastructure automation, deployment, configuration management, and operational tooling.

  • Ability to analyze complex performance and reliability problems and translate findings into practical engineering improvements.

  • Strong incident-management and root-cause-analysis capabilities.

  • Ability to collaborate effectively with software engineers and other technical teams to integrate reliability into platform development.

  • Proactive mindset with a strong focus on automation, self-healing systems, and continuous improvement.

  • Comfortable working independently, taking ownership of critical infrastructure, and operating effectively in a fast-paced technical environment.
Benefits:
  • Competitive compensation.

  • Career growth and continuous learning opportunities.

  • Flexibility and significant ownership in your work.

  • Collaborative and innovative international working environment.

  • Opportunity to work on high-impact AI infrastructure and inference technologies.

  • Exposure to large-scale GPU infrastructure and complex distributed systems.

  • Opportunity to contribute to infrastructure supporting next-generation multimodal AI applications.

  • Diverse and highly technical international teams.

  • Inclusive workplace committed to equal employment opportunities.

  • Workplace accommodations available throughout the application process where required.

  • Employment is subject to authorization to work in the country where the position is based.

Der Markt für diese Art von Stelle

Ähnliche Angebote
76
Ingenieurwesen in Germany
Vollzeit
81%
der Ingenieurwesen-Angebote in Deutschland
Remote möglich
29%
der Ingenieurwesen-Angebote
jobgether

200 offene Stellen · Austria, Belgium, Denmark, France, Germany +11

📊 Ingenieurwesen · Deutschland
1.928
aktive Stellen
30.2%
Remote
Ø 3d
Ø online
Gefragte Skills
ExcelERPISOPythonAWSCI/CDSQLAzureAgileLean

Häufige Fragen

Wie viele Ingenieurwesen-Jobs gibt es in Germany?
Aktuell 76 Stellen im Bereich Ingenieurwesen in Germany auf AlmostHired, bei 25 verschiedenen Unternehmen. Unsere Daten werden täglich aktualisiert.
Bieten Ingenieurwesen-Stellen Home Office an?
29% der Ingenieurwesen-Angebote in Deutschland erlauben Remote-Arbeit, teilweise oder vollständig. Um gezielt nach Remote-Stellen zu filtern, nutze AlmostHired.
Wie erfahre ich, ob ich für diese Stelle passe?
Lad deinen CV hoch — unsere KI vergleicht dein Profil mit den Stellenanforderungen und zeigt dir einen präzisen Match-Score, inklusive passender und fehlender Skills.