via Lever · 18 de septiembre de 2026 ·hace 1 día

Senior Site Reliability Engineer (SRE, Compute Node Team)

jobgether
Spain Tiempo completo
Este anuncio proviene de Lever
Ver oferta original ↗

Accountabilities

  • Ensure the reliability, availability, and performance of compute nodes responsible for running virtual machines.

  • Analyze and debug complex Linux systems across both user space and kernel space.

  • Investigate system capabilities, limitations, dependencies, and trade-offs across different layers of the operating system and infrastructure stack.

  • Troubleshoot complex production issues involving CPU, memory, NUMA, cgroups, and scheduling.

  • Work hands-on with virtualization technologies, primarily QEMU/KVM and Linux-native technologies.

  • Analyze VM lifecycle behavior, performance characteristics, resource utilization, and failure modes.

  • Support and improve containerized workloads using Linux-native mechanisms such as namespaces and cgroups.

  • Design and evolve observability for the compute node layer, including metrics, logs, traces, alerts, SLIs, and SLOs.

  • Build reliability signals that provide clear and actionable insight into system behavior.

  • Lead or contribute to incident response, ensuring production issues are diagnosed and resolved efficiently.

  • Conduct structured root-cause analysis and develop corrective actions for recurring or systemic reliability issues.

  • Lead and contribute to postmortems focused on long-term reliability improvements rather than short-term remediation alone.

  • Identify opportunities to automate operational processes and improve the resilience of compute infrastructure.

  • Collaborate closely with platform, kernel/hypervisor, GPU, and infrastructure teams on system design and operational improvements.

  • Contribute to improving the operability, scalability, and maintainability of node-level services.

  • Investigate performance issues across multiple layers of the compute stack and develop practical engineering solutions.

  • Help establish reliability and observability as core capabilities of the compute platform.

Requirements


  • Significant professional experience in Site Reliability Engineering, Systems Engineering, Linux infrastructure, or a closely related field.

  • Deep expertise in Linux, including strong understanding of both user space and kernel space.

  • Knowledge of important Linux kernel subsystems, including scheduling, memory management, filesystems, cgroups, and namespaces.

  • Strong understanding of system boundaries, constraints, dependencies, and trade-offs across different infrastructure layers.

  • Hands-on experience with QEMU/KVM and a solid understanding of virtualization technologies.

  • Understanding of virtual machine lifecycles, performance characteristics, resource management, and failure modes.

  • Practical experience with containers, Linux namespaces, and cgroups.

  • Strong understanding of resource isolation, allocation, and control in containerized environments.

  • Excellent debugging skills and the ability to reason systematically about complex system failures.

  • Structured, hypothesis-driven approach to incident investigation and troubleshooting.

  • Strong understanding of the SRE discipline, including the relationship between software engineering, operations, reliability, and system design.

  • Experience building and operating observability stacks, rather than simply consuming existing monitoring dashboards.

  • Ability to translate complex system behavior into actionable reliability signals, alerts, SLIs, and SLOs.

  • Strong analytical and problem-solving skills, with the ability to investigate issues across operating-system and infrastructure layers.

  • Experience operating production systems and responding effectively to reliability and performance incidents.

  • Strong communication and collaboration skills when working with multidisciplinary infrastructure and engineering teams.

  • Ability to take ownership of complex technical problems and drive them through investigation, resolution, and long-term improvement.

  • Experience with Kubernetes internals or node-level components is an advantage.

  • Hands-on experience with low-level Linux debugging tools such as perf, eBPF, ftrace, strace, or kernel crash dumps is beneficial.

  • Familiarity with large-scale compute or bare-metal infrastructure is a plus.

  • Contributions to open-source infrastructure or systems software are advantageous.

  • Experience debugging hardware- and driver-level issues, including GPUs, NVLink, or InfiniBand, is a strong plus.
Benefits:
  • Competitive compensation.

  • Career growth and continuous learning opportunities.

  • Flexibility and significant ownership in your work.

  • Collaborative and innovative engineering environment.

  • Opportunity to work on impactful AI and cloud infrastructure projects.

  • Exposure to large-scale compute, Linux systems, virtualization, and distributed infrastructure.

  • Opportunity to collaborate with highly skilled international engineering teams.

  • Inclusive workplace committed to equal employment opportunities.

  • Workplace accommodations available during the application process where required.

  • Employment is subject to authorization to work in the country where the position is based.

El mercado para este tipo de puesto

Ofertas similares
79
puestos de Ingeniería en Spain
Jornada completa
82%
de las ofertas de Ingeniería en España
Teletrabajo posible
31%
de las ofertas de Ingeniería
jobgether

200 open positions · Austria, Belgium, Denmark, France, Germany +11

📊 Ingeniería · España
720
active jobs
32.5%
Remote
Ø 3d
avg. online
Top skills in demand
ExcelERPISOPythonAWSCI/CDSQLAzureAgileLean

Preguntas frecuentes

¿Cuántos empleos de Ingeniería hay disponibles en Spain?
Actualmente 79 puestos de Ingeniería en Spain en AlmostHired, en 26 empresas diferentes. Nuestros datos se actualizan a diario.
¿Los puestos de Ingeniería ofrecen teletrabajo?
31% de las ofertas de Ingeniería en España permiten teletrabajo, parcial o completo. Para filtrar específicamente puestos en remoto, usa AlmostHired.
¿Cómo sé si encajo en esta oferta?
Sube tu CV — nuestra IA compara tu perfil con los requisitos del puesto y te da una puntuación de coincidencia precisa, con habilidades coincidentes y faltantes.