Senior Site Reliability Engineer (SRE, Compute Node Team)
jobgether
Ireland
Full-time
This listing is from Lever
Accountabilities
- Ensure the reliability, availability, and performance of compute nodes responsible for running virtual machines.
- Analyze and debug complex Linux systems across both user space and kernel space.
- Investigate system capabilities, limitations, dependencies, and trade-offs across different layers of the operating system and infrastructure stack.
- Troubleshoot complex production issues involving CPU, memory, NUMA, cgroups, and scheduling.
- Work hands-on with virtualization technologies, primarily QEMU/KVM and Linux-native technologies.
- Analyze VM lifecycle behavior, performance characteristics, resource utilization, and failure modes.
- Support and improve containerized workloads using Linux-native mechanisms such as namespaces and cgroups.
- Design and evolve observability for the compute node layer, including metrics, logs, traces, alerts, SLIs, and SLOs.
- Build reliability signals that provide clear and actionable insight into system behavior.
- Lead or contribute to incident response, ensuring production issues are diagnosed and resolved efficiently.
- Conduct structured root-cause analysis and develop corrective actions for recurring or systemic reliability issues.
- Lead and contribute to postmortems focused on long-term reliability improvements rather than short-term remediation alone.
- Identify opportunities to automate operational processes and improve the resilience of compute infrastructure.
- Collaborate closely with platform, kernel/hypervisor, GPU, and infrastructure teams on system design and operational improvements.
- Contribute to improving the operability, scalability, and maintainability of node-level services.
- Investigate performance issues across multiple layers of the compute stack and develop practical engineering solutions.
- Help establish reliability and observability as core capabilities of the compute platform.
Requirements
- Significant professional experience in Site Reliability Engineering, Systems Engineering, Linux infrastructure, or a closely related field.
- Deep expertise in Linux, including strong understanding of both user space and kernel space.
- Knowledge of important Linux kernel subsystems, including scheduling, memory management, filesystems, cgroups, and namespaces.
- Strong understanding of system boundaries, constraints, dependencies, and trade-offs across different infrastructure layers.
- Hands-on experience with QEMU/KVM and a solid understanding of virtualization technologies.
- Understanding of virtual machine lifecycles, performance characteristics, resource management, and failure modes.
- Practical experience with containers, Linux namespaces, and cgroups.
- Strong understanding of resource isolation, allocation, and control in containerized environments.
- Excellent debugging skills and the ability to reason systematically about complex system failures.
- Structured, hypothesis-driven approach to incident investigation and troubleshooting.
- Strong understanding of the SRE discipline, including the relationship between software engineering, operations, reliability, and system design.
- Experience building and operating observability stacks, rather than simply consuming existing monitoring dashboards.
- Ability to translate complex system behavior into actionable reliability signals, alerts, SLIs, and SLOs.
- Strong analytical and problem-solving skills, with the ability to investigate issues across operating-system and infrastructure layers.
- Experience operating production systems and responding effectively to reliability and performance incidents.
- Strong communication and collaboration skills when working with multidisciplinary infrastructure and engineering teams.
- Ability to take ownership of complex technical problems and drive them through investigation, resolution, and long-term improvement.
- Experience with Kubernetes internals or node-level components is an advantage.
- Hands-on experience with low-level Linux debugging tools such as perf, eBPF, ftrace, strace, or kernel crash dumps is beneficial.
- Familiarity with large-scale compute or bare-metal infrastructure is a plus.
- Contributions to open-source infrastructure or systems software are advantageous.
- Experience debugging hardware- and driver-level issues, including GPUs, NVLink, or InfiniBand, is a strong plus.
- Competitive compensation.
- Career growth and continuous learning opportunities.
- Flexibility and significant ownership in your work.
- Collaborative and innovative engineering environment.
- Opportunity to work on impactful AI and cloud infrastructure projects.
- Exposure to large-scale compute, Linux systems, virtualization, and distributed infrastructure.
- Opportunity to collaborate with highly skilled international engineering teams.
- Inclusive workplace committed to equal employment opportunities.
- Workplace accommodations available during the application process where required.
- Employment is subject to authorization to work in the country where the position is based.
This listing is from Lever. View original listing ↗