Senior Site Reliability Engineer
jobgether
UK
Full-time
This listing is from Lever
Accountabilities
- Operate, maintain, and continuously improve Linux-based infrastructure, with a strong focus on Debian and Ubuntu environments.
- Deploy, manage, and scale production Kubernetes clusters across bare-metal, virtualized, and on-premise environments, overseeing upgrades, node pools, networking, storage, and security hardening.
- Design, implement, and maintain complex networking architectures covering VLANs, L2/L3 routing, VPNs, and multi-site connectivity.
- Build and maintain infrastructure automation using Ansible, Bash, Python, Git-based workflows, and GitOps practices, including automated provisioning through PXE boot, Preseed, and cloud-init.
- Deploy and maintain observability and monitoring platforms such as Prometheus, Grafana, Loki, ELK, and Graylog, ensuring operational data generates actionable insights.
- Lead incident response and escalation activities, troubleshoot complex infrastructure issues, and implement improvements that increase availability and reduce latency.
- Define and implement SLOs and SLIs across physical infrastructure, networking, virtualization, and software services to establish measurable reliability standards.
- Optimize alerting and monitoring pipelines while establishing effective on-call schedules to provide operational coverage across time zones.
- Create and maintain Standard Operating Procedures for recurring infrastructure operations, maintenance, troubleshooting, and incident management.
- Coordinate physical infrastructure maintenance, including hardware issues, periodic maintenance, and data-center operations.
- Manage virtualization and orchestration layers using technologies such as OpenStack, Proxmox, and VMware.
- Contribute to the overall architecture and evolution of infrastructure products, ensuring solutions remain scalable and reliable.
- Plan infrastructure capacity and resources for future initiatives based on projected demand and business growth.
- Partner with development teams to improve system quality, optimize resource utilization, and strengthen engineering practices.
- Collaborate with cross-functional stakeholders to align infrastructure priorities with broader product and customer needs.
Requirements
- Expert-level, hands-on experience operating Kubernetes in production, including cluster lifecycle management, networking, storage, security, and scaling.
- Strong network engineering expertise is essential, particularly across VLANs, L2/L3 routing, VPNs, and multi-site connectivity.
- Strong Linux systems administration skills, particularly with Debian and Ubuntu.
- Solid understanding of networking fundamentals and the ability to design and operate complex network architectures.
- Proven experience developing infrastructure automation using Ansible, Bash and/or Python, Git-based workflows, and GitOps methodologies.
- Practical experience with observability platforms such as Prometheus, Grafana, ELK, Loki, or Graylog.
- Experience working with virtualization technologies including OpenStack, Proxmox, and VMware.
- Experience with bare-metal provisioning and MAAS (Metal as a Service).
- Strong understanding of distributed systems and container orchestration.
- A process-oriented mindset, with the ability to create SOPs and operational procedures from the ground up.
- Experience managing production incidents, escalation processes, and on-call rotations.
- Ability to work independently and make sound technical decisions in a fast-paced, engineering-driven environment.
- Strong communication and collaboration skills, combined with a high level of technical ownership and alignment with team values.
- Fluent English is mandatory.
- Experience with service mesh technologies such as Istio or Linkerd, or advanced CNI implementations, is a plus.
- Knowledge of Cloudflare APIs, DNS automation, or tunnel configurations is advantageous.
- Experience with GPU infrastructure, node preparation, resource scheduling, security practices such as RBAC, firewalls, and network policies is beneficial.
- Familiarity with IT asset management or license tracking workflows is an advantage.
- Experience working across multiple time zones and establishing SRE or reliability frameworks within growing organizations is highly valued.
- 100% remote work within the EU time zone, with CET ±2 hours preferred.
- Flexible working hours designed to support autonomy and effective collaboration.
- High-impact position with significant ownership and the opportunity to influence infrastructure strategy.
- Opportunity to work with a modern technology stack spanning Kubernetes, cloud infrastructure, networking, virtualization, automation, and observability.
- Collaborative and international engineering environment with exposure to complex infrastructure challenges.
- Significant autonomy to shape operational processes, reliability practices, and technical solutions.
- Opportunity to contribute to ambitious cloud infrastructure initiatives with a strong focus on reliability, automation, and continuous improvement.
This listing is from Lever. View original listing ↗