Senior Site Reliability Engineer / Kubernetes
jobgether
Switzerland
Vollzeit
Diese Anzeige stammt von Lever
Accountabilities
- Operate, maintain, and continuously improve Linux-based infrastructure, primarily across Debian and Ubuntu environments, ensuring systems remain secure, stable, and performant.
- Deploy, manage, and scale production Kubernetes clusters across bare-metal, virtualized, and on-premise environments, taking ownership of the full cluster lifecycle including upgrades, node pools, networking, storage, and security hardening.
- Design and maintain complex networking architectures covering VLANs, L2/L3 routing, VPNs, and multi-site connectivity, with a strong focus on reliability and performance.
- Build and maintain automated infrastructure provisioning and operational workflows using Ansible, Bash, Python, Git, and GitOps practices, including PXE boot, Preseed, and cloud-init processes.
- Develop and operate comprehensive observability platforms using technologies such as Prometheus, Grafana, Loki, ELK, and Graylog, improving monitoring quality and ensuring alerts provide actionable operational insights.
- Lead incident response and escalation activities, investigate complex infrastructure issues, reduce recurring failures, and improve overall platform availability and latency.
- Define and implement SLOs and SLIs across physical infrastructure, networking, virtualization, and software services to establish measurable reliability standards.
- Establish and maintain effective on-call practices and schedules that provide reliable operational coverage across multiple time zones.
- Create Standard Operating Procedures and repeatable operational processes for infrastructure maintenance, troubleshooting, incident response, and routine platform activities.
- Coordinate physical infrastructure maintenance, including hardware issues, periodic maintenance activities, and data-center operations.
- Manage virtualization and orchestration technologies including OpenStack, Proxmox, and VMware, while contributing to the broader architecture of the platform and its products.
- Plan infrastructure capacity and resources for future initiatives, taking projected demand, scalability, and growth into account.
- Partner with development and cross-functional teams to improve software and infrastructure quality, optimize resource utilization, and strengthen operational practices across the wider organization.
Requirements
- Expert-level, hands-on experience operating Kubernetes in production environments, with a strong understanding of cluster architecture, lifecycle management, networking, storage, and security.
- Strong network engineering expertise is essential, including VLANs, L2/L3 routing, VPNs, multi-site connectivity, and the ability to design and operate complex network architectures.
- Advanced Linux systems administration experience, particularly with Debian and Ubuntu environments.
- Strong understanding of networking fundamentals, distributed systems, container orchestration, and modern infrastructure architecture.
- Proven experience building and maintaining infrastructure automation using Ansible, Bash and/or Python, Git-based workflows, and GitOps practices.
- Practical experience with observability and monitoring platforms such as Prometheus, Grafana, ELK, Loki, or Graylog.
- Experience with virtualization technologies including OpenStack, Proxmox, and/or VMware.
- Experience with bare-metal provisioning and technologies such as MAAS, PXE, Preseed, and cloud-init.
- Experience managing incidents, escalation processes, on-call rotations, and operational reliability in production environments.
- A process-oriented mindset with the ability to create SOPs and operational procedures from the ground up.
- Strong problem-solving and analytical skills, combined with the ability to work independently in a fast-paced, engineering-driven environment.
- Excellent communication and collaboration skills, with the ability to work effectively with distributed, cross-functional teams.
- Fluent English is mandatory, with the ability to communicate complex technical topics clearly.
- Experience with service mesh technologies such as Istio or Linkerd, or advanced CNI implementations, is a plus.
- Knowledge of Cloudflare APIs, DNS automation, or tunnel configuration is advantageous.
- Experience with GPU infrastructure, node preparation, resource scheduling, security practices such as RBAC, firewalls and network policies, or IT asset management is beneficial.
- Experience working across multiple time zones or establishing SRE and reliability frameworks within growing organizations is a strong advantage.
- 100% remote work within the EU time zone, with a preferred working range of CET ±2 hours.
- Flexible working hours designed to support autonomy and effective collaboration across distributed teams.
- High-impact position with significant ownership, responsibility, and freedom to shape infrastructure and reliability practices.
- Opportunity to work with advanced Kubernetes, cloud infrastructure, networking, virtualization, automation, and observability technologies.
- Collaborative and international engineering environment with strong technical expertise and cross-functional exposure.
- Opportunity to influence architecture, operational standards, automation strategies, and long-term platform scalability.
- Fast-paced environment focused on reliability, engineering excellence, innovation, and continuous improvement.
- Opportunity to contribute to ambitious infrastructure projects while developing deeper expertise across modern cloud and SRE technologies.
Diese Anzeige stammt von Lever. Originalanzeige ansehen ↗