via Lever · 16 september 2026 ·3 dagen geleden

Staff Site Reliability Engineer

jobgether
Netherlands Voltijd
Deze vacature komt van Lever
Bekijk originele vacature ↗

Accountabilities

  • Define and implement SLIs and SLOs for critical production request paths, ensuring reliability objectives are visible, measurable, reviewed, and connected to engineering decisions.

  • Introduce and champion error budgets as a practical framework for balancing reliability investments with product and feature delivery.

  • Establish and maintain the reliability metrics used by engineering leadership to evaluate progress and identify areas requiring investment.

  • Strengthen the complete incident management lifecycle, including detection, response, communication, escalation, postmortems, and follow-up actions.

  • Improve alert quality, anomaly detection, escalation processes, and shared operational tooling in collaboration with infrastructure teams.

  • Lead reliability assessments for high-risk changes and new services, covering production readiness, capacity, failure modes, rollback strategies, and operational risks.

  • Introduce deliberate failure testing, game days, and chaos exercises to identify weaknesses and validate safe operational limits before incidents occur.

  • Work directly with engineering teams on complex reliability challenges through focused engagements, leaving behind stronger practices and clear ownership.

  • Coach Staff and Lead engineers to become reliability advocates within their respective teams and help establish distributed SRE ownership.

  • Develop lightweight, repeatable operational standards covering production readiness, on-call practices, runbooks, change safety, and service operability.

  • Partner with architects and technical leads to ensure reliability and failure tolerance are incorporated into system design rather than addressed after deployment.

  • Remain hands-on during production incidents and investigations, building tooling, dashboards, automation, and reference implementations where appropriate.

  • Promote effective use of AI for incident investigation, telemetry analysis, postmortem development, runbook creation, observability, and reliability tooling.

  • Help structure operational data, alerts, dashboards, and runbooks so that both engineers and AI agents can safely interpret and act on production signals.

  • Contribute production fixes and improvements directly through code and infrastructure changes rather than limiting the role to recommendations and reviews.
Requirements
  • 10+ years of engineering experience, including at least 3 years in SRE, production engineering, or a reliability-focused Staff Engineer role operating across multiple teams.

  • Demonstrated experience owning reliability at a platform or organizational level rather than only for an individual service.

  • Deep practical experience designing and implementing SLIs, SLOs, and error budgets, including successfully driving adoption across product and engineering teams.

  • Strong incident leadership experience, including managing high-severity, customer-facing incidents and leading effective postmortems that result in measurable improvements.

  • Advanced understanding of distributed-system failure modes, including database and cache saturation, cascading failures, retry storms, capacity constraints, graceful degradation, and load shedding.

  • Strong hands-on experience with Kubernetes, AWS, and modern observability platforms such as Datadog or comparable technologies.

  • Ability to read and write production code in Go, TypeScript, or a similar language, as well as work with infrastructure as code.

  • Demonstrated ability to influence teams without direct authority and successfully change engineering practices across an organization.

  • Strong coaching and mentoring skills, with evidence of developing engineers into effective reliability owners.

  • Exceptional written and verbal communication skills, with the ability to clearly communicate incidents, risks, technical trade-offs, and reliability priorities to both engineers and executives.

  • Strong preference for asynchronous, documented decision-making and clear technical communication.

  • Practical experience using AI tools for incident investigation, telemetry analysis, runbook and postmortem development, and engineering tooling.

  • Understanding of how operational data, alerts, dashboards, and runbooks should be structured to support safe AI-assisted diagnosis and operations.

  • Pragmatic approach to reliability, with the ability to balance operational risk, engineering investment, delivery speed, and business priorities.

  • Experience in fraud detection, identity, payments, or other real-time and adversarial environments is an asset.

  • Experience with multi-region architectures, cell-based architectures, or failure-isolation strategies is a plus.

  • Experience operating Elasticsearch, Redis, DynamoDB, or Kafka at scale and understanding their failure modes is beneficial.

  • Familiarity with FinOps and cloud infrastructure cost-versus-reliability trade-offs is an advantage.

  • Must be authorized to work from the hiring location; visa sponsorship is not provided.
Benefits
  • Fully remote working environment.

  • Opportunity to become the first dedicated Site Reliability Engineer and establish organization-wide reliability practices.

  • High level of autonomy and direct influence over engineering standards, operational practices, and platform reliability.

  • Opportunity to work across multiple engineering teams and critical production systems.

  • Close collaboration with engineering leadership, architects, infrastructure teams, and technical leads.

  • Opportunity to shape AI-assisted reliability practices and the future of production operations.

  • Strong focus on professional growth, technical leadership, coaching, and knowledge sharing.

  • Inclusive, globally distributed engineering environment that values diverse perspectives and backgrounds.

  • For US-based employees, the stated cash compensation range is $177,000–$240,000 USD, with actual offers varying according to factors such as experience, skills, education, certifications, and market conditions. Compensation may differ for other hiring locations.

  • Remote work eligibility is subject to applicable regulatory and security requirements in the candidate's location.

De markt voor dit type functie

Vergelijkbare vacatures
59
Engineering-functies in Netherlands
Fulltime
42%
van de Engineering-vacatures in Nederland
Remote mogelijk
18%
van de Engineering-vacatures
jobgether

200 open positions · Argentina, Austria, Belgium, France, Germany +11

📊 Engineering · Nederland
925
active jobs
16.9%
Remote
Ø 3d
avg. online
Top skills in demand
ExcelERPISOPythonAWSCI/CDSQLAzureAgileLean

Veelgestelde vragen

Hoeveel Engineering-banen zijn er in Netherlands?
Momenteel 59 Engineering-functies in Netherlands op AlmostHired, bij 19 verschillende bedrijven. Onze gegevens worden dagelijks bijgewerkt.
Bieden Engineering-functies thuiswerken aan?
18% van de Engineering-vacatures in Nederland staat thuiswerken toe, gedeeltelijk of volledig. Om specifiek op remote functies te filteren, gebruik AlmostHired.
Hoe weet ik of ik bij deze functie pas?
Upload je CV — onze AI vergelijkt je profiel met de functievereisten en geeft je een precieze match score, met overeenkomende en ontbrekende vaardigheden.