Staff Site Reliability Engineer
jobgether
Spain
Tiempo completo
Este anuncio proviene de Lever
Accountabilities
- Define and implement SLIs and SLOs for critical production request paths, ensuring reliability objectives are visible, measurable, reviewed, and connected to engineering decisions.
- Introduce and champion error budgets as a practical framework for balancing reliability investments with product and feature delivery.
- Establish and maintain the reliability metrics used by engineering leadership to evaluate progress and identify areas requiring investment.
- Strengthen the complete incident management lifecycle, including detection, response, communication, escalation, postmortems, and follow-up actions.
- Improve alert quality, anomaly detection, escalation processes, and shared operational tooling in collaboration with infrastructure teams.
- Lead reliability assessments for high-risk changes and new services, covering production readiness, capacity, failure modes, rollback strategies, and operational risks.
- Introduce deliberate failure testing, game days, and chaos exercises to identify weaknesses and validate safe operational limits before incidents occur.
- Work directly with engineering teams on complex reliability challenges through focused engagements, leaving behind stronger practices and clear ownership.
- Coach Staff and Lead engineers to become reliability advocates within their respective teams and help establish distributed SRE ownership.
- Develop lightweight, repeatable operational standards covering production readiness, on-call practices, runbooks, change safety, and service operability.
- Partner with architects and technical leads to ensure reliability and failure tolerance are incorporated into system design rather than addressed after deployment.
- Remain hands-on during production incidents and investigations, building tooling, dashboards, automation, and reference implementations where appropriate.
- Promote effective use of AI for incident investigation, telemetry analysis, postmortem development, runbook creation, observability, and reliability tooling.
- Help structure operational data, alerts, dashboards, and runbooks so that both engineers and AI agents can safely interpret and act on production signals.
- Contribute production fixes and improvements directly through code and infrastructure changes rather than limiting the role to recommendations and reviews.
- 10+ years of engineering experience, including at least 3 years in SRE, production engineering, or a reliability-focused Staff Engineer role operating across multiple teams.
- Demonstrated experience owning reliability at a platform or organizational level rather than only for an individual service.
- Deep practical experience designing and implementing SLIs, SLOs, and error budgets, including successfully driving adoption across product and engineering teams.
- Strong incident leadership experience, including managing high-severity, customer-facing incidents and leading effective postmortems that result in measurable improvements.
- Advanced understanding of distributed-system failure modes, including database and cache saturation, cascading failures, retry storms, capacity constraints, graceful degradation, and load shedding.
- Strong hands-on experience with Kubernetes, AWS, and modern observability platforms such as Datadog or comparable technologies.
- Ability to read and write production code in Go, TypeScript, or a similar language, as well as work with infrastructure as code.
- Demonstrated ability to influence teams without direct authority and successfully change engineering practices across an organization.
- Strong coaching and mentoring skills, with evidence of developing engineers into effective reliability owners.
- Exceptional written and verbal communication skills, with the ability to clearly communicate incidents, risks, technical trade-offs, and reliability priorities to both engineers and executives.
- Strong preference for asynchronous, documented decision-making and clear technical communication.
- Practical experience using AI tools for incident investigation, telemetry analysis, runbook and postmortem development, and engineering tooling.
- Understanding of how operational data, alerts, dashboards, and runbooks should be structured to support safe AI-assisted diagnosis and operations.
- Pragmatic approach to reliability, with the ability to balance operational risk, engineering investment, delivery speed, and business priorities.
- Experience in fraud detection, identity, payments, or other real-time and adversarial environments is an asset.
- Experience with multi-region architectures, cell-based architectures, or failure-isolation strategies is a plus.
- Experience operating Elasticsearch, Redis, DynamoDB, or Kafka at scale and understanding their failure modes is beneficial.
- Familiarity with FinOps and cloud infrastructure cost-versus-reliability trade-offs is an advantage.
- Must be authorized to work from the hiring location; visa sponsorship is not provided.
- Fully remote working environment.
- Opportunity to become the first dedicated Site Reliability Engineer and establish organization-wide reliability practices.
- High level of autonomy and direct influence over engineering standards, operational practices, and platform reliability.
- Opportunity to work across multiple engineering teams and critical production systems.
- Close collaboration with engineering leadership, architects, infrastructure teams, and technical leads.
- Opportunity to shape AI-assisted reliability practices and the future of production operations.
- Strong focus on professional growth, technical leadership, coaching, and knowledge sharing.
- Inclusive, globally distributed engineering environment that values diverse perspectives and backgrounds.
- For US-based employees, the stated cash compensation range is $177,000–$240,000 USD, with actual offers varying according to factors such as experience, skills, education, certifications, and market conditions. Compensation may differ for other hiring locations.
- Remote work eligibility is subject to applicable regulatory and security requirements in the candidate's location.
Este anuncio proviene de Lever. Ver anuncio original ↗