via Lever · 16 septembre 2026 ·il y a 3 jours

Staff Site Reliability Engineer

jobgether
France Temps plein
Cette annonce provient de Lever
Voir l'annonce originale ↗

Accountabilities

  • Define and implement SLIs and SLOs for critical production request paths, ensuring reliability objectives are visible, measurable, reviewed, and connected to engineering decisions.

  • Introduce and champion error budgets as a practical framework for balancing reliability investments with product and feature delivery.

  • Establish and maintain the reliability metrics used by engineering leadership to evaluate progress and identify areas requiring investment.

  • Strengthen the complete incident management lifecycle, including detection, response, communication, escalation, postmortems, and follow-up actions.

  • Improve alert quality, anomaly detection, escalation processes, and shared operational tooling in collaboration with infrastructure teams.

  • Lead reliability assessments for high-risk changes and new services, covering production readiness, capacity, failure modes, rollback strategies, and operational risks.

  • Introduce deliberate failure testing, game days, and chaos exercises to identify weaknesses and validate safe operational limits before incidents occur.

  • Work directly with engineering teams on complex reliability challenges through focused engagements, leaving behind stronger practices and clear ownership.

  • Coach Staff and Lead engineers to become reliability advocates within their respective teams and help establish distributed SRE ownership.

  • Develop lightweight, repeatable operational standards covering production readiness, on-call practices, runbooks, change safety, and service operability.

  • Partner with architects and technical leads to ensure reliability and failure tolerance are incorporated into system design rather than addressed after deployment.

  • Remain hands-on during production incidents and investigations, building tooling, dashboards, automation, and reference implementations where appropriate.

  • Promote effective use of AI for incident investigation, telemetry analysis, postmortem development, runbook creation, observability, and reliability tooling.

  • Help structure operational data, alerts, dashboards, and runbooks so that both engineers and AI agents can safely interpret and act on production signals.

  • Contribute production fixes and improvements directly through code and infrastructure changes rather than limiting the role to recommendations and reviews.
Requirements
  • 10+ years of engineering experience, including at least 3 years in SRE, production engineering, or a reliability-focused Staff Engineer role operating across multiple teams.

  • Demonstrated experience owning reliability at a platform or organizational level rather than only for an individual service.

  • Deep practical experience designing and implementing SLIs, SLOs, and error budgets, including successfully driving adoption across product and engineering teams.

  • Strong incident leadership experience, including managing high-severity, customer-facing incidents and leading effective postmortems that result in measurable improvements.

  • Advanced understanding of distributed-system failure modes, including database and cache saturation, cascading failures, retry storms, capacity constraints, graceful degradation, and load shedding.

  • Strong hands-on experience with Kubernetes, AWS, and modern observability platforms such as Datadog or comparable technologies.

  • Ability to read and write production code in Go, TypeScript, or a similar language, as well as work with infrastructure as code.

  • Demonstrated ability to influence teams without direct authority and successfully change engineering practices across an organization.

  • Strong coaching and mentoring skills, with evidence of developing engineers into effective reliability owners.

  • Exceptional written and verbal communication skills, with the ability to clearly communicate incidents, risks, technical trade-offs, and reliability priorities to both engineers and executives.

  • Strong preference for asynchronous, documented decision-making and clear technical communication.

  • Practical experience using AI tools for incident investigation, telemetry analysis, runbook and postmortem development, and engineering tooling.

  • Understanding of how operational data, alerts, dashboards, and runbooks should be structured to support safe AI-assisted diagnosis and operations.

  • Pragmatic approach to reliability, with the ability to balance operational risk, engineering investment, delivery speed, and business priorities.

  • Experience in fraud detection, identity, payments, or other real-time and adversarial environments is an asset.

  • Experience with multi-region architectures, cell-based architectures, or failure-isolation strategies is a plus.

  • Experience operating Elasticsearch, Redis, DynamoDB, or Kafka at scale and understanding their failure modes is beneficial.

  • Familiarity with FinOps and cloud infrastructure cost-versus-reliability trade-offs is an advantage.

  • Must be authorized to work from the hiring location; visa sponsorship is not provided.
Benefits
  • Fully remote working environment.

  • Opportunity to become the first dedicated Site Reliability Engineer and establish organization-wide reliability practices.

  • High level of autonomy and direct influence over engineering standards, operational practices, and platform reliability.

  • Opportunity to work across multiple engineering teams and critical production systems.

  • Close collaboration with engineering leadership, architects, infrastructure teams, and technical leads.

  • Opportunity to shape AI-assisted reliability practices and the future of production operations.

  • Strong focus on professional growth, technical leadership, coaching, and knowledge sharing.

  • Inclusive, globally distributed engineering environment that values diverse perspectives and backgrounds.

  • For US-based employees, the stated cash compensation range is $177,000–$240,000 USD, with actual offers varying according to factors such as experience, skills, education, certifications, and market conditions. Compensation may differ for other hiring locations.

  • Remote work eligibility is subject to applicable regulatory and security requirements in the candidate's location.

Le marché pour ce type de poste

Offres similaires
88
postes Ingénierie à France
Temps plein
83%
des offres Ingénierie en France
Télétravail possible
3%
des offres Ingénierie
jobgether

200 postes ouverts · Argentina, Austria, Belgium, France, Germany +11

📊 Ingénierie · France
25 310
offres actives
3.4%
Remote
Ø 1d
Ø en ligne
Compétences les plus demandées
ExcelERPISOPythonAWSCI/CDSQLAzureAgileLean

Questions fréquentes

Combien d'offres Ingénierie sont disponibles à France ?
Actuellement 88 postes en Ingénierie à France sur AlmostHired, dans 29 entreprises différentes. Nos données sont mises à jour quotidiennement.
Est-ce que les postes Ingénierie offrent du télétravail ?
3% des offres Ingénierie en France permettent le télétravail, partiel ou total. Pour filtrer spécifiquement les postes en remote, utilisez AlmostHired.
Comment savoir si je corresponds à cette offre ?
Déposez votre CV — notre IA compare votre profil aux exigences du poste et vous donne un score de compatibilité précis, avec les compétences qui correspondent et celles qui manquent.