via Lever · 15 de septiembre de 2026 ·hace 4 días

Incident Operations Lead

jobgether
Spain Tiempo completo
Este anuncio proviene de Lever
Ver oferta original ↗

Accountabilities

  • Build and lead a dedicated Incident Operations function responsible for coordinating the organization’s most critical incidents.

  • Recruit, develop, and certify Incident Commanders and establish a 24x7 follow-the-sun rotation across APAC, EMEA, and AMER.

  • Maintain effective regional handoffs and participate in the incident command roster to remain closely connected to operational realities.

  • Develop and run game days, tabletop exercises, simulations, and certification programs to keep responders prepared for high-severity events.

  • Establish a blameless incident culture focused on identifying process and system improvements rather than individual fault.

  • Own the incident severity model in partnership with Risk, ensuring severity levels appropriately reflect financial and regulatory materiality.

  • Define and maintain escalation paths, response thresholds, leadership escalation criteria, and procedures for unanswered alerts or pages.

  • Keep the service catalogue and service ownership information accurate and actionable so teams can quickly identify the right technical owners during incidents.

  • Coordinate the operational bridge between engineering and technical support teams resolving incidents and partner-facing communications teams providing customer updates.

  • Establish clear communication channels, provide accurate incident facts, and maintain appropriate update cadences throughout major incidents.

  • Lead post-incident retrospectives in collaboration with SRE and ensure follow-up actions are documented, assigned, prioritized, and delivered within defined service levels.

  • Track overdue post-incident reviews and actions by team and incident, ensuring every outstanding item has a clear next step and owner.

  • Define and maintain trusted operational KPIs, including time to respond and time to mitigate, with clear definitions and reliable underlying data.

  • Establish defensible performance baselines before setting improvement targets and use severity-based analysis to drive continuous improvement.

  • Build incident-management processes as scalable, documented, versioned operating products that can support significant organizational growth.

  • Lead an AI and agentic automation roadmap covering incident setup, timeline generation, RCA drafting, action tracking, update reminders, and review follow-ups.

  • Establish clear boundaries for automated workflows, defining where AI agents can operate independently and where human Incident Commander judgment remains essential.

  • Translate incident learnings into concise, actionable guidance and distribute them across engineering teams so lessons from individual failures can drive broader organizational improvement.

Requirements


  • 5+ years of experience in production engineering, Site Reliability Engineering, technical operations, or a closely related technical discipline.

  • Proven hands-on experience commanding high-severity or major production incidents.

  • Demonstrated experience building or standing up an incident command or major-incident management function, rather than only participating in an established process.

  • Experience designing severity models, escalation frameworks, incident-response processes, and operational governance.

  • Proven ability to lead distributed teams across multiple time zones and operate a 24x7 on-call or follow-the-sun rotation.

  • Strong influence and stakeholder-management skills, with the ability to drive action from engineers and teams outside your direct reporting line.

  • Confident judgment when making and defending severity or escalation decisions during high-pressure situations.

  • Strong understanding of reliability metrics and the ability to distinguish improvements in measured performance from genuine improvements in operational outcomes.

  • Excellent scope discipline and the ability to establish clear ownership boundaries while ensuring issues are routed without creating operational gaps.

  • Strong written and verbal communication skills, including the ability to maintain calm and clarity during incident bridges.

  • Ability to communicate effectively with senior and executive stakeholders during high-impact incidents without minimizing or overstating the situation.

  • Strong understanding of FinTech environments and the trust, reliability, and operational requirements associated with API-driven financial platforms.

  • Experience using AI or agentic automation to reduce operational toil and improve incident-management workflows.

  • Excellent organizational skills and a strong focus on documentation, process quality, and continuous improvement.

  • Formal incident command, ITIL, Major Incident Management, or crisis-management training is a plus.

  • Experience running certification programs, game days, simulations, or operational drills is highly valued.

  • Familiarity with modern incident-management and on-call platforms is advantageous.

  • Experience building and maintaining service catalogues or service ownership registries is a plus.

  • Experience partnering with program management or reliability teams to convert incident follow-ups into funded improvement initiatives is desirable.

  • Familiarity with regulatory incident-reporting obligations such as DORA, Reg SCI, FINRA, or equivalent frameworks is advantageous.

  • Experience in online securities trading, capital markets, or another regulated and market-hours-sensitive environment is a plus.

  • Experience deploying an incident-management operating model across multiple regions or entities is advantageous.
Benefits:
  • Competitive salary.

  • Stock options.

  • Health benefits.

  • One-time USD $500 home-office setup allowance for new hires.

  • USD $150 monthly stipend provided through a company card.

  • Opportunity to work in a globally distributed environment.

  • Exposure to a highly technical, reliability-focused operating environment.

  • Opportunity to lead and shape a critical operational function from the ground up.

  • Significant scope to influence incident management, reliability, automation, and organizational readiness.

  • Opportunity to work cross-functionally with engineering, SRE, Risk, communications, and reliability teams.

  • Opportunity to develop and implement AI-powered operational workflows and automation.

  • Inclusive environment that values curiosity, empathy, accountability, and diverse perspectives.

El mercado para este tipo de puesto

Ofertas similares
327
ofertas en Spain
Jornada completa
82%
de las ofertas en España
Teletrabajo posible
13%
de las ofertas
jobgether

200 open positions · Argentina, Austria, Belgium, France, Germany +11

📊 Job market · España
10.820
active jobs
13.7%
Remote
Ø 3d
avg. online

Preguntas frecuentes

¿Cuántas ofertas hay disponibles en Spain?
Actualmente 327 puestos en Spain en AlmostHired, en 109 empresas diferentes. Nuestros datos se actualizan a diario.
¿Los puestos en España ofrecen teletrabajo?
13% de las ofertas en España permiten teletrabajo, parcial o completo. Para filtrar específicamente puestos en remoto, usa AlmostHired.
¿Cómo sé si encajo en esta oferta?
Sube tu CV — nuestra IA compara tu perfil con los requisitos del puesto y te da una puntuación de coincidencia precisa, con habilidades coincidentes y faltantes.