Incident Operations Lead
jobgether
Ireland
Full-time
This listing is from Lever
Accountabilities
- Build and lead a dedicated Incident Operations function responsible for coordinating the organization’s most critical incidents.
- Recruit, develop, and certify Incident Commanders and establish a 24x7 follow-the-sun rotation across APAC, EMEA, and AMER.
- Maintain effective regional handoffs and participate in the incident command roster to remain closely connected to operational realities.
- Develop and run game days, tabletop exercises, simulations, and certification programs to keep responders prepared for high-severity events.
- Establish a blameless incident culture focused on identifying process and system improvements rather than individual fault.
- Own the incident severity model in partnership with Risk, ensuring severity levels appropriately reflect financial and regulatory materiality.
- Define and maintain escalation paths, response thresholds, leadership escalation criteria, and procedures for unanswered alerts or pages.
- Keep the service catalogue and service ownership information accurate and actionable so teams can quickly identify the right technical owners during incidents.
- Coordinate the operational bridge between engineering and technical support teams resolving incidents and partner-facing communications teams providing customer updates.
- Establish clear communication channels, provide accurate incident facts, and maintain appropriate update cadences throughout major incidents.
- Lead post-incident retrospectives in collaboration with SRE and ensure follow-up actions are documented, assigned, prioritized, and delivered within defined service levels.
- Track overdue post-incident reviews and actions by team and incident, ensuring every outstanding item has a clear next step and owner.
- Define and maintain trusted operational KPIs, including time to respond and time to mitigate, with clear definitions and reliable underlying data.
- Establish defensible performance baselines before setting improvement targets and use severity-based analysis to drive continuous improvement.
- Build incident-management processes as scalable, documented, versioned operating products that can support significant organizational growth.
- Lead an AI and agentic automation roadmap covering incident setup, timeline generation, RCA drafting, action tracking, update reminders, and review follow-ups.
- Establish clear boundaries for automated workflows, defining where AI agents can operate independently and where human Incident Commander judgment remains essential.
- Translate incident learnings into concise, actionable guidance and distribute them across engineering teams so lessons from individual failures can drive broader organizational improvement.
Requirements
- 5+ years of experience in production engineering, Site Reliability Engineering, technical operations, or a closely related technical discipline.
- Proven hands-on experience commanding high-severity or major production incidents.
- Demonstrated experience building or standing up an incident command or major-incident management function, rather than only participating in an established process.
- Experience designing severity models, escalation frameworks, incident-response processes, and operational governance.
- Proven ability to lead distributed teams across multiple time zones and operate a 24x7 on-call or follow-the-sun rotation.
- Strong influence and stakeholder-management skills, with the ability to drive action from engineers and teams outside your direct reporting line.
- Confident judgment when making and defending severity or escalation decisions during high-pressure situations.
- Strong understanding of reliability metrics and the ability to distinguish improvements in measured performance from genuine improvements in operational outcomes.
- Excellent scope discipline and the ability to establish clear ownership boundaries while ensuring issues are routed without creating operational gaps.
- Strong written and verbal communication skills, including the ability to maintain calm and clarity during incident bridges.
- Ability to communicate effectively with senior and executive stakeholders during high-impact incidents without minimizing or overstating the situation.
- Strong understanding of FinTech environments and the trust, reliability, and operational requirements associated with API-driven financial platforms.
- Experience using AI or agentic automation to reduce operational toil and improve incident-management workflows.
- Excellent organizational skills and a strong focus on documentation, process quality, and continuous improvement.
- Formal incident command, ITIL, Major Incident Management, or crisis-management training is a plus.
- Experience running certification programs, game days, simulations, or operational drills is highly valued.
- Familiarity with modern incident-management and on-call platforms is advantageous.
- Experience building and maintaining service catalogues or service ownership registries is a plus.
- Experience partnering with program management or reliability teams to convert incident follow-ups into funded improvement initiatives is desirable.
- Familiarity with regulatory incident-reporting obligations such as DORA, Reg SCI, FINRA, or equivalent frameworks is advantageous.
- Experience in online securities trading, capital markets, or another regulated and market-hours-sensitive environment is a plus.
- Experience deploying an incident-management operating model across multiple regions or entities is advantageous.
- Competitive salary.
- Stock options.
- Health benefits.
- One-time USD $500 home-office setup allowance for new hires.
- USD $150 monthly stipend provided through a company card.
- Opportunity to work in a globally distributed environment.
- Exposure to a highly technical, reliability-focused operating environment.
- Opportunity to lead and shape a critical operational function from the ground up.
- Significant scope to influence incident management, reliability, automation, and organizational readiness.
- Opportunity to work cross-functionally with engineering, SRE, Risk, communications, and reliability teams.
- Opportunity to develop and implement AI-powered operational workflows and automation.
- Inclusive environment that values curiosity, empathy, accountability, and diverse perspectives.
This listing is from Lever. View original listing ↗