Team Leader, SRE
jobgether
Italy
Tempo pieno
Questo annuncio proviene da Lever
Accountabilities
- Lead and develop an SRE team, owning the full career lifecycle of direct reports, including onboarding, feedback, performance management, progression, and hiring.
- Coach engineers on both technical craft and interpersonal skills, creating clear opportunities for growth and development.
- Foster a healthy, collaborative team environment by understanding team dynamics, addressing conflicts constructively, and maintaining effective retrospective practices.
- Represent the SRE team across engineering and with senior leadership, communicating priorities, challenges, progress, and technical direction.
- Define and prioritize SRE objectives, balancing operational commitments with project delivery and protecting the team's focus.
- Own the support rotation and on-call model, ensuring sustainable operational coverage and effective incident response.
- Provide technical leadership across core infrastructure, including Kubernetes, AWS, PostgreSQL, DNS, TLS, CI infrastructure, and related platform services.
- Drive the evolution of reliability practices, including Service Level Objectives (SLOs), error budgets, incident management, observability, and post-incident improvements.
- Partner with Security teams on infrastructure threats, patching, security controls, audits, and compliance requirements.
- Oversee infrastructure-related vendor relationships, including renewals and commercial discussions in collaboration with senior leadership.
- Stay hands-on enough to review technical work, challenge architectural decisions, contribute to complex problems, and provide credible technical direction during incidents.
- Help mature the organization's reliability practices by identifying operational gaps and turning lessons from incidents into lasting engineering improvements.
- Balance short-term operational demands with longer-term platform investments and engineering goals.
Requirements
- Proven experience leading an SRE, infrastructure, platform, or DevOps engineering team, with direct responsibility for team members' growth, performance, and career progression.
- Strong people-management and coaching skills, with demonstrated ability to develop engineers and provide clear, constructive feedback.
- Experience hiring engineers and assessing both technical capability and broader engineering judgment.
- Strong conflict-resolution and team-dynamics skills, with the ability to build commitment around shared organizational goals.
- Deep hands-on experience in Site Reliability Engineering, DevOps, cloud infrastructure, or a closely related discipline.
- Production experience with Kubernetes, including operational troubleshooting, reliability, scaling, and real-world failure scenarios.
- Significant experience with AWS and cloud infrastructure at meaningful scale.
- Hands-on experience building, enabling, or scaling AI infrastructure.
- Strong understanding of observability principles and practices.
- Experience with Infrastructure as Code, particularly Terraform.
- Experience with CI/CD platforms such as GitLab CI, GitHub Actions, Jenkins, or comparable technologies.
- Strong knowledge of Docker and shell scripting.
- Experience owning or operating reliability practices including incident response, on-call, SLOs, error budgets, and post-incident improvement processes.
- Previous experience working in regulated environments and understanding the associated operational and compliance requirements.
- Exceptional prioritization skills, particularly when operational workload competes with project delivery.
- Strong written communication skills and comfort working in an asynchronous, globally distributed environment.
- Ability to build strong relationships across engineering and become a trusted partner for teams bringing reliability challenges forward.
- Experience with a backend programming language such as Elixir, Java, Clojure, Node.js, Python, or similar is a plus.
- Familiarity with modern observability technologies such as OpenTelemetry, distributed tracing, or Honeycomb is advantageous.
- Experience with PostgreSQL or Aurora operations, including performance, connection pools, and query optimization, is beneficial.
- Experience administering Linux systems outside cloud environments is a plus.
- Knowledge of infrastructure security from both defensive and offensive perspectives is advantageous.
- Familiarity with cloud cost management and FinOps is beneficial.
- Experience growing an engineering team from a small starting point, including establishing hiring standards, is a plus.
- Fluent English communication skills are required.
- Ability to work effectively in a fully remote and asynchronous environment.
- Fully remote working environment.
- Flexible, asynchronous working model that allows you to organize your schedule around your life.
- Opportunity to work with a globally distributed engineering organization.
- Significant ownership over both technical direction and people development.
- Approximately 60% individual-contributor technical work and 40% leadership responsibilities.
- Opportunity to shape and mature an evolving SRE and reliability practice.
- Exposure to large-scale cloud infrastructure, Kubernetes, AWS, PostgreSQL, observability, CI/CD, security, and AI infrastructure.
- Flexible paid time off.
- Flexible working hours.
- 16 weeks of paid parental leave.
- Budget for coworking spaces, learning, wellness, and gym memberships.
- Mental health support services.
- Stock options.
- Home office budget and IT equipment.
- Annual salary range of $75,450–$169,700 USD, with actual compensation determined by factors such as location, experience, relevant skills, training, business needs, and market conditions.
- Compensation and benefits are structured according to location and local market considerations.
- Start date: As soon as possible.
- Location: Romania, with a fully remote working arrangement.
Questo annuncio proviene da Lever. Vedi l'annuncio originale ↗