Senior Software Engineer - Reliability, Infrastructure, and Tooling
jobgether
UK
Full-time
This listing is from Lever
Accountabilities
- Ramp up on a complex global architecture involving distributed databases, messaging systems, networking infrastructure, Kubernetes, and other core platform technologies, identifying areas of reliability debt and improvement.
- Design and ship reliability-focused engineering work directly within production codebases, including load balancing, load shedding, instrumentation, scalability, and efficiency improvements.
- Build and evolve internal infrastructure and developer tooling that enables product engineering teams to independently operate reliable workloads.
- Partner closely with product development teams to co-design systems and ensure reliability, security, maintainability, and operational readiness are considered throughout development.
- Develop observability capabilities that make system behavior measurable, understandable, and actionable, using appropriate signals and visualization techniques.
- Participate in a shared on-call rotation and contribute to effective incident response, investigation, remediation, and prevention of recurring reliability issues.
- Investigate complex system-level problems across distributed infrastructure, networking, application behavior, and production environments.
- Improve configuration management and infrastructure practices across diverse systems, reducing unnecessary complexity, errors, and technical debt.
- Contribute technical perspectives to architectural discussions and help establish engineering practices that support both short-term delivery and long-term scalability.
- Support systems with demanding workloads, including real-time media, secure customer code execution, advanced networking, and other highly concurrent services.
- Collaborate with engineering partners on potentially contentious reliability and operational decisions with clarity, pragmatism, and strong technical judgment.
- Automate repetitive operational processes wherever possible to improve engineering efficiency and reduce manual intervention.
- Strong professional experience building and operating non-trivial production applications, particularly systems involving high concurrency, distributed workloads, or complex control loops.
- Significant experience with Kubernetes or an equivalent large-scale container orchestration platform.
- Strong understanding of Linux internals and networking, with the ability to investigate issues across multiple layers of a production system.
- Proven experience using observability, monitoring, logging, and related tooling to diagnose difficult production problems.
- Experience operating large-scale, globally distributed systems, including the configuration management, reliability challenges, and technical debt that emerge as systems grow.
- Experience responding to and managing complex production incidents, including identifying root causes and implementing durable corrective actions.
- Experience operating open-source infrastructure technologies such as Kafka, ClickHouse, or comparable distributed systems.
- Strong systems-thinking skills and an ability to reason about infrastructure in terms of signals, feedback, dependencies, and control mechanisms.
- Strong communication and collaboration skills, particularly when working with partner engineering teams and navigating competing priorities.
- A pragmatic approach to engineering that balances immediate delivery needs with long-term maintainability, reliability, and operational cost.
- A strong interest in observability, reliability engineering, clean configuration, automation, and reducing operational complexity.
- Nice to have: experience with data engineering and analytics.
- Nice to have: experience with global Layer 3 networking.
- Nice to have: experience operating systems with long-lived workloads such as real-time media.
- Nice to have: experience in Google SRE or another high-scale reliability engineering environment.
- Nice to have: experience working with compliance frameworks such as PCI.
- $135,000–$300,000 USD compensation range.
- Equity participation as part of the overall compensation package.
- Fully remote work with opportunities for collaboration across a globally distributed organization.
- Health, dental, and vision benefits.
- Flexible vacation policy.
- Opportunity to work on infrastructure supporting large-scale real-time and AI applications.
- Opportunity to contribute to open-source projects alongside experienced engineers.
- Exposure to challenging distributed-systems problems involving real-time media, secure compute, networking, Kubernetes, and globally distributed infrastructure.
- Opportunity to influence reliability architecture, engineering practices, and internal developer tooling.
- Shared on-call practices designed to keep production experience connected across infrastructure and product engineering teams.
- Equal opportunity employment and reasonable accommodation support throughout the hiring process.
This listing is from Lever. View original listing ↗