via Lever · 22 September 2026 ·1 day ago

Senior Software Engineer - Reliability, Infrastructure, and Tooling

jobgether
UK Full-time
This listing is from Lever
View original listing ↗

Accountabilities

  • Ramp up on a complex global architecture involving distributed databases, messaging systems, networking infrastructure, Kubernetes, and other core platform technologies, identifying areas of reliability debt and improvement.

  • Design and ship reliability-focused engineering work directly within production codebases, including load balancing, load shedding, instrumentation, scalability, and efficiency improvements.

  • Build and evolve internal infrastructure and developer tooling that enables product engineering teams to independently operate reliable workloads.

  • Partner closely with product development teams to co-design systems and ensure reliability, security, maintainability, and operational readiness are considered throughout development.

  • Develop observability capabilities that make system behavior measurable, understandable, and actionable, using appropriate signals and visualization techniques.

  • Participate in a shared on-call rotation and contribute to effective incident response, investigation, remediation, and prevention of recurring reliability issues.

  • Investigate complex system-level problems across distributed infrastructure, networking, application behavior, and production environments.

  • Improve configuration management and infrastructure practices across diverse systems, reducing unnecessary complexity, errors, and technical debt.

  • Contribute technical perspectives to architectural discussions and help establish engineering practices that support both short-term delivery and long-term scalability.

  • Support systems with demanding workloads, including real-time media, secure customer code execution, advanced networking, and other highly concurrent services.

  • Collaborate with engineering partners on potentially contentious reliability and operational decisions with clarity, pragmatism, and strong technical judgment.

  • Automate repetitive operational processes wherever possible to improve engineering efficiency and reduce manual intervention.
Requirements
  • Strong professional experience building and operating non-trivial production applications, particularly systems involving high concurrency, distributed workloads, or complex control loops.

  • Significant experience with Kubernetes or an equivalent large-scale container orchestration platform.

  • Strong understanding of Linux internals and networking, with the ability to investigate issues across multiple layers of a production system.

  • Proven experience using observability, monitoring, logging, and related tooling to diagnose difficult production problems.

  • Experience operating large-scale, globally distributed systems, including the configuration management, reliability challenges, and technical debt that emerge as systems grow.

  • Experience responding to and managing complex production incidents, including identifying root causes and implementing durable corrective actions.

  • Experience operating open-source infrastructure technologies such as Kafka, ClickHouse, or comparable distributed systems.

  • Strong systems-thinking skills and an ability to reason about infrastructure in terms of signals, feedback, dependencies, and control mechanisms.

  • Strong communication and collaboration skills, particularly when working with partner engineering teams and navigating competing priorities.

  • A pragmatic approach to engineering that balances immediate delivery needs with long-term maintainability, reliability, and operational cost.

  • A strong interest in observability, reliability engineering, clean configuration, automation, and reducing operational complexity.

  • Nice to have: experience with data engineering and analytics.

  • Nice to have: experience with global Layer 3 networking.

  • Nice to have: experience operating systems with long-lived workloads such as real-time media.

  • Nice to have: experience in Google SRE or another high-scale reliability engineering environment.

  • Nice to have: experience working with compliance frameworks such as PCI.
Benefits
  • $135,000–$300,000 USD compensation range.

  • Equity participation as part of the overall compensation package.

  • Fully remote work with opportunities for collaboration across a globally distributed organization.

  • Health, dental, and vision benefits.

  • Flexible vacation policy.

  • Opportunity to work on infrastructure supporting large-scale real-time and AI applications.

  • Opportunity to contribute to open-source projects alongside experienced engineers.

  • Exposure to challenging distributed-systems problems involving real-time media, secure compute, networking, Kubernetes, and globally distributed infrastructure.

  • Opportunity to influence reliability architecture, engineering practices, and internal developer tooling.

  • Shared on-call practices designed to keep production experience connected across infrastructure and product engineering teams.

  • Equal opportunity employment and reasonable accommodation support throughout the hiring process.

The market for this type of role

Similar openings
68
Engineering roles in UK
Full-time
80%
of Engineering roles in the UK
Remote possible
12%
of Engineering roles
jobgether

200 open positions · Argentina, Austria, Belgium, Denmark, France +12

📊 Engineering · the UK
5,205
active jobs
11.3%
Remote
Ø 2d
avg. online
Top skills in demand
ExcelERPISOPythonAWSCI/CDSQLAzureAgileLean

Frequently asked questions

How many Engineering jobs are available in UK?
Currently 68 Engineering roles in UK on AlmostHired, across 22 different companies. Our data is updated daily.
Do Engineering roles offer remote work?
12% of Engineering roles in the UK allow remote work, either partial or full. To filter specifically for remote positions, use AlmostHired.
How do I know if I match this role?
Upload your CV — our AI compares your profile to the job requirements and gives you a precise match score, with matching and missing skills.