Senior / Lead Site Reliability Engineer

Company: Allwyn UK
Apply for the Senior / Lead Site Reliability Engineer
Location: Watford
Job Description:

At the heart of everything we do is our vision to change lives every day, and our mission to grow The National Lottery responsibly and champion its impact.

We are Allwyn UK, part of the Allwyn Entertainment Group – a multi-national lottery operator with a market-leading presence acrossthe USA (Michigan and Illinois) andEurope, includingCzech Republic, Austria, Greece,Cyprusand Italy.

While the main contribution of The National Lottery to society is through the funds togood causes, at Allwynwe put our purpose and values at the heart of everything we do.Join us as we embark on a once-in-a-lifetime, largescale transformation journey by creating a National Lottery that delivers more money togood causes.

We’lltalk a bit more about us further down the page, but for now -let’stalk about the role and whowe’relooking for…

A bit about the role

At Allwyn, the Senior/Lead Site Reliability Engineer is responsible for technical leadership of reliability engineering across the digital estate, ensuring high availability, performance, and resilience of customer-facing systems during both normal operation and peak lottery events.

The role combines hands-on engineering, incident leadership, and ownership of the SRE improvement backlog and reporting, working across platform, product, and operational teams.

Objectives of the role

  • Own reliability outcomes across services using SLOs, SLIs, and error budgets
  • Improve availability, latency, and scalability across Instant-Win and Draw-based platforms
  • Lead incident response and operational readiness, including peak jackpot events
  • Drive automation and platform maturity, reducing manual operational effort
  • Establish clear reporting on reliability, incidents, and service health trends
  • Supporting with thought-leadership and developing long-term roadmap

Whatyou’llbe doing

Reliability engineering & technical leadership

  • Define and govern SLOs / SLIs / error budgets across critical services
  • Lead reliability design reviews across:
  • Player Identity & Protection systems
  • CMS and Geolocation services
  • Drive architecture improvements for resilience (failover, degradation, scaling patterns)

Production operations & incident leadership

  • Act as incident commander for major incidents and high-severity events
  • Lead 1-in-4 on-call rotation, covering:
  • Peak jackpot proactive monitoring
  • Own end-to-end incident lifecycle:
  • Ensure blameless post-mortems with clear remediation ownership

Observability & service insight

  • Define and evolve observability strategy using:
  • CloudWatch (AWS telemetry)
  • Quantum Metric (user behaviour insight)
  • Standardise:
  • Alerting quality and signal-to-noise ratio
  • Dashboards aligned to SLOs and customer impact
  • Drive correlation between technical signals and user experience
  • Lead automation of operational processes using Terraform and scripting
  • Improve deployment and release safety (CI/CD, progressive delivery patterns)
  • Key contributor to transition strategy for ECS → EKS (Kubernetes adoption)
  • Reduce operational toil through tooling, self-healing, observability and platform improvements
  • Empower Level-1 operational teams with safe, controlled access to the tools they need to operate autonomously

Capacity & performance engineering

  • Own capacity planning for:
  • Lead performance optimisation:
  • Throughput scaling
  • Cost efficiency (AWS utilisation and associated log costs, observability license consumption)
  • Own and prioritise the SRE backlog, balancing:
  • Reliability improvements
  • Automation opportunities to reduce/offload toil
  • Produce structured reporting covering:
  • SLO performance
  • Incident trends and MTTR
  • Platform health and risk areas
  • Provide clear updates to engineering leadership and business stakeholders
  • Embed SRE practices across engineering teams
  • Mentor engineers and SREs on:
  • Reliability engineering
  • Observability
  • Promote a culture of automation, measurement, and continuous improvement
  • Supporting with thought-leadership and developing long-term SRE roadmap for Digital Operations

What experiencewe’relooking for

Technical

  • Strong experience in cloud environments, ideally AWS (ECS, with exposure or experience in EKS/Kubernetes)
  • Hands-on experience with:
  • Terraform (Infrastructure as Code)
  • Observability stacks (Splunk, CloudWatch, Grafana)
  • Strong programming skills (Python, Go, or similar)

SRE practices

  • Observability strategies

Operational

  • Experience in on-call production environments
  • Demonstrated leadership during high-severity incidents
  • Experience migrating container platforms (ECS → EKS/Kubernetes)
  • Experience supporting high-scale consumer platforms
  • Familiarity with real-time analytics / customer experience tooling (e.g., Quantum Metric)
  • Experience in regulated or high-availability environments

About us

At Allwyn, we are dedicated to changing lives and growing the National Lottery responsibly, championing its positive impact on people, places, and the planet.

  • Innovation -We pride ourselves on it!We’reconstantly looking for new ways to excite our customers, bringing new products to market toenjoywhich is all supported by our responsible play values and making them accessible to all.
  • Giving back -Did you know that playing the lottery generates around £30m a week for charities andgood causesin the UK? Our aim is to have doubled this number by the end of the first 10-year license.
  • Sustainability -Our aim is to become a net zero national lottery. We have 2030 targets to decarbonise our operations and energy.We’vealready transitioned to renewable energy providers, made our London and Watford offices zero gas, and ensured our fleet consists of low-emission vehicles. In addition,we’reworking with our value chain partners to develop a net zero target date.
  • Empowering every voice- We believe in creating a culture where everyone feels they belong, can be themselves, has access to opportunities and can thrive for the benefit of good causes.Our diverse teams are working hard to make all parts of The National Lottery inclusive – whether people play a game in a store or online,because when everyone can play, everyonewins..

An inclusive reward offering with wellbeing at the centre

At Allwyn, inclusion is built into how we care for our people. Our benefits and policies support colleaguesand their familiesat every stage of life and career. By prioritising wellbeing and belonging, we create a workplace where everyone feels valued, rewarded, and empowered to succeed.Our people are more than colleagues -they’rewinners, driving positive change and making a real difference in communities.

  • CompanyBonusScheme
  • Matched pension contributions up to 8.5%
  • 26 days annual leave + 2 Life Days (and bank holidays)
  • Single Private Health Cover
  • Complimentary Private Medical
  • Income Protection
  • Flexible Benefits – EV Scheme,Money Coach, Will Writing,Mortgage Advice,Dental and Eye Care Schemes.
  • EnhancedFamily Leave (Maternity,Paternity, Adoption)
  • Employee AssistanceProgramme
  • Discounted Health Assessments
  • Matched Funding

We are a Disability Confident Leader which meanswe’vetaken proactive steps to ensure our workplace is accessible and inclusive for disabled and neurodivergent colleagues and candidates. As part of this we offer an interview to disabled applicants who meet the essential requirements of the job.

If you need anyassistanceor adjustments to this job description or in the application process, please contact a member of the talent team atcareers@allwyn.co.ukandwe’llbe happy to help.

#J-18808-Ljbffr…

Posted: July 31st, 2026