Site Reliability Engineer
Guarda esta oferta y sigue tu búsqueda
Crea una cuenta gratis para guardar empleos, crear alertas y volver a esta oferta desde tu panel.
Al continuar, aceptas nuestros Términos & Política de Privacidad.
The Tyk API Management platform is helping to drive the connected world and power new products and services. Were changing the way that organisations connect any number of their systems and services.Whether internal, external, public or highly encrypted systems, Tyk helps businesses drive value across the retail, finance, telecoms, healthcare, or media industries (to name just a few)
If youve banked online, used an app to check the news, or perhaps even driven a connected car, APIs, and by extension, Tyk, make that possible. Founded in 2015 with offices in London UK, London Ontario, Atlanta and Singapore, we have many thousands of users of our B2B platform across the globe. Brands using Tyk range from Lotte, Bell, T Mobile, to RBS, Capital One and Vinci. We have a varied user base hailing from every continent even Antarctica.
Our Mission
Tyk is on a mission to connect every system in the world. Weve started by building an API Management platform.
Total flexibility, default remote, radical responsibility
We offer unlimited paid holidays and remote working from anywhere in the world , for everyone, Why? Tyk was founded on the principle of offering flexibility and autonomy to our employees, we believe this allows our employees to achieve their best results. It also means we can build the best possible team, location and working hours are no barrier.
If this sounds like an environment that you believe could work for you then read on to find out more.
The role:
Tyk Cloud is our managed API management platform, running on multi-region Kubernetes at scale for customers around the world.
Were looking for an SRE whos as comfortable in the code as in the infrastructure. Youll spend most of your time improving and automating the platform, and when youre on call, youll handle incidents independently. Youll join a small, centralised SRE team that works closely with our product teams.
You dont need to have done everything below. We care most about how you reason through problems and how quickly you learn. Youll have three to four months to get up to speed, shadowing first, before you go on call.
What youll do:
Most of your time
- Deliver the teams planned work each quarter, such as optimising the platform, building self-serve tooling for other teams, and rearchitecting parts of the platform as it grows.
- Help expand Tyk Cloud across regions and clouds, and bring down what it costs to run.
- Automate operations in Go, including building and maintaining our custom Kubernetes operators.
- Run the platforms services and databases, including MongoDB and Redis.
- Improve our observability: find the metrics that matter, and build the dashboards and alerts to act on them.
- Keep runbooks and documentation current, and support security work such as SOC 2 audits.
When youre on call
You'll be doing one week in three initially (one in four as we grow), Monday-Friday, on a 12-hour shift with secondary backup support. Rotas: 14:000:00 UTC
- Be first line for platform alerts and incidents: restore service, escalate or help fix product bugs, and lead post-incident reviews.
- Act as second line for our Customer Success team, on requests that come directly from customers.
- Handle ad hoc requests from other teams across the organisation regarding Tyk Cloud.
Requirements
What youll need:
- 3+ years in SRE, platform or infrastructure roles, across more than one company or production platform.
- Experience owning on-call and leading incidents yourself.
- Hands-on experience running production Kubernetes at scale, ideally EKS: operating, upgrading and debugging large, multi-tenant clusters.
- Experience designing and operating infrastructure on AWS, with Terraform or similar.
- The ability to write, test and ship Go tooling or services.
- Experience with Prometheus and Grafana, and with logging systems.
- Solid Linux and networking fundamentals (DNS, TCP/IP, HTTP, TLS, load balancing).
- Clear communication across time zones and teams.
Our stack
EKS, Terraform/Terragrunt, Helm, GitHub Actions, Argo CD, MongoDB, Redis, Prometheus and Grafana.
Heres why you should join us:
- Everyone has unlimited paid holiday.
- We have total flexibility in hours, as we believe creativity flows better when our people are given freedom to decide when they are most productive. Everyone is unique after all.
- Employee share scheme
- Generous maternity and paternity leave
- Company retreats
We all share the same vision we value authenticity, respect, responsibility, independence, honesty, diversity and inclusion and most importantly treating others how you wish to be treated. We look for like-minded people who bring their personal