Swissquote Bank SA

Site Reliability Engineer - Cloud Operations

📍 1196 Gland, 1196 Gland

Role and responsibilities

Migrate and modernize production applications on Kubernetes, Integrate third-party software into our production platforms and make it fit our operational standards, Work alongside Software and IT Engineers to improve reliability, performance and operational readiness, Design and operate applications on our service mesh platform, Integrate safe deployment patterns such as canary releases and progressive rollouts, Define SLOs, SLIs and useful operational KPIs, then use them to drive improvements, Improve observability across metrics, logs and traces so problems are easier to spot and understand, Explore and integrate AI tools that can help with troubleshooting, incident analysis and remediation, Test how systems behave under load, during failures and when dependencies disappear, Automate repetitive operational work whenever it makes sense, Provide Level-3 support and participate in the on-call rotation.

Team / description

At Swissquote, we’re all in. All in to shake things up. All in to build the bank people actually want to use. All in to make finance less boring — and a lot more powerful. We’re Switzerland’s leading digital bank — 1,400+ people across Europe, the Middle East and Asia, building real financial solutions for over a million clients worldwide. From trading and investing to everyday banking, we cover the full picture. We move fast, but we build things we’re genuinely proud of. Have a look behind the scenes by checking Humans of Swissquote on Instagram. Growing fast creates room. Room to try things, own things, and grow at a pace most places can’t offer. Whether you like the spotlight or prefer to just put your head down and do great work, there’s space for both here. We’ve ditched the dress code — but never the chance to celebrate. Big win or small, we make it count. The kind of place where the atmosphere takes care of itself. As an equal opportunity employer, we welcome candidates from all backgrounds, experiences and perspectives to join our team and contribute to our shared success.

Qualifications and Skills

  • At least 3 years of experience in SRE, DevOps, Platform Engineering or a similar production-focused role

  • Solid hands-on experience running production workloads on Kubernetes, OpenShift, EKS or a similar Kubernetes platform

  • Good knowledge of Helm and how to package, configure and maintain applications with it

  • Experience working with service mesh technologies such as Istio or Linkerd, or strong Kubernetes networking experience

  • A good understanding of service-to-service networking, traffic routing, mTLS and TLS

  • Experience with GitOps and modern deployment strategies such as canary or progressive delivery

  • A practical understanding of SRE concepts such as SLIs, SLOs and error budgets

  • Experience with observability and tracing tooling such as Prometheus, Grafana, Elastic Stack or OpenTelemetry

  • Strong Linux and networking fundamentals, including TCP/IP, DNS and load balancing

  • Comfortable troubleshooting JVM-based applications in production and able to investigate issues related to heap usage, garbage collection or JVM configuration

  • Comfortable automating things with Python, Go, Bash or another programming language

  • Experience or strong interest in applying AI to observability, incident response or operational automation

  • Experience or a strong interest in operating applications that depend on GPU resources or other AI infrastructure

  • Familiarity with Infrastructure as Code tools such as Terraform, Ansible or Puppet