S
Site Reliability Engineer - Cloud Operations
In this role, you will:
• Migrate and modernize production applications on Kubernetes,
• Integrate third-party software into our production platforms and make it fit our operational standards,
• Work alongside Software and IT Engineers to improve reliability, performance and operational readiness,
• Design and operate applications on our service mesh platform,
• Integrate safe deployment patterns such as canary releases and progressive rollouts,
• Define SLOs, SLIs and useful operational KPIs, then use them to drive improvements,
• Improve observability across metrics, logs and traces so problems are easier to spot and understand,
• Explore and integrate AI tools that can help with troubleshooting, incident analysis and remediation,
• Test how systems behave under load, during failures and when dependencies disappear,
• Automate repetitive operational work whenever it makes sense,
• Provide Level-3 support and participate in the on-call rotation.
• At least 3 years of experience in SRE, DevOps, Platform Engineering or a similar production-focused role,
• Solid hands-on experience running production workloads on Kubernetes, OpenShift, EKS or a similar Kubernetes platform,
• Good knowledge of Helm and how to package, configure and maintain applications with it,
• Experience working with service mesh technologies such as Istio or Linkerd, or strong Kubernetes networking experience,
• A good understanding of service-to-service networking, traffic routing, mTLS and TLS,
• Experience with GitOps and modern deployment strategies such as canary or progressive delivery,
• A practical understanding of SRE concepts such as SLIs, SLOs and error budgets,
• Experience with observability and tracing tooling such as Prometheus, Grafana, Elastic Stack or OpenTelemetry,
• Strong Linux and networking fundamentals, including TCP/IP, DNS and load balancing,
• Comfortable troubleshooting JVM-based applications in production and able to investigate issues related to heap usage, garbage collection or JVM configuration,
• Comfortable automating things with Python, Go, Bash or another programming language,
• Experience or strong interest in applying AI to observability, incident response or operational automation,
• Experience or a strong interest in operating applications that depend on GPU resources or other AI infrastructure,
• Familiarity with Infrastructure as Code tools such as Terraform, Ansible or Puppet.
Nice-to-Haves
• Experience with Argo CD, Argo Rollouts or Argo Workflows,
• Deeper experience with Istio, Linkerd or Envoy-based service mesh platforms,
• Experience designing or operating Kubernetes platforms at scale,
• Experience running Java or Spring Boot applications in production,
• Hands-on experience tuning JVM applications for performance or low-latency workloads,
• Experience integrating applications with self-hosted AI platforms such as vLLM,
• Experience troubleshooting AI infrastructure integrations, including model access, GPU availability and NVIDIA MIG configurations.
• Knowledge of Cilium, eBPF or other modern Kubernetes networking technologies,
• Experience with public cloud or large private cloud environments,
• CKAD, CKA, CKS or equivalent hands-on Kubernetes experience,
• A homelab, self-hosted services or side projects where you get to experiment, break things and build them again.
Who You Are
• You like understanding why systems behave the way they do, especially when something goes wrong,
• You automate repetitive work instead of accepting it as part of the job,
• You’re comfortable working across development, infrastructure and operations teams,
• You don’t mind getting deep into software you didn’t build yourself,
• You’re curious about AI and where it can genuinely improve day-to-day operations,
• You are fluent in English and have good conversational French,
• You enjoy keeping up with cloud-native technologies and trying new approaches when they solve a real problem.
 
Please note that Swissquote never requests sensitive personal information or payment of any kind during the recruitment process. Any such request is fraudulent.
SQ2