Sign up to access all features of our service
  • Job search
  • Favorites
  • Create a CV
    New
  • Subscriptions

Senior Site Reliability Engineer

$1000 per year
Full-time

Camunda

Responsibilities

  • Design and maintain Camunda’s Kubernetes-based multi-cloud platform for availability, scalability, and fault tolerance.
  • Establish infrastructure configuration best practices and network services used by engineering teams.
  • Implement and improve monitoring, alerting, and observability across systems.
  • Participate in on-call rotations, diagnose production incidents, perform root cause analysis, and create runbooks.
  • Build automation to reduce repetitive work and improve system reliability and maintainability.
  • Collaborate with product engineering, product management, and support to define and deliver features.
  • Mentor less experienced engineers on complex infrastructure challenges.

Requirements

  • Deep hands-on experience building, deploying, and maintaining production Kubernetes clusters, including workloads, networking, and storage.
  • Strong infrastructure-as-code expertise with Terraform or similar tools, including versioning, testing, and safe deployment practices.
  • Demonstrated experience with monitoring and observability tools such as Prometheus and Grafana.
  • Strong third-level support and incident response skills, including complex troubleshooting, stakeholder communication, and root cause analysis.
  • Experience automating manual work and building reliable, maintainable systems.
  • Ability to use AI tools responsibly for research, code review, documentation, test generation, and automation while protecting confidential data and retaining human accountability.
  • Experience with AWS/EKS, Google Cloud/GKE, or similar managed Kubernetes services is preferred.
  • Experience with ArgoCD or GitOps workflows is preferred.
  • Proficiency in Python, Go, or similar programming languages is preferred.
  • Experience defining service-level objectives and configuring effective alerting frameworks is preferred.

Benefits

  • Fully remote and global work arrangement with flexible working options.
  • Home office budget, co-working space support, and flexible time off.
  • Annual Kickoff, team offsites, Camundi Connection Budgets, meetups, and local gatherings.
  • Locally tailored healthcare, Modern Health mental wellbeing support, and the Live Well Lifestyle Spending Account.
  • Retirement and pension plans, often with company contributions, plus life and disability insurance where relevant.
  • Up to $/€/£1,000 per year for self-directed learning, including courses, certifications, and books.
  • Equity through the Virtual Stock Option Plan where applicable.
Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer in Remote vacancy
  •  ...customers. Cohere is a team of researchers, engineers, designers, and more, who are all...  ...technologies to improve performance and reliability YOU MAY BE A GOOD FIT IF: – You have...  ...email alias. If jobs are viewed on other sites then please verify these through our official... 

    Cohere

    Remote
    1 day ago
  •  ...share the parts that generalize with other teams. – Raise the bar on your team through code and design review, and bring other engineers up on eval practice. – Some of the technical questions in this area are still open. You will help answer them. – Participate... 

    Mapbox

    Remote
    1 day ago
  •  ...the Role We’re looking for a FinOps Engineer to help drive cost optimization and cloud...  ...influencing product decisions when cost to reliability tradeoffs hit. Strong communication...  ...health are important to us. – Annual Off-Sites Once a year, the entire company... 

    Supabase

    Remote
    1 day ago
  •  ...intelligent agents ubiquitous. We build the foundation for agent engineering in the real world, helping developers move from prototypes to...  ...team, working directly with enterprise customers to build reliable, production agents. You’ll translate vague enterprise workflows... 

    LangChain

    Remote
    1 day ago
  •  ...intelligent agents ubiquitous. We build the foundation for agent engineering in the real world, helping developers move from prototypes to...  ...scalable, secure infrastructure deployments and building reliable, well-evaluated agent applications that solve real business problems... 

    LangChain

    Remote
    1 day ago
  • $10 per hour

     ...skills Ability to follow detailed evaluation guidelines consistently Good understanding of spoken language, audio quality, and context Reliable equipment — high-quality headphones are required Ability to complete the evaluation test within 48 hours of receiving the... 

    Rws

    Remote
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!