Site Reliability Engineer

AppLab Systems, Inc
US - California - Sunnyvale
View Company Profile / << Go Back

  • Job Type: Full time
  • Just Posted

Job Description

Investigate and resolve performance and reliability issues across application, infrastructure, database, Kubernetes, and Linux layers, and recommend evidence-based improvements. Design and assess scalable, highly available distributed systems, plan capacity, improve observability, and collaborate with engineering teams and business stakeholders.

Requirements: Requires hands-on experience with distributed systems, performance engineering, Java troubleshooting, Kubernetes, Linux tuning, PostgreSQL optimization, cloud platforms, scripting, deployment automation, and monitoring. Candidates must also have practical LLM experience, strong analytical and communication skills, and the ability to take ownership and work collaboratively.

Key Skills: Distributed Systems, Performance Engineering, Java, Kubernetes, Linux Internals And Tuning, PostgreSQL, Azure, Python, Bash, PowerShell, High Availability, Capacity Planning, Monitoring And Observability, Deployment Pipelines, Large Language Models




Fast Track Upload