Senior Site Reliability Engineer (SRE) Engineer

Umanist Staffing
Other - Pune
View Company Profile / << Go Back

  • Job Type: Full time
  • 26 days ago

Job Description

Senior Site Reliability Engineer (SRE) Engineer

**Location:** Viman Nagar, Pune -- Work From Office

**Experience Overall(must have):** 8 Years

**CTC:** Up to ?25 LPA

**Notice Period:** Immediate Joiners Only within 15d or (if serving max 30days)

**Working Hours:** 3:00 PM -- 12:00 AM, Monday to Friday

**On-Call:** 24/7 Production Support -- On-Call Rotation Required

**Employment Type:** Full-Time

About the Role

We are looking for an experienced **Senior Site Reliability Engineer (SRE) / DevOps Engineer** to manage and improve the reliability, scalability, performance, security, and observability of mission-critical production environments.

The role requires strong hands-on expertise in **Cloud, Kubernetes, DevOps automation, Monitoring \& Observability, Incident Management, and SRE practices**. The ideal candidate should be comfortable handling production incidents while also driving long-term initiatives around reliability, automation, scalability, and reduction of operational toil.

Must-Have Skills \& Experience1. SRE \& Production Operations

* Relevant 7 years of relevant experience in **SRE / DevOps / Cloud Infrastructure / Production Engineering**.

* Hands-on experience with **24/7 production support and on-call operations**.

* Strong experience in **incident management, troubleshooting, RCA, and post-mortems**.

* Good understanding of **SLI, SLO, SLA, Error Budgets, MTTR, and reliability engineering**.

* Experience with **toil reduction, capacity planning, high availability, disaster recovery, and failover strategies**.

* Ability to improve system availability, performance, scalability, and operational reliability.

2. Cloud \& Infrastructure

* Strong hands-on experience with **Microsoft Azure, AWS, and/or GCP**.

* Strong understanding of cloud infrastructure, networking, IAM, storage, compute, and cloud-native services.

* Hands-on experience with at least one major cloud platform and good exposure to multi-cloud environments.

* Experience with:

* Azure: VMs, Networking, Storage, IAM, Azure Monitor, AKS

* AWS: EC2, S3, RDS, IAM, VPC, CloudWatch, EKS

* GCP: Compute Engine, Cloud Storage, IAM, VPC, GKE, Cloud Monitoring

3. Kubernetes \& Containerization

* Strong hands-on experience with **Kubernetes** and containerized workloads.

* Experience with **AKS / EKS / GKE** or equivalent Kubernetes environments.

* Hands-on experience with **Helm** deployments.

* Understanding of Kubernetes troubleshooting, scaling, networking, and workload management.

4. Infrastructure as Code \& DevOps

* Hands-on experience with **Terraform / Infrastructure as Code (IaC)**.

* Experience with Git-based workflows using **GitHub, GitLab, or Azure Repos**.

* Strong DevOps automation and CI/CD understanding.

* Strong scripting skills in **Python and/or Bash**.

5. Monitoring \& Observability

* Strong hands-on experience with **OpenTelemetry**.

* Experience with monitoring and observability tools such as:

* Prometheus

* Grafana

* Datadog

* Azure Monitor

* AWS CloudWatch

* GCP Cloud Monitoring

* Strong understanding of **metrics, logs, distributed tracing, and alerting**.

* Experience implementing monitoring based on **Golden Signals**:

* Latency

* Traffic

* Errors

* Saturation

* Ability to develop symptom-based, user-impact-focused alerting.

6. Linux \& Networking

* Strong knowledge of **Linux system administration**.

* Strong understanding of:

* DNS

* TCP/IP

* Load Balancing

* SSL/TLS

* Networking fundamentals

* Experience supporting highly available production environments.

7. Incident \& Reliability Engineering

* Ability to rapidly diagnose and resolve **high-severity production incidents**.

* Experience driving **MTTR reduction**.

* Strong debugging and analytical problem-solving skills.

* Ability to identify recurring issues and implement permanent corrective/preventive solutions.

Good-to-Have Skills

* Experience working across **Azure AWS GCP** in a multi-cloud environment.

* Knowledge of **Go (Golang)**.

* Experience with **OpenSearch / ELK Stack**.

* Experience supporting **AI/ML workloads** in production.

* Exposure to **Azure AI Services and Azure AI Foundry**.

* Experience supporting **RAG (Retrieval-Augmented Generation)** workloads.

* Experience designing infrastructure for AI/ML platforms.

* Experience building enterprise-wide **OpenTelemetry observability frameworks**.

* Strong understanding of distributed systems architecture.

* Exposure to advanced cloud-native architectures and reliability patterns.

* Experience with security, compliance, vulnerability remediation, secrets management, and network segmentation.

Key ResponsibilitiesProduction \& Incident Management

* Participate in the **24/7 on-call rotation**.

* Diagnose, mitigate, and resolve production incidents.

* Lead RCA and post-incident reviews.

* Implement corrective and preventive actions.

* Continuously improve MTTR and production stability.

Reliability Engineering

* Define and improve **SLIs, SLOs, SLAs, and Error Budgets**.

* Identify and eliminate operational toil.

* Conduct reliability and capacity reviews.

* Improve redundancy, failover, disaster recovery, and system resilience.

Cloud \& Infrastructure

* Manage and optimize cloud infrastructure across Azure, AWS, and/or GCP.

* Manage Kubernetes clusters and containerized applications.

* Implement and maintain Infrastructure as Code using Terraform.

* Support CI/CD and Git-based development workflows.

Observability \& Performance

* Build and improve monitoring, logging, metrics, and tracing.

* Implement **OpenTelemetry and distributed tracing**.

* Establish Golden Signals-based monitoring and alerting.

* Identify and resolve infrastructure and application performance bottlenecks.

Security

* Implement cloud security best practices around **IAM, network segmentation, and secrets management**.

* Support vulnerability remediation and compliance initiatives.

* Collaborate with Development, Security, and Infrastructure teams.

Ideal Candidate

We are looking for someone with:

* Strong **SRE mindset and production ownership**.

* Excellent troubleshooting and incident-management skills.

* Hands-on expertise in **Cloud Kubernetes Terraform Observability**.

* Strong understanding of **OpenTelemetry and Golden Signals**.

* Experience working in highly available, production-critical environments.

* Ability to remain calm and make effective decisions during critical incidents.

* Strong communication and cross-functional collaboration skills.

* Passion for **automation, scalability, reliability, and continuous improvement**.

Important Hiring Criteria

**Must be:**

* 7 years relevant experience

* Immediate joiner

* Willing to work from office in **Viman Nagar, Pune**

* Comfortable with **3:00 PM -- 12:00 AM shift**

* Comfortable with **24/7 on-call rotation**

* Strong hands-on SRE/DevOps experience

* Strong Cloud Kubernetes Observability experience

* Strong production incident management experience

**Good to have:**

* Multi-cloud: Azure AWS GCP

* OpenTelemetry

* AI/ML or RAG production workloads

* Azure AI / AI Foundry

* Go

* OpenSearch / ELK

* Distributed systems




Fast Track Upload