Umanist Staffing
Other - Pune
View Company Profile /
<< Go Back
Senior Site Reliability Engineer (SRE) Engineer
**Location:** Viman Nagar, Pune -- Work From Office
**Experience Overall(must have):** 8 Years
**CTC:** Up to ?25 LPA
**Notice Period:** Immediate Joiners Only within 15d or (if serving max 30days)
**Working Hours:** 3:00 PM -- 12:00 AM, Monday to Friday
**On-Call:** 24/7 Production Support -- On-Call Rotation Required
**Employment Type:** Full-Time
About the Role
We are looking for an experienced **Senior Site Reliability Engineer (SRE) / DevOps Engineer** to manage and improve the reliability, scalability, performance, security, and observability of mission-critical production environments.
The role requires strong hands-on expertise in **Cloud, Kubernetes, DevOps automation, Monitoring \& Observability, Incident Management, and SRE practices**. The ideal candidate should be comfortable handling production incidents while also driving long-term initiatives around reliability, automation, scalability, and reduction of operational toil.
Must-Have Skills \& Experience1. SRE \& Production Operations
* Relevant 7 years of relevant experience in **SRE / DevOps / Cloud Infrastructure / Production Engineering**.
* Hands-on experience with **24/7 production support and on-call operations**.
* Strong experience in **incident management, troubleshooting, RCA, and post-mortems**.
* Good understanding of **SLI, SLO, SLA, Error Budgets, MTTR, and reliability engineering**.
* Experience with **toil reduction, capacity planning, high availability, disaster recovery, and failover strategies**.
* Ability to improve system availability, performance, scalability, and operational reliability.
2. Cloud \& Infrastructure
* Strong hands-on experience with **Microsoft Azure, AWS, and/or GCP**.
* Strong understanding of cloud infrastructure, networking, IAM, storage, compute, and cloud-native services.
* Hands-on experience with at least one major cloud platform and good exposure to multi-cloud environments.
* Experience with:
* Azure: VMs, Networking, Storage, IAM, Azure Monitor, AKS
* AWS: EC2, S3, RDS, IAM, VPC, CloudWatch, EKS
* GCP: Compute Engine, Cloud Storage, IAM, VPC, GKE, Cloud Monitoring
3. Kubernetes \& Containerization
* Strong hands-on experience with **Kubernetes** and containerized workloads.
* Experience with **AKS / EKS / GKE** or equivalent Kubernetes environments.
* Hands-on experience with **Helm** deployments.
* Understanding of Kubernetes troubleshooting, scaling, networking, and workload management.
4. Infrastructure as Code \& DevOps
* Hands-on experience with **Terraform / Infrastructure as Code (IaC)**.
* Experience with Git-based workflows using **GitHub, GitLab, or Azure Repos**.
* Strong DevOps automation and CI/CD understanding.
* Strong scripting skills in **Python and/or Bash**.
5. Monitoring \& Observability
* Strong hands-on experience with **OpenTelemetry**.
* Experience with monitoring and observability tools such as:
* Prometheus
* Grafana
* Datadog
* Azure Monitor
* AWS CloudWatch
* GCP Cloud Monitoring
* Strong understanding of **metrics, logs, distributed tracing, and alerting**.
* Experience implementing monitoring based on **Golden Signals**:
* Latency
* Traffic
* Errors
* Saturation
* Ability to develop symptom-based, user-impact-focused alerting.
6. Linux \& Networking
* Strong knowledge of **Linux system administration**.
* Strong understanding of:
* DNS
* TCP/IP
* Load Balancing
* SSL/TLS
* Networking fundamentals
* Experience supporting highly available production environments.
7. Incident \& Reliability Engineering
* Ability to rapidly diagnose and resolve **high-severity production incidents**.
* Experience driving **MTTR reduction**.
* Strong debugging and analytical problem-solving skills.
* Ability to identify recurring issues and implement permanent corrective/preventive solutions.
Good-to-Have Skills
* Experience working across **Azure AWS GCP** in a multi-cloud environment.
* Knowledge of **Go (Golang)**.
* Experience with **OpenSearch / ELK Stack**.
* Experience supporting **AI/ML workloads** in production.
* Exposure to **Azure AI Services and Azure AI Foundry**.
* Experience supporting **RAG (Retrieval-Augmented Generation)** workloads.
* Experience designing infrastructure for AI/ML platforms.
* Experience building enterprise-wide **OpenTelemetry observability frameworks**.
* Strong understanding of distributed systems architecture.
* Exposure to advanced cloud-native architectures and reliability patterns.
* Experience with security, compliance, vulnerability remediation, secrets management, and network segmentation.
Key ResponsibilitiesProduction \& Incident Management
* Participate in the **24/7 on-call rotation**.
* Diagnose, mitigate, and resolve production incidents.
* Lead RCA and post-incident reviews.
* Implement corrective and preventive actions.
* Continuously improve MTTR and production stability.
Reliability Engineering
* Define and improve **SLIs, SLOs, SLAs, and Error Budgets**.
* Identify and eliminate operational toil.
* Conduct reliability and capacity reviews.
* Improve redundancy, failover, disaster recovery, and system resilience.
Cloud \& Infrastructure
* Manage and optimize cloud infrastructure across Azure, AWS, and/or GCP.
* Manage Kubernetes clusters and containerized applications.
* Implement and maintain Infrastructure as Code using Terraform.
* Support CI/CD and Git-based development workflows.
Observability \& Performance
* Build and improve monitoring, logging, metrics, and tracing.
* Implement **OpenTelemetry and distributed tracing**.
* Establish Golden Signals-based monitoring and alerting.
* Identify and resolve infrastructure and application performance bottlenecks.
Security
* Implement cloud security best practices around **IAM, network segmentation, and secrets management**.
* Support vulnerability remediation and compliance initiatives.
* Collaborate with Development, Security, and Infrastructure teams.
Ideal Candidate
We are looking for someone with:
* Strong **SRE mindset and production ownership**.
* Excellent troubleshooting and incident-management skills.
* Hands-on expertise in **Cloud Kubernetes Terraform Observability**.
* Strong understanding of **OpenTelemetry and Golden Signals**.
* Experience working in highly available, production-critical environments.
* Ability to remain calm and make effective decisions during critical incidents.
* Strong communication and cross-functional collaboration skills.
* Passion for **automation, scalability, reliability, and continuous improvement**.
Important Hiring Criteria
**Must be:**
* 7 years relevant experience
* Immediate joiner
* Willing to work from office in **Viman Nagar, Pune**
* Comfortable with **3:00 PM -- 12:00 AM shift**
* Comfortable with **24/7 on-call rotation**
* Strong hands-on SRE/DevOps experience
* Strong Cloud Kubernetes Observability experience
* Strong production incident management experience
**Good to have:**
* Multi-cloud: Azure AWS GCP
* OpenTelemetry
* AI/ML or RAG production workloads
* Azure AI / AI Foundry
* Go
* OpenSearch / ELK
* Distributed systems
© 2026 engineeringjobs.net, Inc. All Rights Reserved.
Terms of Service | Privacy
Powered by JOBBEX