Job Details

Job Title
Senior Site Reliability Engineer
Location
Guwahati
Job Description
 

Senior Site Reliability Engineer

10463

Guwahati

Virtual

Basis JD, Azure exp is mandatory,

Refer Sanjay SRE JD

E4

 
Job Responsibilities (JR) – Key Areas
1. Reliability Engineering & Production Stability
 Ensure high availability and performance of critical systems and services
 Define, monitor, and manage SLOs / SLIs / SLAs
 Lead incident management, major incident review, and RCA
 Improve MTTR, system resilience, and production stability
2. Infrastructure & Cloud Engineering
 Design and manage cloud infrastructure (AWS / Azure / GCP)
 Implement Infrastructure as Code (IaC) for scalable provisioning
 Manage Kubernetes clusters and container-based deployments
 Ensure HA, DR, and BCP readiness across environments
3. Automation & DevOps Enablement
 Implement and manage CI/CD pipelines for faster releases
 Automate build, deployment, and operational workflows
 Reduce manual efforts through scripting and engineering solutions
 Drive DevOps adoption across application teams
4. Monitoring, Observability & Alerting
 Build and maintain monitoring, logging, and alerting frameworks
 Implement observability using logs, metrics, and traces
 Proactively identify and resolve system anomalies
5. Incident, Problem & Change Management
 Handle P1/P2 incidents and support L2/L3 production escalations
 Perform root cause analysis and implement preventive solutions
 Collaborate with ITSM teams (ServiceNow, etc.)
6. Performance & Capacity Management
 Perform capacity planning and system performance tuning
 Optimise infrastructure utilisation and costs
 Ensure system readiness for peak loads and traffic spikes
7. Engineering Collaboration & Culture
 Work closely with development, QA, Security, and Architecture teams
 Promote SRE practices, reliability culture, and automation mindset
 Mentor junior engineers and contribute to knowledge sharing
✅ Key Skills Required
Technical Skills
 Strong experience in Linux, networking, and distributed systems
 Cloud expertise: AWS / Azure / GCP
 Containerisation: Docker, Kubernetes
 Infrastructure as Code: Terraform / Ansible
 CI/CD tools: Jenkins, GitLab CI, GitHub Actions
Observability & Support
 Monitoring tools: Prometheus, Grafana, ELK, Splunk
 Incident management, RCA, and troubleshooting
 Understanding of SRE concepts (SLO, SLA, error budgets)
Programming / Scripting
 Python / Shell / Go (preferred)
Platform & Tools Exposure
 Git, Jira, Confluence
 ServiceNow (Incident / Change / Problem)
 Exposure to microservices architecture
✅ Educational Qualifications
 B.Tech / BE in Computer Science / IT or related discipline
 Certifications (Preferred): AWS / Azure / Kubernetes
✅ Experience Requirement
 5–10+ years in DevOps / SRE / Cloud engineering
 Experience in enterprise production environments (Banking preferred)
✅ Key Competencies
 Problem-solving and troubleshooting skills
 Ability to manage high-pressure production incidents
 Strong stakeholder communication
 Continuous learning and automation mindset
✅ Major Stakeholders
Internal:
 Application Development Teams
 Platform / Cloud Engineering
 Information Security / Risk / Compliance
 ITSM / Production Support Teams
External:
 Cloud Vendors (AWS, Azure, GCP)
 Tooling / Technology Providers
✅ Success Metrics (KPIs)
 System availability and uptime
 MTTR / MTTD improvement
 Deployment frequency and success rate
 Reduction in production incidents
 Automation coverage

Apply to Job