01Overview
Cloud Architect / DevOps Architect Resiliency & ObservabilityLocation : Remote Anywhere in IndiaExperience : 8+ YearsOpen Positions : 3Joining : Immediate Joiners / Candidates who can join within 15 daysWork Mode : RemoteAbout the Role :We are looking for an experienced Cloud Architect / DevOps Architect with strong hands-on expertise in AWS Cloud, DevOps, Cloud Resiliency, Disaster Recovery and Observability. The ideal candidate should have experience architecting highly available and resilient enterprise cloud platforms, while also working hands-on with Kubernetes, Terraform, CI/CD, monitoring and SRE practices. This is a hands-on technical architecture role involving solution design, implementation, troubleshooting and production reliability.Key Responsibilities :- Design and implement scalable, secure and highly available AWS cloud architectures- Architect cloud environments across compute, networking, storage, security, monitoring and automation- Design HA, DR, failover and replication strategies for critical applications- Build fault-tolerant architectures with automated recovery mechanisms- Design and improve observability across infrastructure, applications and Kubernetes- Work with Prometheus, Grafana, ELK/OpenSearch, Splunk, Datadog, Dynatrace, New Relic or OpenTelemetry- Design and enhance CI/CD and DevOps automation- Provision infrastructure using Terraform, CloudFormation, Ansible or similar tools- Work with Docker and Kubernetes- Apply SRE and reliability engineering practices- Participate in incident management, troubleshooting and Root Cause Analysis (RCA)- Support chaos testing, failure simulations and recovery validation- Ensure cloud environments meet enterprise requirements for security, networking, governance and complianceMust-Have Skills :- 8+ years of experience in Cloud Architecture / DevOps / SRE / Platform Engineering- Strong hands-on AWS Cloud Architecture experience- Strong understanding of High Availability, Resiliency and Reliability- Hands-on experience with HA/DR, failover, replication and disaster recovery- Experience with observability tools such as Prometheus, Grafana, ELK/OpenSearch, Splunk, Datadog, Dynatrace, New Relic or OpenTelemetry- Strong DevOps & CI/CD automation experience- Experience with Terraform / CloudFormation / Ansible- Hands-on experience with Docker & Kubernetes- Understanding of SRE, incident management and RCA- Strong knowledge of cloud networking, security and enterprise production environmentsWhat We Are Specifically Looking For :We are particularly interested in candidates with recent hands-on experience in Cloud Resiliency and Observability. The ideal candidate should be able to :- Architect highly available AWS environments- Design HA/DR and disaster recovery strategies- Define failover, replication and recovery mechanisms- Build strong monitoring and observability frameworks- Identify potential infrastructure failure scenarios- Design automated recovery mechanisms- Work with Kubernetes and cloud-native environments- Apply SRE and reliability engineering practices- Troubleshoot production issues and perform RCAImportant :Candidates whose experience is primarily focused on cloud migration, basic DevOps pipelines or CI/CD management without strong resiliency and observability experience may not be suitable.Joining Details :Immediate Joiners / Candidates who can join within 15 days preferred. (ref:hirist.tech) .