01Overview
About the job
24×7 Platform Monitoring
Monitor platform and application dashboards, alerts and operational mailboxes.
Identify availability, infrastructure, application and service degradation events.
Acknowledge alerts promptly and determine initial severity based on established procedures.
Maintain accurate shift handover and operational records.
First-Level Incident Troubleshooting
Perform basic network connectivity checks including ping, DNS resolution and endpoint connectivity.
Check OpenShift/Kubernetes resource status including nodes, pods, deployments and services.
Review basic application and platform logs to identify common failure conditions.
Check infrastructure and service health using approved dashboards and operational tools.
Collect relevant diagnostic information before escalation.
Incident Recovery
Execute documented recovery procedures and operational runbooks.
Restart or redeploy affected workloads using approved GitOps processes.
Verify service recovery through dashboards, health checks and application endpoints.
Escalate when recovery procedures are unsuccessful or when an incident falls outside the approved operating scope.
Incident Coordination & Escalation
Create and maintain incident tickets with accurate timestamps, symptoms, actions and observations.
Engage the appropriate application, platform, infrastructure or network teams based on established escalation procedures.
Provide clear status updates during active incidents.
Support incident bridges by providing operational information and executing actions requested by L2/L3 engineers.
Ensure effective handover of unresolved incidents between shifts.
Operational Procedures
Follow established Standard Operating Procedures (SOPs), runbooks and change-management processes.
Document newly encountered symptoms and successful troubleshooting steps.
Highlight recurring alerts or operational problems to senior platform engineers.
Participate in operational drills and recovery exercises.
Bachelor's degree in Computer Science, Information Technology, Engineering or related discipline.
Openshift experience is mandatory
5 years of IT operations, infrastructure, cloud or application support experience.
Basic Linux command-line knowledge.
Basic networking knowledge including IP addressing, DNS, ports and connectivity troubleshooting.
Basic understanding of containers and Kubernetes concepts.
Ability to follow technical procedures accurately.
Good written and verbal communication skills.
Willingness and ability to work in a 24×7 shift environment.
Mandatory skills : Platform monitoring,Incident trouble shooting,Incident recovery,openshift,kubernet
Excellent communication, stakeholder management
Job Stability - min 2 years in an organization
Notice Period - Immediate to 45days