01Overview
Qualification
B.E./B.Tech in Computer Science / IT / Electronics or equivalent.3+ years of experience in Server/System Administration, preferably in AI/GPU/HPC infrastructure.
Key Skills
Strong knowledge of Linux/Windows Server Administration.Hands-on experience in server hardware troubleshooting and replacement.Knowledge of GPU servers, NVIDIA/AMD GPUs, GPU drivers and AI workloads.Knowledge of storage, RAID, NVMe, networking and server components.Experience in OS installation, upgrades, patching and driver/software installation.Ability to monitor CPU, memory, storage, network and GPU performance.
Good troubleshooting and log-analysis skillsKnowledge of BMC/IPMI, BIOS/UEFI and hardware health monitoring.Basic knowledge of HPC/AI environments, CUDA, Slurm/Kubernetes will be an advantage.
Key Responsibilities
Provide onsite operational and technical support for AI/GPU server infrastructure.Troubleshoot hardware, OS, driver, application and GPU-related issues.Perform hardware replacement and preventive maintenance.Install/upgrade OS, drivers and software applications.Monitor system health, performance and AI/GPU workloads.Collect logs and performance statistics for troubleshooting and reporting.Maintain daily health-check, incident and activity reports.Coordinate with OEM/L3 technical experts through email and telephone for unresolved issues.Ensure timely escalation and closure of critical incidents.Support RCA and corrective actions for recurring/critical issues.
Desired Candidate Profile
Strong analytical and problem-solving skills.Good communication and customer-handling skills.Ability to work independently at customer locations.Willing to work in 24x7/shift/on-call support, as required.Experience in data center, HPC, AI/GPU or enterprise server environments preferred. .