10,000+ Active Jobs
|
500+ Hiring Companies
|
100% Verified Jobs
India
HiringGo Logo
Companies
Exclusive Jobs
Jobs Login
Homeโ€บCompaniesโ€บDrivetrainโ€บSite Reliability Engineer - SRE
D

Site Reliability Engineer - SRE

๐Ÿ“LOCATIONAll India
๐Ÿ“ˆEXPERIENCE0 to 4 Yrs
๐Ÿ•˜TYPEFull time
๐ŸฅIndustryIT Services & Consulting
๐Ÿ—“POSTED5 Aug 2026

01Overview

Drivetrain is on a mission to empower businesses to make better decisions. Our financial planning & decision-making platform helps companies scale and achieve their targets predictably. Drivetrain is a remote-first company headquartered in the San Francisco Bay Area. Founded in 2021 by a couple of ex-Googlers, Drivetrain is a fast-growing company on a trajectory for success with backing from leading venture capital firms. Drivetrain provides a great culture for its employees to thrive in and be happy. Remote-friendly: Drivetrain brings together the best and the brightest, no matter where they are and provides them a great degree of autonomy. We trust our people. Open & transparent: We know that when our creators have access to all the information they need, their best work will emerge. Idea-friendly: We provide an environment to explore new ideas, to take risks, to make mistakes, and to learn, so you can succeed. Anyone in the company can come up with great ideas and become a catalyst for positive change. We let the best ideas win. Customer-centric: We follow a product-led growth strategy, continuously learning from our customers and collaborating to build the amazing software that Drivetrain is. As a Senior Site Reliability Engineer at Drivetrain, you will be a cornerstone of our engineering organization, ensuring our fast-growing SaaS platform remains highly available, performant, and secure. At this stage of our growth, scaling infrastructure efficiently while maintaining the rigorous security and reliability standards required for financial data is paramount. You will take ownership of our multi-cloud infrastructure, drive automation, champion observability, and collaborate closely with development teams to build a culture of reliability from code commit to production. Key Responsibilities Cloud Infrastructure & Orchestration Multi-Cloud Management: Architect, manage, and continuously optimize highly available cloud infrastructure across both AWS and GCP. Balance workload demands to ensure maximum cost-efficiency, scalability, and strict security compliance across both platforms. Advanced Kubernetes Orchestration: Lead the design, deployment, and management of scalable Kubernetes clusters. Utilize configuration management tools like Kustomize to enforce standardized, repeatable, and automated deployment configurations across all environments. Service Mesh & Security Integration: Implement and maintain service mesh technologies (e.g., Istio, Linkerd) to secure, control, and observe service-to-service communication. Drive container security best practices, including image scanning, runtime protection, and strict RBAC enforcement. CI/CD & Automation Pipeline Engineering: Architect, maintain, and optimize robust CI/CD pipelines using Git and Jenkins. Focus on reducing deployment friction, accelerating release velocity, and enforcing automated testing and security gates. Infrastructure as Code (IaC): Treat infrastructure as software. Write, review, and maintain Terraform modules to provision and manage cloud resources predictably and safely. Operational Automation: Aggressively reduce operational toil. Develop robust Python scripts and tooling to automate routine maintenance, data backups, scaling operations, and system recovery processes. Observability & Reliability Comprehensive Monitoring: Design and enhance our observability stack to provide deep, real-time insights into system health. Manage and scale tools including Prometheus, Grafana, ELK/EFK stack, AWS CloudWatch, and GCP Operations Suite. Reliability Engineering: Spearhead reliability initiatives critical to a scaling SaaS platform. Drive rigorous capacity planning exercises to stay ahead of growth. Incident Management & SLOs: Own the incident response lifecycle. Facilitate blameless postmortems to extract actionable learnings. Define, track, and enforce SLIs, SLOs, and SLAs, ensuring the platform consistently meets its reliability guarantees. Collaboration & Leadership DevOps Culture: Act as an embedded reliability advocate. Collaborate closely with software engineers early in the development lifecycle to ensure applications are designed for deployability, scalability, and resilience. Continuous Improvement: Proactively identify system bot .

02What you'll need

Experience
0 to 4 Yrs
Employment Type
Full time
Programming languages
AWSGCPKubernetesJenkinsPythonCloud InfrastructureIstioLinkerdCICDTerraform

03About DRIVETRAIN

IT Services & ConsultingIndustry
Full timeEmployment Type
All IndiaLocation
Not Disclosed ยท salary hidden by employer
0 to 4 Yrs ยท All India
Applications are reviewed directly by the hiring team.
Role Snapshot
Work ModeNot specified
Visa SponsorshipNot specified
RelocationNot specified
Job TypeFull time
D
DRIVETRAIN
IT Services & Consulting
View all DRIVETRAIN jobs โ†’
Share