01Overview
Job Overview
Role: Senior Data Engineer
Location: Gurugram, Haryana
Experience Required: 4 to 8 Years
Primary Tech Stack: AWS, PySpark, Advanced SQL
Role Overview
We are looking for a dynamic and results-driven Data Engineer with strong expertise in AWS, PySpark, and SQL to join our growing technology team in Gurugram. In this role, you will design, build, and optimize large-scale data pipelines, data warehouses, and modern cloud analytics platforms to handle high-volume data workloads.
Key Responsibilities
Pipeline Development: Design, develop, test, and maintain robust ETL/ELT data pipelines using PySpark and cloud-native services.
Cloud &
Big Data Management: Build, monitor, and optimize scalable data ingestion, transformation, and processing workflows on AWS (e.g., S3, Glue, EMR, Athena, Redshift, Lambda).
SQL Optimization: Write complex SQL queries, perform query tuning, and manage database operations to ensure high performance and low latency.
Data Modeling &
Architecture: Collaborate with cross-functional teams to build data models, schema designs, and data marts supporting analytical reporting.
Performance Tuning: Troubleshoot and resolve production performance bottlenecks in distributed data processing jobs.
Collaboration: Work closely with Data Scientists, Business Analysts, and DevOps teams to align data platform infrastructure with business requirements.
Best Practices: Ensure data quality, security, governance, and CI/CD automation standards are implemented across all deliverables.
Required Qualifications &
Skills
Experience: 4 to 8 years of hands-on experience in Data Engineering, Big Data, or Business Intelligence roles.
Programming Languages: Expert-level proficiency in Python and PySpark.
Cloud Ecosystem: Strong production experience working with AWS cloud services (S3, Glue, EMR, Athena, Redshift, etc.).
Database &
Querying: Strong command over SQL programming, performance tuning, and database design principles.
Big Data Frameworks: Familiarity with distributed computing principles and the Apache Spark ecosystem.
Tools &
Version Control: Experience with orchestration tools (e.g., Apache Airflow), containerization, and version control systems (Git).
Education: Bachelors or Masters degree in Computer Science, Information Technology, Engineering, or a related quantitative field.
Preferred / Good-to-Have Skills
Exposure to modern data lakehouse platforms like Databricks or Snowflake.
Experience working in fast-paced FinTech or Banking domains.
Familiarity with CI/CD deployment models and infrastructure-as-code concepts. .