01Overview
Roles and Responsibilities
Design, develop, test, deploy, and maintain large-scale data processing pipelines using PySpark on AWS.
Collaborate with cross-functional teams to gather requirements and deliver high-quality solutions.
Develop complex ETL processes to extract insights from structured and unstructured data sources.
Ensure scalability, performance, and reliability of the developed applications.
Desired Candidate Profile
2-5 years of experience in PySpark development with expertise in AWS services such as S3, Glue, Lambda etc. .
Strong understanding of SQL concepts including joins, subqueries, aggregations etc. .
Experience working with Python programming language with knowledge of its libraries like NumPy pandas etc. .
Bachelor's degree in Any Specialization (B.C.A. / B.Sc.).
Hands-on experience with DataBricks framework for building scalable data pipelines. .