01Responsibilities
Develop and maintain data processing pipelines using PySpark, SQL, and Hadoop.
Collaborate with data scientists and analysts to optimize data workflows.
Implement data transformation and aggregation processes.
Ensure secure data access and compliance with data governance policies.
Perform Spark job tuning and performance optimization.
Write unit tests and documentation for Spark transformations.
Work on continuous improvement of data processing frameworks.
Skills Required:
Proficiency in PySpark, SQL, and Hadoop.
Experience with big data technologies and frameworks.
Solid problem-solving and analytical skills.
Ability to work in a team-oriented team setting.
Excellent communication skills. .