Job Summary
Responsibilities
- Develop and maintain data pipelines using Apache Airflow to ensure efficient data processing and workflow automation.
- Utilize Hive for data warehousing solutions ensuring data is stored and queried effectively.
- Write and optimize SQL queries to manage and manipulate data across various databases.
- Implement and manage data processing on AWS EMR ensuring scalable and cost-effective solutions.
- Leverage Apache Spark for distributed data processing ensuring high performance and reliability.
- Collaborate with cross-functional teams to understand data requirements and deliver solutions that meet business needs.
- Monitor and troubleshoot data pipelines and workflows to ensure smooth operation and mnimal downtime.
- Provide technical guidance and support to junior developers fostering a collaborative and knowledge-sharing environment.
- Stay updated with the latest industry trends and technologies to continuously improve data processing capabilities.
- Ensure data security and compliance with relevant regulations and standards.
- Document technical specifications and processes to maintain clear and comprehensive records.
- Participate in code reviews to ensure code quality and adherence to best practices.
Qualifications
- Extensive experience with Apache Airflow for workflow automation and scheduling.
- Proficiency in Hive for data warehousing and query optimization.
- Strong SQL skills for data manipulation and management.
- Hands-on experience with AWS EMR for scalable data processing.
- In-depth knowledge of Apache Spark for distributed data processing.
- Excellent problem-solving skills and attention to detail.
- Solid communication and collaboration skills.
- Ability to work in a hybrid model and adapt to changing project requirements.
- Experience with data security and compliance standards.
- Familiarity with code review processes and best practices.
- Commitment to continuous learning and improvement.
- Proven track record .