Company details are confidential
• Develop, test, and maintain scalable data pipelines and ETL processes using Python and PySpark to extract, transform, and load large datasets. • Optimize and fine-tune existing PySpark applications and workflows for improved performance and efficiency. • Work with big data technologies such as Hadoop, Hive, or Kafka, and AWS cloud platforms. • Work with and support solutions on the Databricks Lakehouse platform. • Ensure data quality and integrity throughout the data lifecycle through validation and error-handling mechanisms. • Monitor and troubleshoot data processing jobs in production environments to ensure reliability and timely issue resolution.