01Overview
About the job
Data Engineer – Data Migration Factory
Location: Bangalore – Onsite/ Hybrid
Experience: 8 + Years
Employment Type: Full Time
Budget: 30 - 40 LPA
Job Description
The Data Engineer will be part of the Data Migration Factory team responsible for end-to-end datastore migration from an on-premises Data Lake to an AWS-hosted Lakehouse environment. This is a high-visibility and critical data migration initiative.
Key Responsibilities
1. Pipeline Migration
Refactor and migrate extraction logic and job scheduling from legacy frameworks to the new Lakehouse environment.
Execute physical migration of underlying datasets while ensuring data integrity.
Work with data owners and stakeholders to facilitate hand-off and sign-off of migrated assets.
Act as a technical liaison between engineering teams and internal data stakeholders.
2. Consumption Pattern Migration
Translate and optimize legacy SQL and Spark-based consumption patterns for compatibility with Snowflake and Apache Iceberg.
Analyze data usage patterns and deliver required data products.
Work with internal stakeholders to understand business and technical requirements.
Facilitate hand-off and sign-off conversations with data owners.
3. Data Reconciliation & Quality
Perform rigorous data validation and reconciliation to ensure migrated data is functionally equivalent to production data.
Work with reconciliation frameworks to validate data accuracy and completeness.
Identify and troubleshoot data quality and migration issues.
Collaborate with internal data management platform teams.
Learn and adopt new workflows, technologies and language constructs as required.
Basic Qualifications
Bachelor’s or Master’s degree in Computer Science, Applied Mathematics, Engineering, or a related quantitative field.
3–5 years of hands-on coding experience in a collaborative, team-based environment.
Strong SQL troubleshooting and basic scripting skills.
Professional proficiency in Python or Java.
Strong understanding of the Software Development Life Cycle (SDLC).
Experience with CI/CD best practices.
Experience with Kubernetes (K8s) deployments.
Core Data Engineering Competencies
Temporal Data Modeling: Strong understanding of managing state changes over time, including SCD Type 2.
Schema Management: Experience with schema evolution and enforcement strategies, including Apache Iceberg.
Performance Optimization: Strong understanding of data partitioning and clustering.
Data Architecture: Understanding of normalization vs. denormalization and natural vs. surrogate keys.
Strong understanding of data correctness and reconciliation principles.
Technical Skills
Extraction & Data Processing
Kafka
ANSI SQL
FTP
Apache Spark
Data Formats
JSON
Avro
Parquet
Data Platforms
Hadoop
HDFS
Hive
Snowflake
Apache Iceberg
Sybase IQ
Core Competencies
Strong integrity and ethical decision-making.
Effective collaboration across multiple teams and functions.
Clear and confident communication.
Strong stakeholder management skills.
Ability to work effectively with global teams across time zones and cultures.
Strong ownership and delivery focus.
Ability to drive tasks to closure and meet commitments.
High energy and urgency while maintaining quality and professionalism.
Intellectual curiosity and willingness to learn new technologies and workflows.
Ability to identify risks early, ask thoughtful questions and continuously improve.