Data Engineer PySpark Experiencer
Location: Dallas TX
Job Summary
We are seeking an experienced Data Engineer with strong expertise in PySpark, ETL integration, and large-scale data migration projects. The ideal candidate will be responsible for designing, developing, and optimizing scalable data pipelines, integrating data from multiple enterprise systems, and executing complex data migration initiatives while ensuring data quality, governance, and performance.
Required Experience
10 years of experience in Data Engineering. Experience with PySpark and Apache Spark. Experience designing and developing ETL/ELT pipelines.
Proven experience in data migration projects involving large datasets.
Experience working with cloud-based data platforms (AWS, Azure, or Google Cloud Platform).
Strong SQL programming and data modeling skills.
Key Responsibilities Design, build, and maintain scalable data pipelines using PySpark. Develop and support ETL/ELT processes for ingesting, transforming, and loading data.
Perform end-to-end data migration from legacy systems to modern cloud data platforms.
Integrate data from multiple structured and unstructured sources.
Optimize Spark jobs for performance, scalability, and reliability.
Develop reusable frameworks for data ingestion and transformation.
Validate migrated data for completeness, consistency, and accuracy.
Collaborate with business analysts, architects, data scientists, and application teams.
Troubleshoot production issues and provide root cause analysis.
Implement monitoring, logging, and error-handling mechanisms.
Maintain technical documentation and support release activities.