Mphasis is seeking Senior Data Engineers to support their Legacy Batch Modernization and Rapid Data Ingestion platform enhancement initiatives. The role focuses on modernizing legacy batch processing workloads, migrating existing jobs into AWS Batch, and enhancing cloud-based data ingestion capabilities.
Responsibilities:
- Analyze legacy batch jobs, scripts, database procedures, and data workflows to understand current-state processing, dependencies, business logic, scheduling requirements, and operational constraints
- Rewrite and migrate legacy batch workloads into AWS Batch using modern engineering practices while preserving business functionality and meeting existing service-level agreements
- Develop, enhance, and maintain data pipelines and batch processing solutions using Java Spring Batch, Python, SQL, shell scripting, and AWS services
- Support integration between on-premises legacy systems and AWS-based target platforms, ensuring reliable data movement, synchronization, monitoring, and operational continuity
- Build and optimize data ingestion workflows to support the Rapid Data Ingestion platform, including file intake, validation, transformation, inventory tracking, and downstream consumption
- Implement data quality checks to detect changes in vendor file formats, data integrity issues, missing or incomplete data, and other anomalies that may impact critical business processes
- Enable and support data lineage capabilities by capturing source-to-target data movement, transformations, metadata, and audit-relevant processing details
- Support disaster recovery implementation for the Rapid Data Ingestion platform in AWS to improve resiliency and continuity for critical and audit-sensitive processes
- Develop and maintain job scheduling and orchestration processes using Control-M or equivalent scheduling tools
- Implement secure engineering practices, including application-specific database access, elimination of shared credentials, and enforcement of least-privilege access principles
- Perform performance tuning to ensure rewritten jobs meet or exceed current batch cycle expectations, including completion within strict operational SLAs
- Participate in data validation, reconciliation, defect triage, testing, production deployment, and warranty support activities
- Collaborate with data analysts, systems analysts, architects, application teams, database teams, governance teams, and business stakeholders to ensure successful delivery
- Create and maintain technical documentation covering design, mappings, dependencies, lineage, quality rules, operational procedures, and deployment details
Requirements:
- 9+ years of relevant data engineering experience, preferably in financial services, investment management, banking, or other highly regulated enterprise environments
- Strong experience with Java Spring Batch for batch job development, migration, and modernization
- Hands-on experience with AWS services, especially AWS Batch and Amazon S3
- Strong SQL development experience, including Oracle SQL and PL/SQL concepts
- Experience developing data engineering solutions using Python
- Strong shell scripting experience using Bash, KornShell, or similar Unix/Linux scripting languages
- Experience with Control-M or equivalent enterprise job scheduling and orchestration tools
- Experience working with legacy batch processing environments and modernizing on-premises batch workloads to cloud-based platforms
- Ability to analyze and convert legacy scripts, SQL loaders, stored procedures, and batch workflows into modern, maintainable data processing jobs
- Experience with data ingestion, ETL/ELT development, data validation, reconciliation, and production support
- Understanding of secure access patterns, credential management, least-privilege controls, and audit-focused data processing
- Experience with CI/CD, code deployment, version control, and standard software engineering practices
- Ability to understand business-critical batch cycles and translate legacy processing logic into reliable modern implementations
- Strong analytical and problem-solving skills to investigate data issues, job failures, performance bottlenecks, and production incidents
- Ability to work in regulated environments where data accuracy, lineage, auditability, and operational controls are critical
- Strong communication skills to collaborate with technical teams, business stakeholders, and governance partners
- Ability to work independently and as part of a distributed team across onsite, nearshore, offshore, and remote delivery models
- Strong ownership mindset with the ability to deliver within defined milestones, project timelines, and operational SLAs
- Bachelor's degree in Computer Science, Information Technology, Engineering, Data Engineering, or a related field, or equivalent practical experience
- Experience with Snowflake SQL and cloud data warehouse environments
- Working knowledge of Collibra or similar data governance, metadata management, and data lineage tools
- Experience with file inventory management, metadata capture, and source-to-target lineage reporting
- Experience implementing data quality frameworks, validation routines, exception handling, and notification workflows
- Familiarity with AWS disaster recovery patterns and resilient data platform design
- Exposure to financial services data, compensation processes, intermediary business processes, or audit-sensitive workloads
- Experience with legacy technologies such as ksh, Perl, sqlldr, sqlplus, and Oracle-based batch processing
- Familiarity with production incident management, escalation, monitoring, and operational runbooks