Invoca is the revenue execution and platform leader that helps marketing, commerce, and contact center teams turn every conversation into revenue. In this role, you will build the core data infrastructure that powers AI and ensures the reliability and efficiency of data pipelines for the data science and ML teams.
Responsibilities:
- Design, build, and operate scalable batch and streaming data pipelines that process structured and unstructured data across the Invoca Platform
- Extend and evolve the architecture of Invoca's lakehouse on Databricks (Delta Lake, Unity Catalog, Workflows) as a shared platform, not a one-off system
- Build curated training and fine-tuning corpora, evaluation datasets, and feature pipelines consumed directly by data science and ML engineering
- Treat dataset quality and availability as a direct input to how fast those teams can ship, and how well what they ship performs for customers
- Champion pipeline testing, CI/CD for data code, data quality monitoring, schema management, and dataset versioning for reproducibility
- Reduce the number of data incidents that reach downstream consumers, and shrink the time to detect and resolve the ones that do
- Ensure responsible data handling across the platform: governance, access control, and PII protection for customers in regulated industries
- Evaluate and implement new technologies as needed, and partner with technical leadership to drive adoption
- Mentor others across data science, engineering, and architecture, raising the data engineering bar across the company
- Partner with technical leadership to prioritize data platform investments based on the impact to the product and AI teams who depend on you
Requirements:
- 5+ years of professional experience in data engineering or closely related software engineering roles, with clear ownership of production data systems
- Advanced proficiency in Python for data engineering, including PySpark and modern data processing libraries (e.g., Pandas, Polars, PyArrow), with solid software engineering fundamentals: testing, code review, and CI/CD
- Advanced proficiency in SQL, including performance tuning, complex analytical queries, and data modeling for both relational databases (e.g., MySQL, PostgreSQL) and lakehouse/warehouse environments
- Strong experience with the Databricks platform (Delta Lake, Unity Catalog, Workflows, Jobs/Compute) or an equivalent lakehouse/warehouse stack (e.g., Snowflake, BigQuery) with willingness to go deep on Databricks
- Strong experience with Apache Spark at scale: partitioning, skew, join strategies, and cost/performance optimization
- Experience with workflow orchestration (e.g., Databricks Workflows, Airflow, Dagster), including dependency management, backfills, and idempotent pipeline design
- Experience with data quality and pipeline reliability practices: expectations/validation frameworks, monitoring and alerting, incident response, and schema evolution
- Experience with streaming data technologies (e.g., Kafka, Kinesis, Spark Structured Streaming) for real-time or near-real-time processing
- Working knowledge of cloud infrastructure, with a preference for AWS (S3, IAM, and containers), including security fundamentals and cost awareness
- Familiarity with the data needs of ML and AI systems: preparing training/fine-tuning and evaluation datasets, working with unstructured text, and supporting model-serving infrastructure (e.g., FastAPI services, Databricks Model Serving, SageMaker endpoints)
- Understanding of data governance and privacy practices; experience handling sensitive data (PII, HIPAA/PCI contexts) is a strong plus
- Bachelor's Degree or equivalent experience required