Relativix is a vehicle intelligence platform that connects fleet operators and repair shops through shared diagnostic data. The Data Engineer will design and maintain data pipelines, optimize ETL processes, and collaborate with teams to ensure high-quality datasets for analytics and machine learning.
Responsibilities:
- Design, build, and maintain large-scale data pipelines on the Databricks platform
- Implementing and optimizing ETL processes at scale
- Developing and refining data models
- Supporting data warehousing and lakehouse solutions
- Collaborating closely with software engineers, data scientists, and product teams
- Preparing and serving data for machine learning models
- Enabling advanced predictive capabilities
- Contributing to monitoring, troubleshooting, and continuously improving data systems
Requirements:
- Solid, hands-on understanding of the Databricks platform, including Spark, Delta Lake, and lakehouse architecture for large-scale data processing
- Proven experience designing, running, and optimizing large-scale ETL pipelines, including data ingestion, transformation, and integration from multiple sources
- Strong data engineering fundamentals, including building and maintaining scalable data pipelines and distributed data systems
- Solid understanding of AWS and its data services (e.g., S3, Glue, Kinesis, Redshift, Lambda) for building production data infrastructure
- Good understanding of machine learning models and their data requirements, including feature engineering, training/serving data pipelines, and supporting data scientists in deploying models to production
- Experience with data modeling and data warehousing to design robust schemas and storage solutions for analytics and reporting
- Proficiency in a modern programming language commonly used for data engineering (e.g., Python, Scala, or Java) and strong SQL skills
- Understanding of best practices in data quality, data governance, and security for production environments
- Bachelor's degree in Computer Science, Engineering, Data Science, or a related field, or equivalent practical experience
- A self-starter mentality with the drive to own projects end-to-end and thrive in the ambiguity of an early-stage startup
- Ability to work independently in a remote environment, communicate clearly with cross-functional teams, and manage priorities in a fast-paced setting
- Experience with IoT data manipulation, vehicle telemetry, automotive data, or real-time streaming analytics (e.g., Kafka, Kinesis, Spark Structured Streaming)