Job Description
About The Role
The role is responsible for designing, building, and optimizing the core data infrastructure and pipelines that power business intelligence and machine learning models. This team processes terabytes of structured and unstructured data daily, turning raw transactional and clickstream events into reliable, high-performing data assets.
The engineer will collaborate closely with data scientists, software developers, and business stakeholders to ensure data availability, reliability, and security. The ideal candidate cares deeply about data quality, pipeline latency, and engineering best practices like CI/CD, version control, and comprehensive testing.
Key Responsibilities
- Design, build, and maintain scalable batch and real-time data pipelines using Apache Spark, dbt, and Airflow
- Optimize data warehouse performance in Snowflake, BigQuery, or Redshift by managing clustering keys, partitioning, and query optimization
- Implement robust data quality checks, anomaly detection, and schema validation throughout the ingestion and transformation lifecycle
- Collaborate with backend engineering teams to integrate source system event streams via Kafka, Kinesis, or Pub/Sub into the central lakehouse
- Establish infrastructure-as-code deployments for data pipelines using Terraform and maintain CI/CD pipelines in GitHub Actions
- Monitor pipeline health, system latency, and cloud compute costs, implementing optimizations to reduce infrastructure spend
What We Are Looking For
- 3–6 years of dedicated data engineering experience in a cloud-native production environment (AWS, GCP, or Azure)
- Advanced, production-grade SQL and Python programming skills, with a strong understanding of software engineering patterns
- Hands-on experience building complex data transformation pipelines using dbt (data build tool) and orchestrating them with Airflow or Prefect
- Proven experience with distributed compute frameworks such as Apache Spark, Databricks, or AWS EMR
- Strong understanding of data modeling principles, including Kimball dimensional modeling, Data Vault, or wide-table strategies
- Bonus: Experience with real-time streaming architectures (Kafka, Flink) or managing infrastructure using Terraform
Looking for more opportunities?
Browse thousands of graduate jobs and entry-level positions.