Location
Remote
Salary
Not specified
Type
fulltime
Posted
Today
Job Description
Job Description
Strong SQL and Python.
Data modelling, warehousing concepts, partitions, clustering, indexing, optimisation.
ETL/ELT pipeline design and production support.
Batch and streaming pipeline fundamentals.
Orchestration using Airflow, Composer, Control-M, ADF, dbt, or equivalent.
Big data processing using Spark, PySpark, Beam, Dataflow, Dataproc, Databricks, EMR, or equivalent.
Cloud storage and data lake concepts.
Data quality, reconciliation, schema validation, lineage, monitoring, and alerting.
CI/CD for data pipelines and code versioning.
Ability to explain scale scenarios, not just tool definitions.
Generic Skills (Must Have)
Python and advanced SQL.
Data quality checks, reconciliation, alerting, retries, and backfills.
Production pipeline monitoring and troubleshooting.
GCP Skills (Must Have)
Mandatory: BigQuery: datasets, tables, views, partitioning, clustering, SQL optimisation, cost awareness.
Option 1: Dataflow / Apache Beam or equivalent streaming/batch framework.
Option 2: Dataproc / PySpark or strong Spark experience.
Pub/Sub or equivalent event streaming.
Cloud Composer / Airflow DAG development.
Nice to have (Trainable)
Dataform, dbt, CI/CD for data.
BigQuery ML.
Dataflow templates.
Looker / LookML basics.
Terraform for data infrastructure.
Skills: etl,sql,gcp,elt,data engineer,pyspark,python
Looking for more opportunities?
Browse thousands of graduate jobs and entry-level positions.