Location
New York, United States
Salary
$140,000 - $190,000 /yearly
Type
fulltime
Posted
Today
Job Description
Role overview
The Data Engineer builds the pipelines and data platforms that support analytics, reporting, AI development, and operational decisions. Working across ingestion, transformation, storage, modeling, and delivery, this position turns complex business needs into reliable production systems, including batch and real-time pipelines, OLTP and OLAP stores, semantic models, and BI-ready datasets. The work combines cloud engineering, automation, observability, and infrastructure-as-code in a regulated environment.
Building and Operating the Data Foundation
- Design, develop, and maintain data pipelines connecting internal systems, third-party platforms, files, APIs, event streams, and databases. Build scalable ETL and ELT workflows for batch and real-time processing.
- Make pipelines testable, observable, reliable, and extensible. Develop reusable integration patterns for increasing data volumes, new sources, and consumers across analytics, applications, and AI.
- Design and manage architectures for transactional, analytical, and reporting workloads. Optimize data models, warehouse schemas, and curated datasets for analytics and business intelligence.
- Help design and operate warehouses, lakehouses, streaming systems, and orchestration frameworks, with practical approaches to storage, partitioning, performance, retention, and lifecycle management.
- Deploy and operate data pipelines and stores on AWS, GCP, or Azure. Use infrastructure-as-code, CI/CD, and automation to improve deployment consistency, speed, and reliability.
- Monitor production systems through logging, alerting, and observability tools; identify and resolve issues while supporting secure, resilient, cost-conscious cloud infrastructure.
- Implement data-quality checks, validation, reconciliation, and monitoring. Maintain lineage, documentation, metadata, schema-evolution standards, and operational runbooks.
- Improve data accessibility, consistency, and usability with appropriate controls. Support security, privacy, auditability, and regulatory compliance.
- Partner with Product, Engineering, AI, Analytics, business stakeholders, and domain subject matter experts to understand workflows and constraints and translate requirements into maintainable data solutions.
- Support analysts, researchers, product teams, and operational users. Explain data availability, quality, tradeoffs, and delivery timelines to technical and nontechnical stakeholders.
- Improve performance, reliability, scalability, and developer productivity; simplify architecture, reduce operational toil, and extend the value of shared data platforms.
- Move iteratively from problem definition to implementation and improvement, strengthening engineering quality through design, testing, documentation, and operational discipline.
Required Technical Background
- 2–4\+ years building and operating production-grade data pipelines and systems.
- Strong experience with standard ETL/ELT, orchestration, warehousing, streaming, and BI tools and platforms.
- Experience with OLTP and OLAP systems and an understanding of transactional versus analytical workload tradeoffs.
- Experience integrating databases, APIs, files, message queues, SaaS platforms, and event streams through flexible pipelines supporting batch and real-time processing.
- Experience deploying and operating cloud data infrastructure on AWS, GCP, or Azure.
- Strong SQL capabilities, including data modeling, transformation frameworks, and performance optimization.
- Experience building LLM-based AI capabilities, including orchestration, evaluation, and data integration patterns.
- Experience with data-engineering programming languages such as Python, Java, Scala, or Go.
- Comfort with CI/CD, infrastructure-as-code, observability, and production operations.
- Sound judgment when requirements are ambiguous or changing, balancing speed, reliability, and flexibility; clear communication with technical and nontechnical colleagues.
Preferred Experience
- Orchestration and transformation platforms such as Airflow, Dagster, or dbt.
- Cloud warehouses or lakehouses such as Snowflake, BigQuery, Redshift, or Databricks.
- Streaming and real-time platforms such as Kafka, Kinesis, or SQS.
- Curated datasets, semantic layers, and self-service analytics using Looker, Power BI, Tableau, or similar reporting tools.
- Fintech, mortgage, lending, payments, insurance, or other regulated-domain experience.
- Data platforms supporting AI, machine learning, or decisioning workflows.
- Improving data quality, reliability, cost efficiency, and scalability as platforms grow.
Prior fintech or finance experience is not required. Candidates with strong data-engineering skills, technical judgment, and systems thinking are encouraged to apply even if their background does not match every listed qualification.
Work and Compensation
This is a regular full-time, fully remote position. The posting lists New York, New York, United States, and Toronto, Canada. The salary range is USD $140,000–$190,000, plus an annual bonus. Compensation may vary with experience, location, and other job-related factors.
Benefits include medical coverage beginning on the first day and a company-matched 401(k).
Employment decisions are based on merit, competence, and qualifications, without regard to race, color, religion, sex, national origin, age, disability, veteran status, sexual orientation, or any other status protected by applicable federal, state, or local law.
Looking for more opportunities?
Browse thousands of graduate jobs and entry-level positions.