Data Engineer

Programming.com

remote globalPosted Jul 10, 2026
Posting intelligenceActively listedReposted 28×, possible evergreen/ghost posting

Skills

mlkubernetesdatabricksairflowdockergithubpythonazuresparkkafkaflinkcicdgooglecloudaws

About the role

Role: Data Engineer

Location: Remote (Requires occasional travel)

Responsibilities:

Build and operate robust data pipelines for ingestion, cleaning, and transformation using Databricks, Airflow, or Dagster.

Develop efficient ETL/ELT workflows in Python and SQL to support both batch and streaming workloads.

Collaborate with ML and AI teams to deliver high-quality datasets for training, evaluation, and production features.

Model and maintain structured data assets (Delta, Parquet, Iceberg) for reliability, versioning, and lineage tracking.

Implement orchestration and monitoring - schedule jobs, track dependencies, and automate recovery from failures.

Ensure data quality and compliance through validation frameworks, schema enforcement, and audit logging.

Contribute to data platform evolution - evaluate tools, standardize best practices, and improve developer experience.

Support performance and cost optimization across compute, storage, and orchestration systems.

Qualifications:

3–6 years of experience as a Data Engineer or ETL Developer in a production environment.

Proficiency in Python and SQL; strong familiarity with Databricks, Spark, or equivalent big-data frameworks.

Experience with workflow orchestration tools such as Airflow, Dagster, Luigi or Prefect.

Deep understanding of data modeling, data warehousing, and distributed data processing.

Knowledge of modern data lakehouse architectures (Delta, Parquet, Iceberg).

Familiarity with CI/CD, GitHub Actions, and data pipeline testing frameworks.

Comfort working in a cross-functional environment with ML, product, and analytics teams.

Nice to Have:

Experience with sports, telemetry, or sensor data pipelines.

Familiarity with streaming frameworks (Kafka, Spark Structured Streaming, Flink).

General knowledge of American football, the NFL, and college football

Background in data governance, lineage, and observability tools (Monte Carlo, Great Expectations, Unity Catalog, OpenLineage).

Experience with cloud infrastructure (AWS, GCP, or Azure) and containerization (Docker, Kubernetes).

Exposure to best practices in machine-learning model management and MLOps

Questions about this role

Click "Apply with AI Applyd" above. We auto-fill the application from your resume and answer screening questions in seconds. No copy and paste, no juggling tabs.

Compensation varies by seniority, employer size, and location. When this listing publishes a salary band you'll see it in the badge row above the description.

Most applications complete in under 90 seconds. You can track the status in your dashboard and watch the screenshot proof land the moment the application submits.

AI Applyd supports Greenhouse, Lever, Ashby, Workday, iCIMS, SmartRecruiters, Personio, Teamtailor and other major ATS platforms. If we can submit through the platform, we do.

Want AI Applyd to auto-apply to roles like this?

We tailor your resume per posting, fill the forms, and track replies for you.