Data Engineer

Virtusa

New York City, USonsitePosted Jun 8, 2026
Posting intelligenceActively listed

Skills

databrickspythonazurespark

About the role

Role Summary

The Data Engineer will design and implement scalable, distributed data pipelines and enterprise data lakes, leveraging Spark, Databricks, Python, and SQL. The role focuses on building high-performance ETL/ELT pipelines, metadata-driven architectures, and governed analytical data stores supporting advanced analytics and machine learning workloads in cloud environments.

Key Responsibilities

Pipeline Design & Optimization Design high-performance ETL/ELT pipelines using PySpark and Spark SQL to translate business requirements into optimized data pipelines, demonstrating an ability to reduce processing latency by up to 50%.

Cloud Data Architecture Design and implement data models for Medallion architecture (Bronze, Silver, Gold) using Delta Lake, enabling scalable and reusable data processing.

Data Ingestion & Orchestration Orchestrate data pipelines using Azure Data Factory (ADF) to reliably ingest, transform, and load enterprise datasets. Implement data ingestion pipelines, including those connecting on-premises HDFS with Azure Data Factory and Databricks, to create curated Gold-layer datasets supporting Microsoft Fabric analytics.

Data Governance & Security Implement centralized data governance using Unity Catalog for managing catalogs, schemas, role-based access controls (RBAC), and fine-grained permissions.

Quality Assurance & Cost Management Build scenario-based test frameworks in Databricks using PySpark for data validation. Optimize storage costs (e.g., 25% reduction in Azure Storage) by managing required history/versions of Delta tables.

Operational Monitoring Generate an automated email reporting framework to set up pipeline failure alerts, reducing manual support efforts by 40%.

Questions about this role

Click "Apply with AI Applyd" above. We auto-fill the application from your resume and answer screening questions in seconds. No copy and paste, no juggling tabs.

Compensation for Data Engineer roles in United States varies widely by seniority, employer size, and remote vs onsite arrangement. Check the salary range on this listing when published, or browse our Data Engineer hub for United States medians across recent openings.

Most applications complete in under 90 seconds. You can track the status in your dashboard and watch the screenshot proof land the moment the application submits.

AI Applyd supports Greenhouse, Lever, Ashby, Workday, iCIMS, SmartRecruiters, Personio, Teamtailor and other major ATS platforms. If we can submit through the platform, we do.

Want AI Applyd to auto-apply to roles like this?

We tailor your resume per posting, fill the forms, and track replies for you.