
Databricks Developer
Skills
About the role
As a Databricks Developer, you will design and build the enterprise data pipelines that power analytics, reporting and AI initiatives for a leading company in the energy sector. Join a fully remote data engineering team working hands-on with cutting-edge Lakehouse technology.
HIGH-IMPACT DATA PROJECTS
LATEST LAKEHOUSE TECH
FULLY REMOTE
LEARNING & GROWTH
Don't tick every box? If you meet around 70% of the requirements above, we'd still encourage you to apply.
ABOUT THE ROLE
We are looking for a highly skilled Data Engineer with 5+ years of experience to design, build and optimize enterprise data pipelines on the Databricks Lakehouse platform for a leading energy sector company. In this role, you will be the hands-on technical driver responsible for transforming raw data into high-quality, actionable datasets. You will build and maintain a Medallion architecture, optimize Spark workloads, and ensure the data infrastructure seamlessly supports advanced analytics, BI dashboards and emerging Generative AI applications.
KEY RESPONSIBILITIES
Data Pipeline Engineering
Design, build and maintain scalable, robust ETL/ELT pipelines using Python, SQL and Apache Spark within the Databricks environment.
Implement and manage a robust Medallion architecture (Bronze, Silver, Gold layers) to process and refine data from diverse sources.
Develop and maintain the Gold semantic layer specifically optimized for high-performance consumption by BI tools (e.g., Power BI).
Platform Optimization & Architecture
Optimize Databricks workloads, cluster configurations and Spark queries to ensure high performance and cost efficiency.
Work extensively with open table formats, specifically Delta Lake and Apache Iceberg, to ensure ACID compliance, time travel and efficient data storage.
Execute complex data migrations, including transitioning legacy workloads from traditional cloud data warehouses (e.g., AWS Redshift) into the Databricks Lakehouse.
Data Governance & Automation
Implement data governance and access control policies at the table, row and column levels using Databricks Unity Catalog.
Automate deployment processes and pipeline orchestration using Databricks Workflows, CI/CD pipelines (e.g., GitHub Actions, Azure DevOps) and tools like Terraform.
Embed data quality checks and monitoring directly into pipelines to ensure strict Master Data Management (MDM) standards are upheld.
AI & Advanced Analytics Support
Collaborate closely with Data Scientists and AI Engineers to provision clean, structured data for machine learning model training and inference.
Support the data foundations required for GenAI frameworks, autonomous agents and AI observability platforms.
REQUIRED SKILLS & EXPERIENCE
5+ years of dedicated data engineering experience in an enterprise environment.
Expert-level proficiency in Python and SQL.
Extensive hands-on experience with Databricks, Apache Spark and Delta Lake.
Strong understanding of distributed systems, big data architecture and data modeling techniques (e.g., Kimball, Data Vault).
Deep familiarity with cloud-native data services (AWS, Azure or GCP), specifically cloud storage (S3/ADLS) and compute provisioning.
Proven experience with version control (Git), CI/CD methodologies and agile software development life cycles.
NICE TO HAVE
Experience evaluating and working with Apache Iceberg alongside Delta Lake.
Familiarity with streaming data architectures (e.g., Structured Streaming, Kafka).
Experience building backend frameworks or internal tools using lightweight libraries like Streamlit.
EDUCATION
Bachelor's or Master's degree in Computer Science, Information Technology, Data Engineering or a related field.
PREFERRED CERTIFICATIONS
Databricks Certified Data Engineer Associate or Professional.
AWS, Azure or GCP data/cloud certifications.
WORKING MODEL
Fully remote position.
Questions about this role
Want AI Applyd to auto-apply to roles like this?
We tailor your resume per posting, fill the forms, and track replies for you.