Big Data Engineer
Skills
About the role
Based on the image, here's a professionally formatted Job Description for a Big Data Engineer (EL3).
Job Title: Big Data Engineer (EL3)
Experience: 4+ Years
Employment Type: Full-Time
Location: [Location]
Job Summary
We are looking for a skilled Big Data Engineer with 4+ years of experience in designing, developing, and optimizing scalable big data solutions. The ideal candidate should have strong expertise in Scala, PySpark, AWS, and AWS Glue, with hands-on experience building high-performance data pipelines and data validation frameworks. The candidate will work closely with cross-functional teams to deliver scalable, high-quality data engineering solutions while leveraging AI-powered development tools to improve productivity.
Key Responsibilities
Design, develop, and optimize highly scalable, rule-based data validation engines capable of validating data across multiple data pipelines.
Build and maintain high-performance big data pipelines using Scala, PySpark, AWS, and AWS Glue.
Develop clean, efficient, and well-tested code while utilizing optimized data formats such as Parquet for high-performance storage and retrieval.
Perform data validation, profiling, and quality assessments to ensure consistency and accuracy across enterprise data systems.
Leverage enterprise-approved AI coding assistants and code-generation tools to improve software development efficiency and automate repetitive tasks.
Collaborate with business stakeholders and cross-functional engineering teams to translate business requirements into scalable technical solutions.
Evaluate emerging big data technologies and AI-driven engineering practices to enhance platform capabilities.
Participate in code reviews, testing, debugging, and performance tuning to ensure high-quality software delivery.
Continuously optimize data processing performance, scalability, and reliability across distributed systems.
Required Qualifications
Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, or a related field.
4+ years of software engineering experience developing big data pipelines.
3+ years of hands-on programming experience in Scala and PySpark.
3+ years of experience building big data solutions on AWS.
Strong experience with AWS Glue for ETL development.
Experience working with big data storage formats such as Parquet.
At least 1 year of experience in data validation, data profiling, or building rule-based data validation engines.
Experience using AI coding assistants or code-generation tools (such as GitHub Copilot or Codex) during software development.
Required Technical Skills
Scala
PySpark
Apache Spark
AWS
AWS Glue
Parquet
ETL Development
Data Validation
Data Profiling
Big Data Pipelines
Distributed Data Processing
SQL
Git
AI Coding Assistants (GitHub Copilot, Codex, etc.)
Preferred Skills
Experience with large-scale distributed data processing.
Knowledge of data quality frameworks and governance practices.
Strong analytical and problem-solving abilities.
Familiarity with CI/CD pipelines and software engineering best practices.
Excellent communication and collaboration skills.
What We Offer
Opportunity to work on enterprise-scale big data platforms.
Exposure to AI-assisted software development practices.
Collaborative and innovation-driven work environment.
Career growth with modern cloud and big data technologies.
* Competitive compensation and benefits package.
Work Location: Hybrid remote in Remote
Questions about this role
Want AI Applyd to auto-apply to roles like this?
We tailor your resume per posting, fill the forms, and track replies for you.