Senior Data Engineer
Skills
About the role
Job Description
AWS SageMaker Unified Studio Expertise:
Proficiency with AWS SageMaker Unified Studio, including its various components Discover, Build and Govern.
Experience with development in Sagemaker IDE and Application (Jupyterlab, Spaces and Partner AI Apps).
Data Analysis and Integrations ( Query Editor, Visual ETL Jobs and Data Processing jobs)
Orchestration of workflows and ML Pipelines
Understanding and expertise to work with ML and Gen AI tools available in unified studio are ad-on values.
Setting up projects and data governance in SageMaker Unified studio
Advanced Data Ingestion & Processing (Real-time & Batch):
Framework Development: Proven ability to design, develop, and implement highly reusable and adaptable data ingestion frameworks capable of handling diverse source types (e.g., databases, APIs, message queues, file systems).
Low-Latency Real-time: Deep expertise in building real-time data pipelines with stringent sub-second latency requirements, utilising services like AWS Kinesis (Data Streams/Firehose), Apache Kafka, or similar streaming technologies.
Batch Processing: Experience with robust batch ingestion patterns and tools (e.g., AWS Glue, Apache Spark) for efficient processing of larger datasets.
Data Transformation: Stron g skills in designing and implementing efficient data transformation logic for both streaming and batch data.
Programming and Scripting:
Advanced proficiency in Python, particularly for developing scalable data ingestion and export frameworks, API integration, and extensive use of the AWS SDK (Boto3).
Experience with performance optimisation techniques for Python applications in data-intensive environments.
Familiarity with other relevant languages (e.g., Scala, Spark) for high-performance streaming applications is beneficial.
AWS Services for Data and Infrastructure:
In-depth knowledge of core AWS services: Lambda, S3, DynamoDB, CloudWatch, SQS, SNS, and API Gateway.
Strong understanding of AWS networking (VPC, security groups, private endpoints) and IAM for secure, fine-grained access control.
Key Skillset:
Previous experience in developing framework for batch and real time data ingestions in relation and no-sql databases or filesystems.
Previous experience in Real time data ingestion with low latency with experience on AWS Kinesis Stream, Apache Kafka or similar streaming technologies
Proven experience in various database technologies (e.g. Oracle, Teradata, MongoDB, Snowflake etc)
Sound knowledge and experience in building Data Warehouses and Data Lakehouses.
Data Modeling experience is a plus point
Location
Sydney
Job Function
TECHNOLOGY
Role
Engineer
Job Id
422253
Desired Skills
AWS | Python | Apache
Questions about this role
Want AI Applyd to auto-apply to roles like this?
We tailor your resume per posting, fill the forms, and track replies for you.