Big Data Engineer (Cancer Science Institute)

National University of Singapore

unknownPosted Jul 1, 2026
Posting intelligenceActively listedReposted 24×, possible evergreen/ghost posting

Skills

cloudformationelasticsearchpostgreskubernetesterraformdynamodbpythonaws

About the role

Job Title: Big Data Engineer (Cancer Science Institute)

University-Level Unit: Cancer Science Institute of Singapore

Faculty/Department-Level Unit: Research

Employee Category: Research Staff

Location_ONB: Kent Ridge Campus

Posting Start Date: 01/07/2026

Job Description

The Cancer Science Institute of Singapore – a part of National University of Singapore – is seeking a skilled Big Data Engineer to join the Genomics and Data Analytics Core (GeDaC). We are operating a petabyte-scale "Data Nexus" that serves as the foundation for a production AI Factory in cancer and human disease research.

You do not need a background in biology . We are looking for a pure engineer who understands data logistics, infrastructure, and scale.

The Team & Leadership

You will join a highly specialized technical team comprising an experienced Cloud/HPC Architect, an agile Full-Stack Developer, and a senior IT Manager.

Crucially, you will report to a Facility Head with deep, hands-on expertise in petabyte-scale data-intensive computing and DataOps. This ensures you will work in an environment where technical complexity is understood, architectural decisions are respected, and job scope is managed with engineering reality in mind.

Key Responsibilities

Data Ingestion & Logistics: Architect and maintain robust automation for ingesting raw data from sequencing instruments to our hybrid storage systems. You will own the "handshakes" that ensure data moves reliably from edge to cloud.

Infrastructure as Code (IaC): Manage and deploy AWS resources (S3, Lambda, DynamoDB, RDS) using AWS CloudFormation, ensuring our infrastructure is reproducible, version-controlled, and follows DevSecOps best practices.

Technical Compliance & Provenance: Implement the technical controls for data governance. This includes designing immutable audit logs, automated access control policies, and lineage tracking systems to satisfy regulatory requirements (no manual report writing required).

Hybrid Cloud Synchronization: Manage the lifecycle of data moving between on-premise HPC and AWS S3 Intelligent-Tiering/Glacier to balance high-performance availability with long-term cost optimization.

Pipeline Integration: Work closely with the Senior HPC Engineer to ensure data is correctly staged for Nextflow/Kubernetes processing pipelines, and capture the outputs back into the data lake/warehouse.

Database Management: Maintain the SQL and NoSQL databases that serve as the "source of truth" for file metadata, ensuring the Full-Stack team has low-latency API access to query file status.

Requirements

Education : Bachelor’s degree in Computer Science, Information Systems, Engineering, or a related field.

Experience :

2+ years of experience in Data Engineering, Backend Development, or DevOps.

Demonstrable experience working with commercial cloud infrastructure (AWS preferred)

Technical Skills :

Core Logic: Strong proficiency in Python (data tooling, automation scripts).

Infrastructure: Experience with Infrastructure as Code (IaC) tools such as AWS CloudFormation and/or Terraform is essential.

Data Management: Proficiency with SQL (PostgreSQL/Aurora) and object storage (S3).

Environment: Beyond comfortable working in Linux/Unix environments.

Attributes :

Meticulous: You care deeply about data integrity. A missing file or a broken checksum bothers you.

Ownership-driven: You take responsibility for systems you build and operate.

Collaborative: You can work effectively within an established technical team, integrating your work with existing APIs and processing pipelines.

Preferred Experience

Experience with workflow managers like Nextflow or container orchestration via Kubernetes.

Experience with hybrid-cloud data transfer tools (e.g., AWS DataSync, Storage Gateway).

Knowledge of searching/indexing tools like Elasticsearch or OpenSearch.

Questions about this role

Click "Apply with AI Applyd" above. We auto-fill the application from your resume and answer screening questions in seconds. No copy and paste, no juggling tabs.

Compensation varies by seniority, employer size, and location. When this listing publishes a salary band you'll see it in the badge row above the description.

Most applications complete in under 90 seconds. You can track the status in your dashboard and watch the screenshot proof land the moment the application submits.

AI Applyd supports Greenhouse, Lever, Ashby, Workday, iCIMS, SmartRecruiters, Personio, Teamtailor and other major ATS platforms. If we can submit through the platform, we do.

Want AI Applyd to auto-apply to roles like this?

We tailor your resume per posting, fill the forms, and track replies for you.