Synthetic Data Engineer (AI Data/Training)

Hyphen Connect

Hong Kong, HKonsitePosted Apr 24, 2026
Posting intelligenceMay be filled, listed long ago

Skills

airflowspark

About the role

We are seeking a talented and innovative Synthetic Data Engineer. In this role, you will design and implement domain-specific synthetic data generation pipelines, ensuring high-quality data management for training loops. Your expertise will drive the success of data processing and model training within the organization.

Responsibilities:

Design domain-specific synthetic data generation (SDG) pipelines via self-instruct and constitutional prompting.

Implement automated quality scoring and de-duplication systems.

Manage data pipelines that feed directly into SFT and DPO training loops.

Qualifications:

Proven experience building large-scale data pipelines (Airflow, Spark, Ray).

Deep knowledge of prompt engineering for data generation.

Familiarity with dataset distillation and bias mitigation.

Questions about this role

Click "Apply with AI Applyd" above. We auto-fill the application from your resume and answer screening questions in seconds. No copy and paste, no juggling tabs.

Compensation for Data Engineer roles in Hong Kong varies widely by seniority, employer size, and remote vs onsite arrangement. Check the salary range on this listing when published, or browse our Data Engineer hub for Hong Kong medians across recent openings.

Most applications complete in under 90 seconds. You can track the status in your dashboard and watch the screenshot proof land the moment the application submits.

AI Applyd supports Greenhouse, Lever, Ashby, Workday, iCIMS, SmartRecruiters, Personio, Teamtailor and other major ATS platforms. If we can submit through the platform, we do.

Want AI Applyd to auto-apply to roles like this?

We tailor your resume per posting, fill the forms, and track replies for you.