Senior Big Data Engineer

Citi

Pune, INonsitePosted Jul 22, 2026
Posting intelligenceActively listed

Skills

hadoopsparkscala

About the role

Discover your future at Citi

Working at Citi is far more than just a job. A career with us means joining a team of approximately 219,000 dedicated people from around the globe. At Citi, you’ll have the opportunity to grow your career, give back to your community and make a real impact.

Job Overview

Citi is looking for a Senior Big Data Engineer to design, build, and optimize large-scale data pipelines and distributed data systems that power critical business intelligence across the organization. Based in Pune and operating in a hybrid model, you will work within a high-performing engineering team where your expertise in PySpark, the Hadoop ecosystem, and streaming data platforms will directly shape the reliability and performance of Citi's data infrastructure.

Responsibilities

Build and maintain scalable data pipelines using PySpark within a Big Data environment to process and transform large volumes of structured and unstructured data.

Design and develop solutions across the Hadoop ecosystem - including Hive, HDFS, Sqoop, Spark, Impala, and Scala - to enable efficient data ingestion, processing, and storage.

Develop and manage real-time and batch data workflows using streaming data platforms, ensuring high availability and low-latency data delivery.

Write complex SQL queries to extract, validate, and analyze data across distributed systems, supporting data-driven decision-making.

Design and implement data models and data architecture patterns aligned with data warehouse principles, ensuring scalability, accuracy, and consistency.

Automate pipeline scheduling and orchestration using shell scripting and Autosys, reducing manual intervention and improving operational reliability.

Independently identify, assess, and resolve technical risks and data issues in a timely manner, maintaining system integrity across the data platform.

Required Qualifications & Skills

4 -7 years of relevant experience.

Hands-on expertise in PySpark and Big Data processing, with the ability to build and optimize distributed data workflows at scale.

Practical knowledge of the Hadoop ecosystem, including Hive, HDFS, Sqoop, Spark, Impala, and Scala, applied in a production environment.

Proficiency in complex SQL query development for data analysis, transformation, and validation across large datasets.

Solid understanding of distributed systems architecture and how data flows across interconnected processing layers.

Demonstrated knowledge of data modelling and data design, with familiarity in data warehouse concepts and dimensional modelling techniques.

Competence in shell scripting and job scheduling using Autosys or equivalent workflow automation tools.

Strong analytical and problem-solving ability, with a track record of working independently to diagnose and resolve complex data engineering challenges.

Clear and effective communication skills, with the ability to articulate technical concepts to both technical and non-technical audiences.

-

Job Family Group:

Technology

-

Job Family:

Applications Development

-

Time Type:

Full time

-

Most Relevant Skills

Please see the requirements listed above.

-

Other Relevant Skills

For complementary skills, please see above and/or contact the recruiter.

-

Questions about this role

Click "Apply with AI Applyd" above. We auto-fill the application from your resume and answer screening questions in seconds. No copy and paste, no juggling tabs.

Compensation for Data Engineer roles in India varies widely by seniority, employer size, and remote vs onsite arrangement. Check the salary range on this listing when published, or browse our Data Engineer hub for India medians across recent openings.

Most applications complete in under 90 seconds. You can track the status in your dashboard and watch the screenshot proof land the moment the application submits.

AI Applyd supports Greenhouse, Lever, Ashby, Workday, iCIMS, SmartRecruiters, Personio, Teamtailor and other major ATS platforms. If we can submit through the platform, we do.

Want AI Applyd to auto-apply to roles like this?

We tailor your resume per posting, fill the forms, and track replies for you.