Data Engineer – GCP Java & Big Data
Skills
About the role
Gurgaon, Bengaluru
Qualification
:
Job Summary:
We are looking for skilled GCP Data Engineers with 6–9 years of hands-on experience in building scalable, high-performance data solutions. The ideal candidate will have strong expertise in Java-based big data processing frameworks and deep exposure to Google Cloud Platform (GCP) services for modern data engineering workloads.
Key Skills & Experience:
Strong programming expertise in Java, with experience in building distributed data processing applications
Hands-on experience with Big Data technologies such as Apache Spark (Java/Scala APIs), Hadoop, and Hive
Experience with Spark (DataFrame/Spark SQL) using Java or Scala (PySpark knowledge is a plus but not primary)
Solid understanding of data structures, algorithms, and object-oriented programming in Java
Strong knowledge of SQL, data modeling, and data warehousing concepts
Experience working with Linux/Unix environments and scripting (Bash or similar)
Proven analytical and problem-solving skills, especially in debugging and optimizing data pipelines
Ability to design and build scalable, fault-tolerant data processing systems
Important to Have:
Hands-on experience with GCP services such as BigQuery, Dataflow (Apache Beam with Java), Dataproc, Cloud Storage, Pub/Sub, and IAM
Experience with workflow orchestration tools like Airflow or Cloud Composer
Exposure to cloud migration projects, especially transitioning from on-premise Hadoop ecosystems to GCP
Familiarity with streaming data pipelines using Pub/Sub and Dataflow
Understanding of CI/CD pipelines and DevOps practices in a cloud environment
Roles & Responsibilities:
Design, develop, and maintain scalable ETL/ELT pipelines using Java-based big data frameworks on GCP
Build and optimize batch and streaming data processing solutions using Dataproc, Dataflow, and Spark
Ensure high-quality, efficient, and maintainable code by following best practices and coding standards
Perform unit testing, integration testing, and troubleshoot complex data pipeline issues
Collaborate with cross-functional teams to understand data requirements and deliver robust solutions
Estimate development efforts and contribute to sprint planning and delivery
Participate in code reviews and mentor junior team members where required
Design cost-optimized and performance-efficient architectures leveraging GCP-native services
Skills Required
:
GCP, Pyspark, Java, SQL
Role
:
Skills : Java, Bigdata ,GCP
Core Java: Well versed with OOP, Data Structures, Generics, Collections, Basic Regular Expressions, IO, Basic Concurrency. Java Spring: Core, REST API's Knowledge of: Basics Shell scripting, Postman, JSON, MYSQL Big Data: Hadoop, Map reduce, Basic Spark, HBase(M7) GCP skillset, Big query
We are looking for a skilled Software Engineer / Data Engineer with strong expertise in Core Java, Big Data technologies, and GCP to design, develop, and maintain scalable data processing systems and microservices.
Primary Skills / Technical Expertise
Core Java
Strong knowledge of OOP concepts, Data Structures & Algorithms
Expertise in Collections, Generics, Exception Handling
Experience in Multithreading & Basic Concurrency
Hands-on with Java IO & file processing
Understanding of Regular Expressions
Java & Spring Framework
Experience in Spring Core, Spring Boot
Strong exposure to REST API design and development
Knowledge of Microservices architecture
Big Data Technologies
Hands-on experience with:
Hadoop Ecosystem
MapReduce
Apache Spark (Core & Basics of Spark SQL/PySpark)
Hive (good to have)
HBase or other NoSQL systems
Understanding of:
Distributed computing concepts
Batch & real-time data processing pipelines
Ability to handle large-scale data (GBs to TBs) (typical in Big Data roles) [Round 2 MI...Recording | Video]
GCP (Google Cloud Platform)
Hands-on experience with:
BigQuery
Cloud Storage (GCS)
Good to have:
Dataflow / Dataproc
Pub/Sub
Understanding of cloud-based data pipelines and deployment
Database / Tools
Strong knowledge of:
SQL (MySQL / Hive / BigQuery queries)
JSON handling & API integration
Tools:
Postman (API testing)
Shell scripting (basic automation)
Version control tools (Git – optional but preferred)
Key Responsibilities
Design and develop scalable data processing applications
Build and optimize ETL/data pipelines for large datasets
Develop and maintain REST APIs and microservices
Work on Big Data transformations and storage solutions
Integrate systems with GCP cloud services
Perform data validation, testing, and performance tuning
Collaborate with cross-functional teams for end-to-end delivery
Experience
:
6.0 to 9.0 years
Job Reference Number
:
13860
Questions about this role
Want AI Applyd to auto-apply to roles like this?
We tailor your resume per posting, fill the forms, and track replies for you.