Senior Data Engineer -GCP, Python, PySpark
Astra-North Infoteck Inc. ~ Conquering today’s challenges, achieving tomorrow’s vision!
Skills
About the role
Data Engineer – Google Cloud Platform (GCP), BigQuery, Python & Data Pipelines
Location: Hybrid – 3 Days per Week On-Site
Job Summary
• Skilled Data Engineer with hands-on Google Cloud Platform (GCP) experience to design, build, and maintain scalable data pipelines and cloud-based data solutions • Expertise in data warehousing, ETL/ELT development, big data technologies, and GCP services to support enterprise analytics and business intelligence initiatives
Key Responsibilities
• Design, develop, and maintain scalable batch and real-time data pipelines on GCP • Build and optimize data ingestion frameworks from multiple structured and unstructured data sources • Develop ETL/ELT processes using GCP-native services and modern data engineering tools • Design and implement data models for analytics, reporting, and machine learning use cases • Manage and optimize data storage solutions using BigQuery and Cloud Storage • Monitor data quality, performance, security, and governance standards • Collaborate with business stakeholders, data analysts, architects, and data scientists to deliver enterprise data solutions • Implement CI/CD, automation, and DevOps practices for data engineering workloads • Troubleshoot data processing issues and optimize pipeline performance • Ensure compliance with data security and privacy requirements
Required Skills
• Strong experience in Data Engineering and data pipeline development • Hands-on experience with Google Cloud Platform (GCP) services including: • BigQuery • Cloud Storage • Dataflow • Pub/Sub • Dataproc • Cloud Composer (Airflow) • Cloud Functions • Cloud Run • Strong programming skills in: • Python • SQL • PySpark • Experience with ETL/ELT frameworks and data integration tools • Strong knowledge of data warehousing concepts and dimensional modeling • Experience working with relational and NoSQL databases • Knowledge of real-time and streaming data architectures • Understanding of CI/CD pipelines, GitHub, and DevOps practices • Experience with data quality, metadata management, and governance
Preferred Skills
• Experience with Apache Spark, Hadoop, Kafka, and Airflow • Exposure to machine learning data pipelines and MLOps • Experience with Terraform or Infrastructure as Code (IaC) • Knowledge of containerization technologies such as Docker and Kubernetes • Experience in Banking, Financial Services, Insurance, or large enterprise environments
Required Qualifications
• Bachelor’s or Master’s degree in Computer Science, Information Technology, Engineering, or a related field • Excellent analytical and problem-solving skills • Strong communication and stakeholder management skills • Experience working in Agile/Scrum environments
Essential Skills
• Google Cloud Platform (GCP) • BigQuery • Python • SQL • PySpark • ETL/ELT • Dataflow • Cloud Composer (Airflow) • Cloud Storage • CI/CD
Questions about this role
Want AI Applyd to auto-apply to roles like this?
We tailor your resume per posting, fill the forms, and track replies for you.