Senior Deveops Platform Engineer
Skills
About the role
We are seeking a Senior / Principal DevOps / Platform Engineer to own and operate the full infrastructure lifecycle for a production-grade, cloud-native platform operating under strict security and compliance requirements. This is an end-to-end platform ownership role. You will design and provision cloud infrastructure, build and maintain CI/CD pipelines, support production operations and continuously improve reliability, scalability and security. The platform supports modern application workloads running on Kubernetes, integrates with managed data services, and services business-critical use cases. The candidate will work closely with senior architects and engineers in a small, experienced team. This role has direct responsibility for production systems with minimal bureaucracy between engineering and operations.
Responsibilities
Cloud Infrastructure & Platform Engineering
Design, provision, and manage cloud infrastructure across production, staging, and development environments.
Operate and optimize managed Kubernetes clusters, including:
Node pools
Autoscaling
Cluster upgrades
Maintenance windows
Pod Disruption Budgets (PDBs)
Maintain Kubernetes manifests using Kustomize.
Implement cloud-native workload identity and eliminate credential-based authentication.
Configure and manage networking components including:
Virtual Networks (VNets)
Subnets
Security Groups
Private Endpoints
DNS
Network Peering
CI/CD & Release Engineering
Design, build, and maintain GitHub Actions CI/CD pipelines.
Implement secure branch protection and environment promotion strategies.
Manage container image publishing and artifact versioning.
Enforce security scanning, code quality gates, and container vulnerability checks.
Maintain release workflows and controlled deployment processes.
Create local development environments aligned with production configurations.
Observability
Own platform monitoring, logging, and distributed tracing.
Implement OpenTelemetry instrumentation.
Build dashboards for:
Platform health
API performance
Database metrics
Cache efficiency
CI/CD pipeline performance
Define alerting strategies and incident response workflows.
Maintain long-term log retention for compliance.
Develop production troubleshooting documentation and runbooks.
Secrets & Configuration Management
Manage centralized secrets using HSM-backed key management.
Maintain Kubernetes ConfigMaps and Secrets through Git-based workflows.
Implement managed identities and role-based access control (RBAC).
Ensure strict separation of environments and secure configuration management.
Container Platform Management
Build and maintain hardened enterprise container images.
Optimize Docker builds using multi-stage builds and caching.
Enforce immutable production containers.
Manage enterprise container registries, retention policies, and cleanup processes.
Database & Data Infrastructure
Administer production PostgreSQL databases.
Implement:
Connection pooling
Secure authentication
Zero-downtime migrations
Backup & Point-in-Time Recovery (PITR)
Manage Redis clusters for caching and session management.
Configure cloud object storage lifecycle and access policies.
Support data orchestration workflows and secure analytics platform connectivity.
Security & Compliance
Configure Web Application Firewalls (WAF) and edge security services.
Manage API Gateway authentication, authorization, throttling, and monitoring.
Enforce Kubernetes Network Policies.
Conduct infrastructure and container vulnerability assessments.
Maintain encryption, audit logging, and compliance controls.
Participate in security architecture reviews and threat modeling.
Observability & Monitoring
Own platform monitoring, logging, and distributed tracing.
Implement OpenTelemetry instrumentation.
Build dashboards for:
Platform health
API performance
Database metrics
Cache efficiency
CI/CD pipeline performance
Define alerting strategies and incident response workflows.
Maintain long-term log retention for compliance.
Develop production troubleshooting documentation and runbooks.
Required Skills & QualificationsPlatform Engineering
5+ years of experience in DevOps, Platform Engineering, or Cloud Infrastructure.
Strong hands-on experience with Kubernetes in production environments.
Extensive experience with Microsoft Azure (AWS knowledge is a plus).
Expertise in Infrastructure-as-Code using Terraform, ARM Templates, or Bicep.
Experience with Git-based infrastructure management.
CI/CD & Automation
Strong experience with GitHub Actions.
Container image management and enterprise registries.
Secure CI/CD implementation.
Automated testing, scanning, and release pipelines.
Cloud & Containers
Docker and Kubernetes expertise.
Container security best practices.
Registry lifecycle management.
Production container optimization.
Databases
PostgreSQL administration.
Redis deployment and optimization.
Database migrations, backups, and high availability.
Data orchestration support.
Observability
Monitoring and logging platforms.
OpenTelemetry implementation.
Production incident troubleshooting.
SLO/SLA monitoring and reliability engineering.
Security & Networking
WAF/CDN implementation.
API Gateway management.
Kubernetes Network Policies.
Secrets management and identity services.
Infrastructure security and compliance.
Programming & Automation
Python automation.
Bash/Shell scripting.
YAML and JSON.
Strong troubleshooting and debugging skills.
Preferred Qualifications
Experience working with regulated or compliance-driven environments.
Hands-on experience with GitOps tools such as ArgoCD or Flux.
Helm chart development.
Identity and Access Management (IAM) solutions.
Advanced Kubernetes troubleshooting.
RBAC design and security hardening.
Experience integrating analytics and data platforms.
Day-to-Day Responsibilities
Monitor and support production infrastructure.
Resolve infrastructure incidents and platform issues.
Review and merge infrastructure pull requests.
Plan and execute platform upgrades and maintenance.
Design scalable cloud platform enhancements.
Participate in architecture and security reviews.
Develop operational documentation, runbooks, and post-incident reports.
Mentor engineering teams on DevOps, Kubernetes, and cloud best practices.
Continuously improve engineering standards, automation, and operational excellence.
Work Location: Hybrid remote in Noida, Uttar Pradesh (Noida)
Questions about this role
Want AI Applyd to auto-apply to roles like this?
We tailor your resume per posting, fill the forms, and track replies for you.