Site Reliability Engineer
Skills
About the role
Remote 5–9 yrs Apply by Aug 31, 2026
Site Reliability Engineer
Join Techila's team of Salesforce experts. We build senior-led transformations that deliver measurable outcomes for clients worldwide.
Experience
5–9 yrs
Employment Type
Full-time
Openings
1 position
Apply By
Aug 31, 2026
Required Skills
Site Reliability engineer
Job Description
Site Reliability Engineer Pune · Full Time · Remote
About the Role The Site Reliability Engineer will be responsible for ensuring the reliability and performance of .NET applications and AWS infrastructure. The role involves partnering with application engineers, embedding reliability into new feature design, and optimizing infrastructure for consistent and reproducible environments. Success in this role means driving data-driven reliability improvements, reducing toil, and freeing the team for higher-impact engineering work.
Key Responsibilities
Read, debug, and contribute to production C#/.NET code to diagnose and fix app-level reliability issues.
Identify and resolve memory leaks, thread pool exhaustion, and GC pressure before they manifest as incidents.
Partner with application engineers to embed reliability into new feature design and deployment practices.
Instrument .NET services with distributed tracing and structured logging to surface runtime anomalies early.
Operate and optimize EC2 Auto Scaling, ECS Fargate, and Lambda workloads.
Build and maintain infrastructure-as-code using CloudFormation or CDK for consistent, reproducible environments.
Automate operational tasks, deployment pipelines, and disaster recovery procedures.
Manage RDS SQL Server deployments including Multi-AZ failover configuration and read replica setup.
Diagnose and resolve performance issues: slow queries, missing indexes, and blocking chains.
Build and maintain observability stacks using CloudWatch metrics, log insights, and alarms; AWS X-Ray for distributed tracing.
Own service health dashboards, SLOs/SLIs, and drive data-driven reliability improvements.
Design alerts that surface signal - not noise - and ensure on-call responders have the context to act quickly.
Conduct root cause analysis (RCA) on incidents and lead blameless post-mortems to capture lessons and prevent recurrence.
Participate in on-call rotation to respond to production incidents and drive swift resolution.
Requirements
Have 4+ years of experience in SRE, DevOps, platform engineering, or a systems-focused software engineering role.
Possess C#/.NET engineering ability - can read, debug, and contribute to production code; experience diagnosing memory leaks, thread exhaustion, and GC pressure.
Have AWS compute fluency: hands-on depth across EC2 Auto Scaling, ECS Fargate, and Lambda, with informed opinions on when to use each.
Have RDS SQL Server operational experience: Multi-AZ failover, read replicas, backup/PITR, slow query analysis, and blocking chain resolution.
Have native AWS observability proficiency: CloudWatch (metrics, logs, alarms), X-Ray, and infrastructure-as-code via CloudFormation or CDK.
Have AWS networking and security competence: VPCs, security groups, ALB/NLB, Route 53, TLS/ACM, and least-privilege IAM.
Have SLO discipline: experience defining SLIs/SLOs against real metrics, running blameless postmortems, and carrying an on-call pager.
Have strong scripting ability (PowerShell, Python, or Bash) for automation and operational tooling.
Have excellent communication skills and a collaborative, blameless engineering mindset.
Good to Have
Genuine openness to adopting AI tools and a willingness to experiment with new technology to work smarter and faster.
What We Offer
Opportunity to work with a collaborative and blameless engineering team.
Chance to drive data-driven reliability improvements and reduce toil.
Freedom to experiment with new technology and tools to work smarter and faster.
At a Glance
Work ModeRemote
EmploymentFull-time
Experience5–9 yrs
Openings1
DeadlineAug 31, 2026
[ Hiring process ]
What to expect
Four stages, typically completed within 2–3 weeks. We respect your time - every stage has a clear purpose and timely feedback.
STEP 0130 min
Screening Call
Introductory conversation with our talent team to understand your background and motivations.
STEP 0260–90 min
Technical Round
Live problem-solving with a senior architect on Salesforce design, integrations, or domain depth.
STEP 0345 min
Culture Fit
Conversation with practice leadership covering working style, ownership, and how you collaborate.
STEP 04Within 5 days
Offer
Formal offer with full compensation breakdown, start date, and onboarding plan.
Questions about this role
Want AI Applyd to auto-apply to roles like this?
We tailor your resume per posting, fill the forms, and track replies for you.