Site Reliability Engineer

Techila Global Services

INremote countryPosted Aug 4, 2026
Posting intelligenceActively listed

Skills

cloudformationsalesforcefargatepythonswiftawsecsc#

About the role

Remote 5–9 yrs Apply by Aug 31, 2026

Site Reliability Engineer

Join Techila's team of Salesforce experts. We build senior-led transformations that deliver measurable outcomes for clients worldwide.

Experience

5–9 yrs

Employment Type

Full-time

Openings

1 position

Apply By

Aug 31, 2026

Required Skills

Site Reliability engineer

Job Description

Site Reliability Engineer Pune · Full Time · Remote

About the Role The Site Reliability Engineer will be responsible for ensuring the reliability and performance of .NET applications and AWS infrastructure. The role involves partnering with application engineers, embedding reliability into new feature design, and optimizing infrastructure for consistent and reproducible environments. Success in this role means driving data-driven reliability improvements, reducing toil, and freeing the team for higher-impact engineering work.

Key Responsibilities

Read, debug, and contribute to production C#/.NET code to diagnose and fix app-level reliability issues.

Identify and resolve memory leaks, thread pool exhaustion, and GC pressure before they manifest as incidents.

Partner with application engineers to embed reliability into new feature design and deployment practices.

Instrument .NET services with distributed tracing and structured logging to surface runtime anomalies early.

Operate and optimize EC2 Auto Scaling, ECS Fargate, and Lambda workloads.

Build and maintain infrastructure-as-code using CloudFormation or CDK for consistent, reproducible environments.

Automate operational tasks, deployment pipelines, and disaster recovery procedures.

Manage RDS SQL Server deployments including Multi-AZ failover configuration and read replica setup.

Diagnose and resolve performance issues: slow queries, missing indexes, and blocking chains.

Build and maintain observability stacks using CloudWatch metrics, log insights, and alarms; AWS X-Ray for distributed tracing.

Own service health dashboards, SLOs/SLIs, and drive data-driven reliability improvements.

Design alerts that surface signal - not noise - and ensure on-call responders have the context to act quickly.

Conduct root cause analysis (RCA) on incidents and lead blameless post-mortems to capture lessons and prevent recurrence.

Participate in on-call rotation to respond to production incidents and drive swift resolution.

Requirements

Have 4+ years of experience in SRE, DevOps, platform engineering, or a systems-focused software engineering role.

Possess C#/.NET engineering ability - can read, debug, and contribute to production code; experience diagnosing memory leaks, thread exhaustion, and GC pressure.

Have AWS compute fluency: hands-on depth across EC2 Auto Scaling, ECS Fargate, and Lambda, with informed opinions on when to use each.

Have RDS SQL Server operational experience: Multi-AZ failover, read replicas, backup/PITR, slow query analysis, and blocking chain resolution.

Have native AWS observability proficiency: CloudWatch (metrics, logs, alarms), X-Ray, and infrastructure-as-code via CloudFormation or CDK.

Have AWS networking and security competence: VPCs, security groups, ALB/NLB, Route 53, TLS/ACM, and least-privilege IAM.

Have SLO discipline: experience defining SLIs/SLOs against real metrics, running blameless postmortems, and carrying an on-call pager.

Have strong scripting ability (PowerShell, Python, or Bash) for automation and operational tooling.

Have excellent communication skills and a collaborative, blameless engineering mindset.

Good to Have

Genuine openness to adopting AI tools and a willingness to experiment with new technology to work smarter and faster.

What We Offer

Opportunity to work with a collaborative and blameless engineering team.

Chance to drive data-driven reliability improvements and reduce toil.

Freedom to experiment with new technology and tools to work smarter and faster.

At a Glance

Work ModeRemote

EmploymentFull-time

Experience5–9 yrs

Openings1

DeadlineAug 31, 2026

[ Hiring process ]

What to expect

Four stages, typically completed within 2–3 weeks. We respect your time - every stage has a clear purpose and timely feedback.

STEP 0130 min

Screening Call

Introductory conversation with our talent team to understand your background and motivations.

STEP 0260–90 min

Technical Round

Live problem-solving with a senior architect on Salesforce design, integrations, or domain depth.

STEP 0345 min

Culture Fit

Conversation with practice leadership covering working style, ownership, and how you collaborate.

STEP 04Within 5 days

Offer

Formal offer with full compensation breakdown, start date, and onboarding plan.

Questions about this role

Click "Apply with AI Applyd" above and you are done. Your resume is rewritten for this advert, the screening questions are answered, and it is submitted on Techila Global Services's own hiring system. No retyping your history, no fourteen tabs, no evening lost.

Compensation for DevOps / SRE roles in India varies widely by seniority, employer size, and remote vs onsite arrangement. Check the salary range on this listing when published, or browse our DevOps / SRE hub for India medians across recent openings.

You never touch the form - the application is filled and submitted for you on Techila Global Services's own hiring system. It is not marked sent when we press submit. It is marked sent when a confirmation from their system arrives at the address we apply with, and your dashboard shows which stage each application is at until then.

Twelve applicant tracking systems have a real apply path: Workday, Greenhouse, Lever, Ashby, Workable, iCIMS, Personio, Recruitee, Teamtailor, Rippling, Breezy and SmartRecruiters. Your application goes in on the employer's own hiring system, never into an aggregator queue.

Want AI Applyd to auto-apply to roles like this?

We tailor your resume per posting, fill the forms, and track replies for you.