Senior DevOps Engineer
Skills
About the role
To get the best candidate experience, please consider applying for a maximum of 3 roles within 12 months to ensure you are not duplicating efforts.
Job Category
Software Engineering
Job Details
About Salesforce
Salesforce is the #1 AI CRM, where humans with agents drive customer success together. Here, ambition meets action. Tech meets trust. And innovation isn’t a buzzword - it’s a way of life. The world of work as we know it is changing and we're looking for Trailblazers who are passionate about bettering business and the world through AI, driving innovation, and keeping Salesforce's core values at the heart of it all.
Ready to level-up your career at the company leading workforce transformation in the agentic era? You’re in the right place! Agentforce is the future of AI, and you are the future of Salesforce.
The Salesforce Cloud Operations Team consists of smart and tenacious engineers dedicated to excellence through platform expertise and uncompromising integrity. While we understand that delivery is important, we consider quality to be our highest priority. The code we deliver must be functional, scalable, secure, and maintainable.
Our team is responsible for the Cloud Infrastructure Platform, covering both first-party and public cloud environments. As our platform grows in size and complexity, we must consistently provide a top-tier platform for running the business today and support tremendous scale in the near future. We are looking to grow our team to develop and manage our next generation of software/hardware systems designed to maintain our world-class data platform.
As a Senior Operations Engineer, you will contribute to the DevOps culture on the global team by enhancing our automation and CI/CD toolset to orchestrate thousands of servers in perfect harmony. We are agile, fast-paced, work hard, and play hard.
Your Impact
Build scalable, elastic services, capable of running in private Salesforce data centers and in public clouds
Design high-quality systems to improve platform security, reliability, availability, and scalability
Design, build, configure, test, update, document, implement scripts, and automate operational support in multiple infrastructure environments at world-scale
Collaborate with development, architecture, and release teams, providing architectural design recommendations and driving standards for effective transition to production operations.
Build efficient components/algorithms/services to serve a high volume of requests with low latencies. Conduct performance tuning, performance monitoring, capacity planning, and other related monitoring.
Create, document, and test backup and recovery procedures including Operational and Disaster Recovery scenarios
Design, document, implement smart thorough testing and environmental upgrade strategies to ensure quality delivery
Continuously integrate and deploy your code, monitor live services and supporting hardware, and set up alerting mechanisms to ensure system health and availability. Learn through data-driven feedback loops to improve it
Participate in Code Reviews, cross train and mentor other team members
Participate in the team's on-call rotation to address complex problems in real-time and keep services operational and highly available, but have a life outside of work - we value balance.
Be able to upscale your team and other teams in the same organization proactively with ideas and learnings from experience in the relevant topics
Participate in and lead the development of team/organization long range plans and roadmaps
Demonstrate strong technical expertise in your team's domain area and focus on the most challenging deliveries
Advocate for strategic improvements in response to changing service conditions
Synthesize information coming from outside your organization (architectural guidelines, workgroup results, best practices) and guide your team to put this information into practice
Basic Qualifications
BS degree and 6+ years of prior relevant experience or Master's with 5+ years of prior relevant experience
Strong knowledge of basic programming concepts and respective scripting languages. One or more of the following: Bash, Go, Ruby, Python
Broad knowledge of Linux Operating Systems internals and Virtualization platforms
Strong experience in automating a variety of infrastructure and administration tasks
Working Knowledge of CI/CD, DevOps, Configuration Management and Infrastructure as Code principles
Experience with Source Code Management concepts and tools such as Git
Experience working with and troubleshooting large, highly available and high performance systems
Experience with Agile development methodology
Be available for after hours on-call support and will be required to apply production packages during non-peak hours
Experience with leading projects that involve more than the immediate scrum team
Preferred Qualifications
Experience with monitoring solutions, such as - Prometheus, Grafana, or Splunk
Ability to manage system issues, updates and upgrades for both first-party and public cloud
Experience with public cloud development (e.g. AWS, GCP, Alibaba, and/or Azure platforms)
Experience with containers (Docker) and related orchestration (Kubernetes)
Experience with Terraform, Github, Cinc (Chef), Jenkins, Spinnaker
Knowledge of data concepts related to relational databases and NoSQL stores.
Strong infrastructure engineering background
Experience with integrating AI for automated anomaly detection, predicting system degradation, triggering self-healing workflows, reducing alert fatigue etc.
Unleash Your Potential
When you join Salesforce, you’ll be limitless in all areas of your life. Our benefits and resources support you to find balance and be your best , and our AI agents accelerate your impact so you can do your best . Together, we’ll bring the power of Agentforce to organizations of all sizes and deliver amazing experiences that customers love. Apply today to not only shape the future - but to redefine what’s possible - for yourself, for AI, and the world.
Accommodations
If you need a reasonable accommodation during the application or the recruiting process, please submit a request via this Accommodations Request Form .
Please note that Salesforce uses artificial intelligence (AI) tools to help our recruiters assess and evaluate candidates’ resumes and qualifications throughout the recruiting process. Humans will always make any candidate selection and hiring decisions. Please see our Candidate Privacy Statement for more information about how we use your personal data and your rights, including with regard to use of AI tools and opt out options.
Posting Statement
Questions about this role
Want AI Applyd to auto-apply to roles like this?
We tailor your resume per posting, fill the forms, and track replies for you.