
AgentOps Engineer – AI Managed Services
Skills
About the role
Location: Singapore
About the Role
Keep AI-powered services running at scale. As an AgentOps Engineer, you will manage the reliability, performance, and operational health of enterprise AI agents in production. You will play a critical role in monitoring agent behavior, optimizing costs, managing releases, and ensuring AI solutions remain secure, reliable, and business-ready after go-live.
Key Responsibilities
Operate, monitor, and support AI agents, copilots, and Agentic Managed Services platforms in production environments.
Design and implement observability, monitoring, incident response, release management, and operational support frameworks for AI solutions.
Monitor agent accuracy, latency, availability, exception rates, user feedback, token consumption, and operational performance.
Manage prompt, workflow, and agent releases, including rollback strategies, production controls, and readiness checks.
Support AI FinOps activities through usage tracking, cost visibility, performance optimization, and consumption reporting.
Develop operational runbooks, reliability standards, and lifecycle controls to ensure scalable and supportable AI services.
Partner with Agentic Engineers, AI Quality Engineers, Solution Architects, and AMS teams to continuously improve reliability, security, and cost efficiency.
Required Qualifications
Proven experience in DevOps, Site Reliability Engineering (SRE), or production operations.
Experience with monitoring, observability, and operational support tools.
Experience managing incidents, releases, production support, and service operations.
Strong understanding of reliability, performance, automation, and operational excellence practices.
Excellent problem-solving and stakeholder management skills.
Preferred Qualifications
Experience supporting AI, Generative AI, copilots, agents, or intelligent automation platforms.
Knowledge of cloud platforms such as AWS, Microsoft Azure, or Google Cloud Platform (GCP).
Experience with CI/CD pipelines, Infrastructure as Code (Terraform, Bicep, or equivalent), and platform engineering practices.
Familiarity with AI observability, AI FinOps, cost optimization, and performance monitoring.
Experience working with enterprise platforms such as SAP, Salesforce, ServiceNow, or SuccessFactors.
Questions about this role
Want AI Applyd to auto-apply to roles like this?
We tailor your resume per posting, fill the forms, and track replies for you.