Asso. Manager, SRE (5+yrs, English)
YUM! Digital & Technology
Mô tả công việc
Team Leadership and Development Lead and develop a team of Site Reliability Engineers (Level 6-7), initially 3 direct reports and growing as the Vietnam site consolidates, owning performance, coaching, retention, and day-to-day execution Build individual development plans that grow Level 6 engineers toward independent Level 7 scope Establish a high-ownership culture where engineers are accountable for outcomes, not tasks Run team rituals: 1:1s, shift retrospectives, and development check-ins Follow-the-Sun Operations Own the Vietnam shift within GRE’s global follow-the-sun coverage model, including schedule design, coverage planning, and holiday/leave management Ensure clean, structured handoffs to and from US and India teams, with clear ownership transfer on open incidents and in-flight work Maintain shift readiness: runbooks current, alerts actionable, escalation paths clear Serve as escalation point for the Vietnam shift during complex or high-severity incidents Reliability Practice Execution Own production reliability outcomes for the markets, platforms, and services within the Vietnam team’s scope Drive SRE operating standards within the team: incident response rigor, SLO ownership, runbook quality, and post-incident follow-through Ensure monitoring coverage, dashboards, and alerting remain accurate and effective across owned services Enforce GRE-wide reliability standards, ensuring the Vietnam practice operates in alignment with the global model Platform Engineering, Automation, and AI Own the Vietnam team’s contribution to GRE’s reliability modernization roadmap: auto-healing, auto-remediation, and self-service issue mitigation Treat automation as a delivery commitment, not a byproduct: plan, prioritize, and track toil-reduction and self-service work alongside operational coverage Ensure the team builds with platform engineering practices: infrastructure as code, runbooks as code, GitOps workflows, and reusable tooling over one-off fixes Champion AI-first engineering practices within the team, ensuring engineers develop with and through modern AI tooling Partner with GRE’s Foundations and Intelligence pillars on observability, signal detection, and automated response infrastructure Stakeholder Engagement Represent the Vietnam team in East GRE planning and reliability forums Communicate reliability outcomes, coverage status, and team health clearly to GRE leadership Partner with the Associate Director on headcount planning, team scope, and Vietnam-specific delivery
Yêu cầu công việc
5+ years in site reliability engineering, DevOps, infrastructure, or production operations roles 1+ years of people management experience, or 2+ years as a senior technical lead with demonstrated coaching and delivery ownership Hands-on credibility across incident response, observability, and automation, with the technical depth to guide Level 6-7 engineers Experience operating in shift-based, on-call, or follow-the-sun coverage models Working knowledge of at least one major cloud provider (AWS preferred) and modern observability tooling (e.g., Datadog, Prometheus, Grafana) Proficiency in at least one scripting or programming language sufficient to review and guide automation work Understanding of SLI/SLO frameworks and reliability engineering fundamentals Strong written and verbal English communication skills for cross-region collaboration with US and India teams Preferred Requirements Experience building or standing up a new team, site, or shift operation Experience managing engineers across the early-to-mid career range with a track record of promotions or level progression Kubernetes, container orchestration, and infrastructure as code experience (e.g., Terraform) Familiarity with AI-assisted operations tooling and automation-first reliability approaches, including auto-healing and auto-remediation patterns Exposure to platform engineering and internal developer platform concepts: self-service tooling, developer portals (e.g., Port, Backstage), GitOps Experience in multi-region or globally distributed team models Relevant certifications (AWS, CKA, or similar)
Quyền lợi
100% salary during the probation period 18 days of annual leave per year Five “Recharge Days” – extra days in addition to company holidays One day of paid leave for your birthday 2 days WFH/ week Half-day Fridays Full salary insurance 13th-month bonus Advanced health insurance (Generali) Regular team and engagement activities LinkedIn learning and training courses, based on company policy MacBook and monitors