Mid/Senior DevOps Engineer (Kubernetes, Linux)
SkyLab
Mô tả công việc
DevOps Engineer (Mid / Senior) Datacenter Infrastructure & Observability Location: Ho Chi Minh City, Vietnam (onsite) | Level: Mid / Senior | Type: Full-time | Experience: Mid 3+ years, Senior 5+ years About SkyLab SkyLab is a Singapore-headquartered AI infrastructure company operating across Southeast and North Asia. We design, build and operate GPU-as-a-Service (GPUaaS) environments on the latest NVIDIA platforms for AI model developers and enterprises, whose training and inference workloads depend on every node and link being healthy, 24 hours a day. About the role The DevOps team builds and runs the platform that keeps it that way. Our monitoring and automation stack watches hundreds of GPU servers across every site, covering hardware health, network fabric and alerting that gets the right people involved quickly. As new sites come online across the region, each one joins this platform from day one. You will join the team that owns it end to end. Everything is infrastructure-as-code, reviewed through merge requests, and runs on-premises rather than on public cloud. You will work closely with our Data Center Operations team and with engineers overseas, so clear written English is an important part of the job. Key responsibilities Observability: operate and improve a multi-site monitoring platform, where each site keeps working on its own and a central NOC gives fleet-wide visibility. Hardware and network health: use out-of-band server data and network telemetry to spot issues with GPUs, nodes and links early. Automation: automate provisioning, configuration and deployment as code, and keep the CMDB accurate so it drives discovery and alerting. AI-augmented operations: apply AI tools and agents to reduce manual work, such as alert triage, incident summaries and runbook automation, and use AI assistants to write and review code faster. Alerting: tune alerts to be clear and actionable, and keep the runbooks behind them up to date. Troubleshooting: investigate issues across hardware, network, OS and application layers, find the root cause, and share what was learned. Reliability and security: help maintain highly available services, backup and recovery plans, secrets, and regular patching. CI/CD and tooling: improve pipelines and build internal tools that make changes easier and safer for other teams. Collaboration: document your work and partner with Data Center Operations and overseas engineers on design and problem-solving. Our toolset: Prometheus/Grafana stack, Terraform, Ansible, Kubernetes (k3s), Proxmox, GitLab CI, NetBox (CMDB), Jira and PagerDuty.
Yêu cầu công việc
Requirements (must-have) Experience: Mid 3+ years in DevOps, SRE or infrastructure roles. Senior 5+ years, including running production systems. Strong Linux administration and troubleshooting skills. Hands-on experience with Ansible and Terraform. Experience with Docker and Kubernetes in production. Experience with Prometheus-based monitoring (PromQL, alert rules, exporters) and Grafana. Solid networking fundamentals: TCP/IP, DNS, TLS, load balancing, VLANs and basic routing. Scripting in Python and Bash. Virtualization experience, ideally Proxmox or another KVM-based hypervisor. Git-based workflow with code review and CI/CD, ideally GitLab CI. Comfortable communicating in English, written and spoken. Nice to have You don't need all of these. Any of them will help you ramp up faster. Server hardware and BMC management (Redfish, IPMI, iDRAC), ideally with GPU/HPC servers. Datacenter networking, especially Juniper. Thanos, Loki or Grafana at multi-site scale. NetBox or another CMDB. PostgreSQL high availability (Patroni, etcd) or object storage (MinIO, S3). Incident tooling (PagerDuty, Jira Service Management) and on-call experience. Secrets management (SOPS, Vault) and CVE remediation. Experience building or using AI agents and LLM-based tools in operations (AIOps). Public cloud (AWS or Azure), Go, or certifications such as CKA/CKAD. For Senior level, we also expect you to Lead the design of new components and site rollouts. Help set standards for alert quality, runbooks and code review. Take a lead role in incident response and follow-up. Mentor and support mid-level engineers. Soft skills Calm and methodical when things go wrong. Proactive, and comfortable working with teammates across time zones. Enjoys improving how things are done. Writes clearly, so others can follow your docs and runbooks.
Quyền lợi
Competitive package Professional working environment Opportunities to challenge and develop your career Social insurance, health insurance, unemployment insurance as labor law stipulated Premium Healthcare Opportunity to participate in stock option program. Public holiday and Annual leave in accordance with the Vietnamese labour law