AI Infrastructure Engineer GPU & Colocation
Công ty TNHH Sapawoo
Mô tả công việc
We are building our own GPU infrastructure to run AI models, support AI-driven software development, and power AI capabilities in our software products. We are looking for a hands-on engineer to build this infrastructure in a colocation data centre and take responsibility for its ongoing operation. You will turn workload requirements into a practical infrastructure setup—from selecting servers and coordinating installation to configuring Linux, deploying model-serving environments, and keeping systems secure and reliable. Your Responsibilities Plan and Build the Infrastructure • Translate AI workload requirements into GPU, compute, storage, and networking configurations. • Evaluate hardware compatibility, supplier proposals, and costs. • Coordinate procurement, delivery, installation, and commissioning. • Confirm rack space, power, cooling, connectivity, and access requirements with the colocation provider. • Install and configure servers, networking equipment, cabling, and remote management. • Maintain infrastructure documentation and asset inventories. Configure GPU and AI Systems • Configure and maintain Linux servers, GPU drivers, CUDA environments, and container runtimes. • Deploy and maintain model-serving infrastructure in collaboration with the AI and engineering teams. • Diagnose GPU, memory, storage, and network performance issues. • Benchmark workloads and improve resource utilisation. • Automate deployment and configuration to make environments reproducible. Manage Networking and Security • Configure network segmentation, firewalls, VPNs, and secure administrative access. • Implement access controls, credential management, system hardening, and patching. • Separate workloads and environments according to product and security requirements. • Troubleshoot connectivity issues with infrastructure providers. Own Reliability and Operations • Establish monitoring and alerting for hardware health, GPU utilisation, capacity, and service availability. • Implement backup and recovery procedures and test restoration. • Maintain operational runbooks and troubleshoot infrastructure incidents. • Coordinate hardware replacements, maintenance windows, and remote-hands support. • Identify single points of failure and propose practical improvements. • Establish clear support coverage and escalation procedures with the team. Manage Capacity and Costs • Track utilisation, operating costs, and capacity constraints. • Recommend upgrades based on measured demand and performance. • Assess trade-offs between owned hardware, rented GPU capacity, and cloud services. • Plan expansion while avoiding unnecessary complexity and overprovisioning.
Yêu cầu công việc
What You Bring • Hands-on experience deploying and operating physical servers in a data centre or colocation environment. • Strong Linux administration and troubleshooting skills. • Practical experience with GPU servers, NVIDIA drivers, CUDA, and AI workloads. • Solid networking knowledge, including switching, routing, VLANs, firewalls, and VPNs. • Experience with containers, monitoring, backups, and infrastructure automation. • An understanding of rack power, cooling, hardware compatibility, and remote server management. • A disciplined approach to security, documentation, maintenance, and recovery. • The ability to work independently and coordinate with vendors and data centre providers. • Professional English for technical documentation and collaboration. • Willingness to perform on-site installation and maintenance when required. Useful Additional Experience • Operating AI inference and model-serving systems. • Managing multi-GPU workloads and high-speed networking. • Using infrastructure-as-code and configuration-management tools. • Supporting environments with workload isolation and availability requirements. How You Will Work • You will own the implementation and reliable operation of our AI infrastructure, working closely with the AI and software engineering teams to align infrastructure with application and model requirements. • This is a hands-on role covering both the initial build and ongoing operations, with direct involvement in model deployment, performance optimisation, capacity planning, and infrastructure decisions. What Success Looks Like • Infrastructure meets agreed workload, security, and operational requirements. • AI workloads run reliably with measurable performance and utilisation. • Monitoring, tested recovery procedures, and operational documentation are in place. • Infrastructure costs are transparent and expansion decisions are based on evidence. • Incidents are resolved effectively, with clear coordination across suppliers and service providers.
Quyền lợi
Theo quy định công ty