Job Title: Nvidia Stack Specialist
Job Summary
We are looking for an experienced Infrastructure Operations Engineer to manage and support high-performance computing and storage systems, including NVIDIA A100 DGX servers, network
switches, and DDN storage appliances. The role involves ensuring optimal performance, availability, and reliability of critical infrastructure components within a hybrid physical and virtual environment.
Key Responsibilities
- Monitor resource management system (SLURM) to keep resource allocation efficient and
aligned with organizational priorities.
- Automate configuration management, software updates, and maintenance of system
availability using modern DevOps tools (Ansible, Salt, Gitlab, etc.)
- Plan and maintain new systems that support the NVIDIA Software stack.
- Work directly with developers and hardware architects to debug issues, identify new
requirements, and improve workflows.
- Provide technical support, administration, and monitoring of Linux systems, Nvidia DGX1
and A100 servers within a physical and virtual environment.
- Maintain scripts, security updates, patches, and configurations for the proper
functioning of servers.
- Coordinate with customers and stakeholders to collect data, conduct analysis, develop, and
implement solutions associated with incident tickets and requirements.
- Seek opportunities for continuous improvement to support effective and
- Work on NVIDIAs, SLURM, DGX, Superpod next-generation CPU Linux Administration Skills
- Verify micro-architecture/architecture features at the unit level or subsystem or full
chip test benches including FPGA/Silicon
with CPU architects in creating verifiable designs.
- Hands-on experience with HPC cluster job schedulers such as SLURM, LSF
- Hands-on experience with Software Defined Networking and HPC cluster networking.
- Hands-on experience with InfiniBand with IBOP and RDMA.
- Understanding of MLPerf benchmarking.
- Hands-on experience analysing and tuning performance for a variety of HPC workloads.
- Deploy, configure and manage NFS Storage + DDN Storage to Nvidia PODs.
Qualifications
- 5+ years of experience in infrastructure operations or HPC environments.
- Hands-on experience with NVIDIA DGX systems and DDN storage.
- Strong understanding and workings on below components
- NVIDIA
- A100 DGX servers
- Network
- switches
- AI400
- DDN storage systems
- DDN
- N6100 file storage
- BCM and UFM appliances
- Familiarity with BCM, UFM, and related monitoring tools.
- Excellent problem-solving and communication skills.
Preferred Certifications
- NVIDIA Certified Systems Engineer
- DDN Storage Certification
- Linux Professional Institute Certification (LPIC)
Flexible
for travel to KSA