HPC Infrastructure Administrator Job at NVIDIA, Santa Clara, CA

MG1Eckp0WitxWW5aSXQ3RTBDbHFRUWRUVXc9PQ==
  • NVIDIA
  • Santa Clara, CA

Job Description

HPC Infrastructure Administrator Location Santa Clara, CA : We are now seeking a HPC Infrastructure Engineer! NVIDIA's Compute Architecture Group is growing our team of HPC Infrastructure Engineers who run our internal cluster for accelerated AI and HPC software development. As part of this team, you will help to manage a diverse cluster of GPU-accelerated systems. Your contributions will enable engineers to work efficiently with a wide variety of forward-looking hardware configurations as they vigilantly seek out opportunities for performance optimization and continuously deliver high quality software. Our ideal candidate is versatile enough to apply expertise from many domains: system administration, performance analysis, automation, and architecture. Your work will enable the ground breaking experimentation that allows us to design the world's most powerful systems for the most demanding computing applications. You will have a meaningful impact at a fast-moving company that is spearheading the next wave in computing technology. Join our technically diverse team of GPU architects, software engineers and infrastructure experts to unlock unprecedented performance in every domain! What you'll be doing:
  • Administer an HPC cluster composed of Linux systems ranging from the world's most powerful servers to embedded systems
  • Maintain the configuration of our resource management system (SLURM) to keep resource allocation efficient and aligned with organizational priorities
  • Automate configuration management, software updates, and maintenance of system availability using modern DevOps tools (Ansible, Gitlab, etc.)
  • Plan and maintain new systems that support the NVIDIA Software stack
  • Work directly with developers and hardware architects to debug issues, identify new requirements, and improve workflows
  • Actively communicate with users and management regarding resource planning and allocation
  • What we need to see:
    • 5+ years of previous experience deploying and administering HPC clusters
    • BA, BS, or MS in CS, EE, CE or equivalent experience
    • Deep knowledge of distributed resource scheduling systems (Slurm (preferred), LSF, etc.)
    • Demonstrated ability to script in bash, and at least one high-level language (Python preferred)
    • Experience with container technologies (Docker, Singularity, etc.)
    • Deep understanding of operating systems, computer networks, and high-performance hardware
    • Ability to work well with developers, hardware architects, & test engineers
    • Passionate dedication to providing quality support for users
    Ways to stand out from the crowd:
    • Prior work experience managing high performance fabrics and parallel file systems
    • Familiarity with CUDA and managing GPU-accelerated computing systems
    • Basic knowledge of deep learning frameworks and algorithms
    The base salary range is 118,400 USD - 224,250 USD. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. You will also be eligible for equity and benefits . NVIDIA accepts applications on an ongoing basis. NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Job Tags

Full time, Work experience placement,

Similar Jobs

Amtrak

Mgr Material Control - 90403717 - Washington Job at Amtrak

 ...Your success is a train ride away! As we move Americas workforce toward the future, Amtrak connects businesses and communities across the country. We employ more than 20,000 diverse, energetic professionals in a variety of career fields throughout the United States... 

CEDENT

Angular.js Developer - Senior UI (Chicago, IL) Job at CEDENT

 ...including core concepts, components, modules, and services. Experience with Angular 16, 17, or 18- the newer, the better Solid understanding of HTML, CSS, and JavaScript Experience integrating with OIDC systems is a must. OIDC = OpenID Connect ************... 

DSV - Global Transport and Logistics

Freight Forwarder, Gateway Exports Job at DSV - Global Transport and Logistics

 ...is local and close to our customers. Read more at Location: Grapevine, TX Division:Air & Sea Job Posting Title: Freight Forwarder, Gateway Exports Time Type: Full Time Summary An Air Export Gateway Freight Forwarder will be responsible for managing... 

Express Employment Professionals - Lynchburg

MIG Welder Job at Express Employment Professionals - Lynchburg

 ...specifications Inspect welds to ensure quality and consistency Maintain a clean and safe work area Follow all safety and production procedures Pay and Schedule Pay: $18.00 to $25.00 per hour Schedule: Monday through Thursday, 7:00 AM to 5:30 PM, no Fridays... 

SP Associates

1359-Extrusion Polyurethane Process Engineer Job at SP Associates

 ...Extrusion Polyurethane Process Engineer- Charleston, SC .Travel will be required to India 40-50%.Salary is open with 8-10 years of extrusion process improvement experience required.Great benefits!!Relocation is available.Bonus plan will be included. #1359 Requirements...