AMD is seeking an AI Research Infrastructure Engineer to operate, scale, and continually improve the shared GPU and HPC compute platform that supports AI, ML, and HPC research. You will own the day-to-day health of the SLURM and GPU clusters and work hands-on with researchers to run demanding workloads—including large-scale multi-GPU and multi-node training—reliably and efficiently. This is a research-enablement role, not a traditional systems-administration position, and requires research literacy to partner with researchers as a technical peer and accelerate their work. The role covers internal and externally visible compute clusters across the research organization, including AMD's university program clusters and collaboration with external academic partners. You may coordinate contractors and systems administrators supporting the environment. Familiarity with agentic engineering workflows—tools such as Claude Code, Codex, or Cursor—is expected.
AMD is seeking an AI Research Infrastructure Engineer to operate, scale, and continually improve the shared GPU and HPC compute platform that supports AI, ML, and HPC research. You will own the day-to-day health of the SLURM and GPU clusters and work hands-on with researchers to run demanding workloads—including large-scale multi-GPU and multi-node training—reliably and efficiently. This is a research-enablement role, not a traditional systems-administration position, and requires research literacy to partner with researchers as a technical peer and accelerate their work. The role covers internal and externally visible compute clusters across the research organization, including AMD's university program clusters and collaboration with external academic partners. You may coordinate contractors and systems administrators supporting the environment. Familiarity with agentic engineering workflows—tools such as Claude Code, Codex, or Cursor—is expected.
Key responsibilities: - Own day-to-day operations of the SLURM-managed GPU and HPC clusters, ensuring high availability, utilization, and performance across a multi-user research environment. - Partner directly with researchers to run and optimize demanding multi-GPU and multi-node workloads. - Operate and improve the broader compute platform—shared storage, networking, containers, monitoring—and build automation and self-service workflows that reduce friction for researchers. - Support GPU platform and hardware bring-up: validation, enablement, debugging, and operational readiness. - Manage AMD's university program clusters and
Track similar jobs
Get email alerts when new roles like this are posted.
Based on: AI Research Infrastructure Engineer
See how this role matches your resume
Sign in to get an AI match score and personalized feed.
Key responsibilities: - Own day-to-day operations of the SLURM-managed GPU and HPC clusters, ensuring high availability, utilization, and performance across a multi-user research environment. - Partner directly with researchers to run and optimize demanding multi-GPU and multi-node workloads. - Operate and improve the broader compute platform—shared storage, networking, containers, monitoring—and build automation and self-service workflows that reduce friction for researchers. - Support GPU platform and hardware bring-up: validation, enablement, debugging, and operational readiness. - Manage AMD's university program clusters and
Track similar jobs
Get email alerts when new roles like this are posted.
Based on: AI Research Infrastructure Engineer
See how this role matches your resume
Sign in to get an AI match score and personalized feed.
Aftermarket Account Manager
HTS Engineering Ltd.
Full Stack Developer (Java + Angular)
Tecdata Engineering
Accounting and Financial Analyst
Linamar/McLaren Engineering
Network Automation Engineer
DLS Engineering
Quality Engineer
Answer Engineering / Re:Build Manufacturing
SOVT Technical Writer
Akima Systems Engineering
Aftermarket Account Manager
HTS Engineering Ltd.
Full Stack Developer (Java + Angular)
Tecdata Engineering
Accounting and Financial Analyst
Linamar/McLaren Engineering
Network Automation Engineer
DLS Engineering
Quality Engineer
Answer Engineering / Re:Build Manufacturing
SOVT Technical Writer
Akima Systems Engineering