Quantiphi, an AI-first digital engineering and consulting company, is seeking a highly skilled Senior Platform Engineer to design, optimize, and scale infrastructure for GenAI and large language model workloads. This role focuses on GPU profiling, distributed training, and high-performance compute environments, and offers the opportunity to build GenAI platform foundations and support production deployments in collaboration with data science, MLOps, and application teams. Responsibilities: Design and implement scalable infrastructure for LLM and GenAI workloads across multi-GPU environments; Perform GPU profiling, benchmarking, and performance optimization for distributed training; Manage compute jobs using Slurm clusters and OpenShift/Kubernetes; Enable and optimize the NVIDIA GPU stack (CUDA, cuDNN, NCCL, Triton, RAPIDS); Collaborate with cross-functional teams to deploy models in research and production; Build and support GenAI pipelines (fine-tuning, RAG, multi-modal inferencing, LLMOps); Develop reusable infrastructure templates using Terraform and Helm; Contribute to internal PoCs/workshops and support client-facing delivery. Qualifications: Strong experience with Slurm and distributed training; Hands-on experience with Red Hat OpenShift and/or Kubernetes; Deep knowledge of NVIDIA GPU ecosystem; Linux systems, performance tuning, and multi-GPU optimization; Experience deploying GenAI workloads (LLM tuning, RAG, multi-modal); Infra-as-code tools (Terraform, Ansible); Cloud GPU environments (GCP, Azure, AWS, OCI) and/or on-prem GPU clusters; Nice-to-have: NVIDIA NIMs/DG
Quantiphi, an AI-first digital engineering and consulting company, is seeking a highly skilled Senior Platform Engineer to design, optimize, and scale infrastructure for GenAI and large language model workloads. This role focuses on GPU profiling, distributed training, and high-performance compute environments, and offers the opportunity to build GenAI platform foundations and support production deployments in collaboration with data science, MLOps, and application teams. Responsibilities: Design and implement scalable infrastructure for LLM and GenAI workloads across multi-GPU environments; Perform GPU profiling, benchmarking, and performance optimization for distributed training; Manage compute jobs using Slurm clusters and OpenShift/Kubernetes; Enable and optimize the NVIDIA GPU stack (CUDA, cuDNN, NCCL, Triton, RAPIDS); Collaborate with cross-functional teams to deploy models in research and production; Build and support GenAI pipelines (fine-tuning, RAG, multi-modal inferencing, LLMOps); Develop reusable infrastructure templates using Terraform and Helm; Contribute to internal PoCs/workshops and support client-facing delivery. Qualifications: Strong experience with Slurm and distributed training; Hands-on experience with Red Hat OpenShift and/or Kubernetes; Deep knowledge of NVIDIA GPU ecosystem; Linux systems, performance tuning, and multi-GPU optimization; Experience deploying GenAI workloads (LLM tuning, RAG, multi-modal); Infra-as-code tools (Terraform, Ansible); Cloud GPU environments (GCP, Azure, AWS, OCI) and/or on-prem GPU clusters; Nice-to-have: NVIDIA NIMs/DG
Track similar jobs
Get email alerts when new roles like this are posted.
Based on: Senior Platform Engineer
See how this role matches your resume
Sign in to get an AI match score and personalized feed.
Track similar jobs
Get email alerts when new roles like this are posted.
Based on: Senior Platform Engineer
See how this role matches your resume
Sign in to get an AI match score and personalized feed.
D365 Solutions Architect
Diversified
Senior Director, Customer Support
Dayforce
Senior Software Developer
Dayforce
Engineering Manager, Monetisation
Pet Media Group (PMG)
Strategic Account Executive
Optum
Strategic Account Executive
Optum
D365 Solutions Architect
Diversified
Senior Director, Customer Support
Dayforce
Senior Software Developer
Dayforce
Engineering Manager, Monetisation
Pet Media Group (PMG)
Strategic Account Executive
Optum
Strategic Account Executive
Optum