Based in Cambridge, Massachusetts, this on-site role at Lila Sciences seeks a Staff/Principal DevOps Engineer, AI Inference to design, implement, and optimize infrastructure for serving machine learning models at scale. The role bridges platform engineering, site reliability engineering, and ML infrastructure to build systems that deliver low-latency, high-throughput inference across GPU clusters and cloud accelerators. You will collaborate with ML engineers, research scientists, and software engineers to create inference platforms that reliably serve models to production users while maximizing compute efficiency.
Based in Cambridge, Massachusetts, this on-site role at Lila Sciences seeks a Staff/Principal DevOps Engineer, AI Inference to design, implement, and optimize infrastructure for serving machine learning models at scale. The role bridges platform engineering, site reliability engineering, and ML infrastructure to build systems that deliver low-latency, high-throughput inference across GPU clusters and cloud accelerators. You will collaborate with ML engineers, research scientists, and software engineers to create inference platforms that reliably serve models to production users while maximizing compute efficiency.
What you'll be building: - GPU/accelerator infrastructure on Kubernetes: scheduling, resource isolation, multi-tenant GPU sharing, device plugins, and topology-aware placement for inference workloads - Model serving platforms using frameworks such as vLLM, Triton Inference Server, TGI, or custom stacks with optimized batching, caching, and request routing - Intelligent request routing and load balancing across heterogeneous accelerator fleets (NVIDIA GPUs, AWS Inferentia/Trainium) to maximize utilization and minimize latency - Autoscaling systems that dynamically match inference compute supply with demand across production, research, and experimental workloads - Production-grade deployment pipelines for ML models: canary rollouts, A/B testing, model versioning, and safe rollback across multi-region deployments - Infrastructure-as-code with Terraform and Helm for GPU-accelerated EKS clusters, including node pools, spot/on-demand strategies, and accelerator-specifi
Track similar jobs
Get email alerts when new roles like this are posted.
Based on: Staff/Principal DevOps Engineer, AI Inference
See how this role matches your resume
Sign in to get an AI match score and personalized feed.
What you'll be building: - GPU/accelerator infrastructure on Kubernetes: scheduling, resource isolation, multi-tenant GPU sharing, device plugins, and topology-aware placement for inference workloads - Model serving platforms using frameworks such as vLLM, Triton Inference Server, TGI, or custom stacks with optimized batching, caching, and request routing - Intelligent request routing and load balancing across heterogeneous accelerator fleets (NVIDIA GPUs, AWS Inferentia/Trainium) to maximize utilization and minimize latency - Autoscaling systems that dynamically match inference compute supply with demand across production, research, and experimental workloads - Production-grade deployment pipelines for ML models: canary rollouts, A/B testing, model versioning, and safe rollback across multi-region deployments - Infrastructure-as-code with Terraform and Helm for GPU-accelerated EKS clusters, including node pools, spot/on-demand strategies, and accelerator-specifi
Track similar jobs
Get email alerts when new roles like this are posted.
Based on: Staff/Principal DevOps Engineer, AI Inference
See how this role matches your resume
Sign in to get an AI match score and personalized feed.
© 2026 JobMatcher. Все права защищены.
© 2026 JobMatcher. Все права защищены.