Scale is seeking an AI Infrastructure Engineer to design and build fault-tolerant, high-performance backend systems for serving large language models at scale. You will contribute to an internal platform that enables LLM capability discovery and collaborate with researchers and engineers to optimize models for production and research use cases. You will conduct architecture reviews to uphold best practices in system design and scalability, develop monitoring and observability solutions to ensure system health and performance, and lead projects end-to-end from requirements gathering to implementation in a cross-functional environment. The ideal candidate has 4+ years of experience building large-scale backend systems, strong programming skills in Python, Go, Rust, or C++, and hands-on experience with LLM serving and routing fundamentals (rate limiting, token streaming, load balancing, budgets). Familiarity with containers and orchestration tools (Docker, Kubernetes), cloud infrastructure (AWS, GCP), and infrastructure as code (Terraform) is required. Nice-to-haves include experience with modern LLM serving frameworks such as vLLM, SGLang, TensorRT-LLM, or text-generation-inference. This is an on-site role based in London, United Kingdom.
Scale is seeking an AI Infrastructure Engineer to design and build fault-tolerant, high-performance backend systems for serving large language models at scale. You will contribute to an internal platform that enables LLM capability discovery and collaborate with researchers and engineers to optimize models for production and research use cases. You will conduct architecture reviews to uphold best practices in system design and scalability, develop monitoring and observability solutions to ensure system health and performance, and lead projects end-to-end from requirements gathering to implementation in a cross-functional environment. The ideal candidate has 4+ years of experience building large-scale backend systems, strong programming skills in Python, Go, Rust, or C++, and hands-on experience with LLM serving and routing fundamentals (rate limiting, token streaming, load balancing, budgets). Familiarity with containers and orchestration tools (Docker, Kubernetes), cloud infrastructure (AWS, GCP), and infrastructure as code (Terraform) is required. Nice-to-haves include experience with modern LLM serving frameworks such as vLLM, SGLang, TensorRT-LLM, or text-generation-inference. This is an on-site role based in London, United Kingdom.
Track similar jobs
Get email alerts when new roles like this are posted.
Based on: AI Infrastructure Engineer
See how this role matches your resume
Sign in to get an AI match score and personalized feed.
Track similar jobs
Get email alerts when new roles like this are posted.
Based on: AI Infrastructure Engineer
See how this role matches your resume
Sign in to get an AI match score and personalized feed.
© 2026 JobMatcher. All rights reserved.
© 2026 JobMatcher. All rights reserved.
Cloud Engineer
ST Engineering
Android Software Developer
ACKtive Engineering Solutions LLC
Senior Software Developer - Network Devices
ACKtive Engineering Solutions LLC
Senior Data Scientist, Agentic AI & Multi-Cloud Architecture
Akima Systems Engineering (ASE)
Senior HR Generalist
Partner Engineering & Science, Inc.
Accountant
HCL Engineering & Surveying LLC
Cloud Engineer
ST Engineering
Android Software Developer
ACKtive Engineering Solutions LLC
Senior Software Developer - Network Devices
ACKtive Engineering Solutions LLC
Senior Data Scientist, Agentic AI & Multi-Cloud Architecture
Akima Systems Engineering (ASE)
Senior HR Generalist
Partner Engineering & Science, Inc.
Accountant
HCL Engineering & Surveying LLC