Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
We are seeking a MLOps Engineer to design, build, and operate high-performance, reliable inference platforms for serving large machine learning models in production. The role focuses on the systems engineering side of AI deployment, including request routing, batching, caching, autoscaling, GPU utilization, and end-to-end observability across diverse model workloads. The ideal candidate brings strong distributed systems and performance engineering expertise, has shipped serving systems at scale, and understands the trade-offs between latency, throughput, cost, and quality in ML serving.
Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
We are seeking a MLOps Engineer to design, build, and operate high-performance, reliable inference platforms for serving large machine learning models in production. The role focuses on the systems engineering side of AI deployment, including request routing, batching, caching, autoscaling, GPU utilization, and end-to-end observability across diverse model workloads. The ideal candidate brings strong distributed systems and performance engineering expertise, has shipped serving systems at scale, and understands the trade-offs between latency, throughput, cost, and quality in ML serving.
Key Responsibilities - Design and operate model serving platforms supporting diverse workloads including LLMs, vision models, and recommendation systems. - Optimize inference performance using continuous batching, paged attention, speculative decoding, and request multiplexing. - Implement multi-tenant routing, rate limiting, and quality-of-service policies across model endpoints. - Build autoscaling and capacity management systems that balance latency, throughput, and cost. - Tune GPU utilization, memory management, and KV cache strategies for LLM serving workloads. - Integrate model serving with API gateways, identity systems, and observability platforms. - Implement caching, prompt deduplication, and response reuse strategies where appropriate. - Drive end-to-end observability including latency histograms, queue dynamics, G
Track similar jobs
Get email alerts when new roles like this are posted.
Based on: MLOps Engineer
See how this role matches your resume
Sign in to get an AI match score and personalized feed.
Key Responsibilities - Design and operate model serving platforms supporting diverse workloads including LLMs, vision models, and recommendation systems. - Optimize inference performance using continuous batching, paged attention, speculative decoding, and request multiplexing. - Implement multi-tenant routing, rate limiting, and quality-of-service policies across model endpoints. - Build autoscaling and capacity management systems that balance latency, throughput, and cost. - Tune GPU utilization, memory management, and KV cache strategies for LLM serving workloads. - Integrate model serving with API gateways, identity systems, and observability platforms. - Implement caching, prompt deduplication, and response reuse strategies where appropriate. - Drive end-to-end observability including latency histograms, queue dynamics, G
Track similar jobs
Get email alerts when new roles like this are posted.
Based on: MLOps Engineer
See how this role matches your resume
Sign in to get an AI match score and personalized feed.
D365 Solutions Architect
Diversified
Senior Director, Customer Support
Dayforce
Senior Software Developer
Dayforce
Engineering Manager, Monetisation
Pet Media Group (PMG)
Strategic Account Executive
Optum
Strategic Account Executive
Optum
D365 Solutions Architect
Diversified
Senior Director, Customer Support
Dayforce
Senior Software Developer
Dayforce
Engineering Manager, Monetisation
Pet Media Group (PMG)
Strategic Account Executive
Optum
Strategic Account Executive
Optum