Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. We are seeking an ML Infrastructure Engineer to design, build, and operate the platform layer that powers large-scale AI training and inference workloads, with a focus on GPU clusters, distributed training frameworks, scheduling, storage performance, and developer experience for ML engineers and researchers.
Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. We are seeking an ML Infrastructure Engineer to design, build, and operate the platform layer that powers large-scale AI training and inference workloads, with a focus on GPU clusters, distributed training frameworks, scheduling, storage performance, and developer experience for ML engineers and researchers.
Job Type: Full-time, Direct W2 Salary Range: $100,000–$150,000 annually Experience Required: 6+ years
Sponsorship: U.S. citizens, green card holders, EAD holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B petitions for this position.
Key Responsibilities - Design and operate GPU and accelerator infrastructure for training and inference across on-prem clusters, cloud-managed services, and hybrid configurations. - Build scheduling, queuing, and resource-sharing systems to maximize accelerator utilization across multiple teams. - Integrate PyTorch, JAX, DeepSpeed, FSDP, Megatron-LM, and Ray Train into a unified AI platform. - Operate high-performance storage systems and data pipelines to feed training workloads at near-line-rate. - Design networking architectures supporting RDMA, InfiniBand, NCCL, and high-bandwidth communication. - Build observability for AI workloads, including utilization, throughput, training stability, and failure analytics. - Implement checkpointing, restart, and fault-tolerance patterns for long-running traini
Track similar jobs
Get email alerts when new roles like this are posted.
Based on: ML Infrastructure Engineer
See how this role matches your resume
Sign in to get an AI match score and personalized feed.
Job Type: Full-time, Direct W2 Salary Range: $100,000–$150,000 annually Experience Required: 6+ years
Sponsorship: U.S. citizens, green card holders, EAD holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B petitions for this position.
Key Responsibilities - Design and operate GPU and accelerator infrastructure for training and inference across on-prem clusters, cloud-managed services, and hybrid configurations. - Build scheduling, queuing, and resource-sharing systems to maximize accelerator utilization across multiple teams. - Integrate PyTorch, JAX, DeepSpeed, FSDP, Megatron-LM, and Ray Train into a unified AI platform. - Operate high-performance storage systems and data pipelines to feed training workloads at near-line-rate. - Design networking architectures supporting RDMA, InfiniBand, NCCL, and high-bandwidth communication. - Build observability for AI workloads, including utilization, throughput, training stability, and failure analytics. - Implement checkpointing, restart, and fault-tolerance patterns for long-running traini
Track similar jobs
Get email alerts when new roles like this are posted.
Based on: ML Infrastructure Engineer
See how this role matches your resume
Sign in to get an AI match score and personalized feed.
D365 Solutions Architect
Diversified
Senior Director, Customer Support
Dayforce
Senior Software Developer
Dayforce
Engineering Manager, Monetisation
Pet Media Group (PMG)
Strategic Account Executive
Optum
Strategic Account Executive
Optum
D365 Solutions Architect
Diversified
Senior Director, Customer Support
Dayforce
Senior Software Developer
Dayforce
Engineering Manager, Monetisation
Pet Media Group (PMG)
Strategic Account Executive
Optum
Strategic Account Executive
Optum