Staff MLOps & GPU Infrastructure Engineer
About The Role
Architect large-scale GPU infrastructure for training and serving high-throughput machine learning models. You will manage Ray clusters, Kubernetes orchestration, and model registries across multi-cloud environments.
Key Requirements & Scope
- 6+ years in DevOps and MLOps engineering environments
- Extensive experience with Kubernetes (K8s), Ray, Kubeflow, and Triton Inference Server
- Experience managing H100/A100 GPU clusters on AWS, GCP, CoreWeave, or Lambda Labs
- Deep expertise in automated CI/CD for model weights and telemetry monitoring
Required Skills & Technologies
Ready to Apply?
Applications are reviewed directly by the mercor talent team. Click below to submit your profile or resume.
Apply on mercorSimilar Positions in AI/ML
Lead the generation of realistic 3D simulation environments and assets for training embodied AI and robotics models. You will translate real-world physical dynamics into simulation suites (Isaac Sim, Gazebo, Unreal Engine) to accelerate autonomous agent learning.
Help develop computer vision foundation models by inspecting, annotating, and evaluating complex imagery datasets. You will apply domain knowledge to label high-resolution aerial, medical, or spatial photography with pinpoint accuracy.
Assist leading robotics AI teams by evaluating sensor telemetry, joint kinematic trajectories, and robotic actuation logs. You will translate mechanical engineering knowledge into clear dataset evaluations.