Responsibilities:
Build components to enable large-scale distributed training with focus on usability, performance and resiliency.
Collaborate with partner teams, dependencies (e.g., Google Compute Engine (GCE), Google Kubernetes Engine (GKE), Cloud Tensor Processing Unit (Cloud TPU), Core Infra, CoreML) to develop pre-training/post-training software to enable customer run reliable and performant workloads.
Work with Product Team, Account teams and customers to identify key pain points, define scope of the problems, translate them into projects and lead them to execution.
Contribute to product excellence, and improve overall product quality and user experience.
Minimum qualifications:
Bachelor’s degree or equivalent practical experience.
2 years of experience with software development in one or more programming languages (e.g., Python, C++ or Java).
Experience in machine learning infrastructure.
Experience in distributed machine learning.
Experience with distributed computing and Cloud APIs.
Preferred qualifications:
Experience with large-scale distributed machine learning, ML performance optimizations, etc.
Experience with ML frameworks (e.g., PyTorch).
Experience working on Cloud and ML products.
Experience in back-end system development and familiarity with common Google Infrastructure.
Experience with OSS AI/ML Job scheduling software.
Ready to apply?
Sign up first — takes a minute — and you get Hiro’s take on this role, a resume tailored to it, and (if available) a referral from a real employee at Google. All free with your Pro gift.
Sign up to apply