
Job Overview
Location
United States
Job Type
Full-time
Category
Software Engineering
Date Posted
July 4, 2026
Full Job Description
📋 Description
- • Design and optimize large-scale pre-training pipelines for foundation models, focusing on maximizing throughput and computational efficiency.
- • Implement distributed training strategies across GPU and TPU clusters to enable scalable training of generative AI models.
- • Collaborate directly with research teams to convert experimental AI models into robust, production-ready training workflows.
- • Develop and maintain monitoring, logging, and fault-tolerance systems to ensure reliability during extended large-scale training runs.
- • Continuously benchmark model training performance across diverse hardware configurations and software stacks to identify and resolve bottlenecks.
- • Optimize GPU memory usage, data loading, and parallelism techniques to reduce training time and improve resource utilization.
- • Work with containerized environments using Docker and orchestration tools like Kubernetes to deploy and manage training infrastructure.
- • Analyze high-performance computing and networking limitations to improve system-level performance in distributed training environments.
- • Tune hyperparameters and system configurations to achieve optimal convergence rates and stability during pre-training phases.
- • Document training pipelines, performance metrics, and optimization strategies for cross-team knowledge sharing and reproducibility.
- • Stay current with advancements in ML infrastructure, distributed systems, and hardware accelerators to inform architectural decisions.
- • Participate in code reviews and technical design discussions to uphold engineering standards across the pre-training platform.
- • Troubleshoot and resolve failures in multi-node training jobs, including hardware errors, network partitioning, and software crashes.
- • Contribute to the development of internal tools that automate repetitive tasks in the pre-training lifecycle.
- • Ensure training systems are scalable, maintainable, and aligned with enterprise-grade reliability requirements.
- • Support the integration of new model architectures into existing pre-training frameworks with minimal disruption.
🎯 Requirements
- • Bachelor’s, Master’s, or PhD in Computer Science, Engineering, or related field—or equivalent experience
- • 2+ years of experience with large-scale model training and distributed systems
- • Strong coding skills in Python and familiarity with ML frameworks (PyTorch, TensorFlow, JAX)
- • Experience with GPU scheduling, memory optimization, and parallelism strategies
- • Comfort with containerized and orchestrated environments (Docker/Kubernetes)
- • Understanding of high-performance computing and networking bottlenecks
🏖️ Benefits
- • Opportunity to work on cutting-edge generative AI infrastructure with global impact
- • Collaborative environment with researchers and engineers pushing the boundaries of open-source AI
- • Exposure to state-of-the-art hardware and large-scale computing clusters
- • Culture focused on curiosity, openness, and empowering technical innovation
Skills & Technologies
See exactly how your profile matches this role — strengths, skill gaps, and what to do about them.
About Mindbeam AI
Mindbeam AI is a New York City–based startup specializing in next-generation AI infrastructure. Its flagship product, Litespark, is a framework designed to accelerate the pre-training and fine-tuning of large language models (LLMs). Litespark utilizes advanced algorithms to significantly reduce training times—from months to days—while minimizing costs and energy consumption. The framework is compatible with industry-standard machine learning frameworks like PyTorch, TensorFlow, and JAX, and is optimized for NVIDIA GPU hardware. Mindbeam's solutions are utilized by Fortune 100 enterprises and are available on AWS Marketplace.
Subscribe to the weekly newsletter for similar remote roles and curated hiring updates.
Newsletter
Weekly remote jobs and featured talent.
No spam. Only curated remote roles and product updates. You can unsubscribe anytime.
Similar Opportunities

Distro, Inc.
3 months ago
2 months ago

Equifax Inc.
20 days ago

Springer Nature Limited
1 month ago
