Mindbeam AI logo

Machine Learning Engineer - Pre Training

Job Overview

Location

United States

Job Type

Full-time

Category

Software Engineering

Date Posted

July 4, 2026

Full Job Description

📋 Description

  • Design and optimize large-scale pre-training pipelines for foundation models, focusing on maximizing throughput and computational efficiency.
  • Implement distributed training strategies across GPU and TPU clusters to enable scalable training of generative AI models.
  • Collaborate directly with research teams to convert experimental AI models into robust, production-ready training workflows.
  • Develop and maintain monitoring, logging, and fault-tolerance systems to ensure reliability during extended large-scale training runs.
  • Continuously benchmark model training performance across diverse hardware configurations and software stacks to identify and resolve bottlenecks.
  • Optimize GPU memory usage, data loading, and parallelism techniques to reduce training time and improve resource utilization.
  • Work with containerized environments using Docker and orchestration tools like Kubernetes to deploy and manage training infrastructure.
  • Analyze high-performance computing and networking limitations to improve system-level performance in distributed training environments.
  • Tune hyperparameters and system configurations to achieve optimal convergence rates and stability during pre-training phases.
  • Document training pipelines, performance metrics, and optimization strategies for cross-team knowledge sharing and reproducibility.
  • Stay current with advancements in ML infrastructure, distributed systems, and hardware accelerators to inform architectural decisions.
  • Participate in code reviews and technical design discussions to uphold engineering standards across the pre-training platform.
  • Troubleshoot and resolve failures in multi-node training jobs, including hardware errors, network partitioning, and software crashes.
  • Contribute to the development of internal tools that automate repetitive tasks in the pre-training lifecycle.
  • Ensure training systems are scalable, maintainable, and aligned with enterprise-grade reliability requirements.
  • Support the integration of new model architectures into existing pre-training frameworks with minimal disruption.

🎯 Requirements

  • Bachelor’s, Master’s, or PhD in Computer Science, Engineering, or related field—or equivalent experience
  • 2+ years of experience with large-scale model training and distributed systems
  • Strong coding skills in Python and familiarity with ML frameworks (PyTorch, TensorFlow, JAX)
  • Experience with GPU scheduling, memory optimization, and parallelism strategies
  • Comfort with containerized and orchestrated environments (Docker/Kubernetes)
  • Understanding of high-performance computing and networking bottlenecks

🏖️ Benefits

  • Opportunity to work on cutting-edge generative AI infrastructure with global impact
  • Collaborative environment with researchers and engineers pushing the boundaries of open-source AI
  • Exposure to state-of-the-art hardware and large-scale computing clusters
  • Culture focused on curiosity, openness, and empowering technical innovation

Skills & Technologies

Python
Docker
Kubernetes
TensorFlow
PyTorch
Onsite
Degree Required

Ready to Apply?

You will be redirected to an external site to apply.

AI Job Fit Analysis
Pro

See exactly how your profile matches this role — strengths, skill gaps, and what to do about them.

Mindbeam AI logo
Mindbeam AI
Visit Website

About Mindbeam AI

Mindbeam AI is a New York City–based startup specializing in next-generation AI infrastructure. Its flagship product, Litespark, is a framework designed to accelerate the pre-training and fine-tuning of large language models (LLMs). Litespark utilizes advanced algorithms to significantly reduce training times—from months to days—while minimizing costs and energy consumption. The framework is compatible with industry-standard machine learning frameworks like PyTorch, TensorFlow, and JAX, and is optimized for NVIDIA GPU hardware. Mindbeam's solutions are utilized by Fortune 100 enterprises and are available on AWS Marketplace.

Get more remote jobs like this

Subscribe to the weekly newsletter for similar remote roles and curated hiring updates.

Newsletter

Weekly remote jobs and featured talent.

No spam. Only curated remote roles and product updates. You can unsubscribe anytime.

Similar Opportunities

Expired
Distro, Inc. logo

Distro, Inc.

Argentina
Full-time
Expired Jun 22, 2026
Azure
REST
Data Science
+1 more

3 months ago

Expired
Philippines
Full-time
Expired Jul 21, 2026
Remote

2 months ago

Equifax Inc. logo

Equifax Inc.

GBR-Contract-Home-Based-Remote
Full-time
Expires Sep 7, 2026
Remote
Degree Required

20 days ago

Springer Nature Limited logo

Springer Nature Limited

Remote, Switzerland
Full-time
Expires Aug 17, 2026
Junior
Remote
Degree Required

1 month ago