
Job Overview
Location
United States
Job Type
Full-time
Category
Software Engineering
Date Posted
June 22, 2026
Full Job Description
đź“‹ Description
- • Optimize LLM inference across multiple modalities (text, vision, audio) to drive business value and meet customer performance and cost-efficiency goals
- • Provide technical support for supervised fine-tuning (SFT) and LoRA-based fine-tuning to enhance model quality for customer-specific use cases
- • Design and implement production-ready LLM-based applications using Nebius Token Factory’s serverless inference and fine-tuning platform
- • Build scalable AI solutions leveraging multimodal models and domain-specific open-source LLMs through the company’s proprietary APIs
- • Deliver expert guidance in prompt engineering, RAG architectures, and model selection to help customers transition from prototype to scalable production
- • Collaborate with product and engineering teams to translate customer feedback into platform improvements and roadmap priorities
- • Guide clients through end-to-end deployment of LLM workflows, ensuring reliability, low-latency inference, and efficient resource utilization
- • Conduct LLM evaluation using task-specific benchmarks, offline/online evaluation pipelines, and LLM-as-a-judge methodologies to validate model performance
- • Work closely with backend engineering teams to refine inference optimizations such as speculative decoding, quantization, and cache-aware routing based on real-world customer needs
- • Support customers in selecting and integrating open-source LLMs and commercial APIs (e.g., OpenAI, Anthropic) into their applications using Nebius Token Factory
- • Maintain deep technical alignment with the evolving LLM ecosystem, including model architectures, fine-tuning techniques, and inference frameworks
🎯 Requirements
- • 5+ years of experience in ML/AI systems, with at least 2 years focused on LLMs and generative AI
- • Deep knowledge of the LLM ecosystem, including model architectures and fine-tuning approaches
- • Hands-on experience running LLMs in production, including deploying and operating inference workloads
- • Hands-on experience with LLM fine-tuning, including supervised fine-tuning (SFT/LoRA) and data preparation/curation
- • Strong Python programming skills
- • Excellent communication skills, with the ability to clearly explain technical concepts to diverse audiences
🏖️ Benefits
- • 100% company-paid medical, dental, and vision coverage for employees and families
- • 401(k) Plan with up to 4% company match and immediate vesting
- • 20 weeks paid parental leave for primary caregivers, 12 weeks for secondary caregivers
- • Remote work reimbursement of up to $85/month for mobile and internet
- • Company-paid short-term, long-term, and life insurance coverage
- • Competitive compensation ranging from $210,000 to $260,000 USD
Skills & Technologies
See exactly how your profile matches this role — strengths, skill gaps, and what to do about them.
About Nebius Group N.V.
Nebius Group N.V. is a Netherlands-based technology company that operates a full-stack cloud platform designed for AI and machine learning workloads. It provides scalable GPU and CPU infrastructure, managed Kubernetes, object storage, and specialized AI services to enterprises and research organizations worldwide. The company was formed from the restructuring of Yandex N.V. and continues to serve global markets with data centers across Europe and North America.
Subscribe to the weekly newsletter for similar remote roles and curated hiring updates.
Newsletter
Weekly remote jobs and featured talent.
No spam. Only curated remote roles and product updates. You can unsubscribe anytime.
Similar Opportunities

xAI Corporation
3 months ago

Ramp Business Corporation
2 months ago

Hims & Hers Health, Inc.
2 months ago
2 months ago
