
Job Overview
Location
Remote; Singapore
Job Type
Full-time
Category
Software Engineering
Date Posted
June 22, 2026
Full Job Description
đź“‹ Description
- • Optimize LLM inference across multiple modalities (text, vision, audio) to drive business value and meet customer performance and cost-efficiency goals
- • Provide expert support in supervised fine-tuning (SFT) and LoRA-based fine-tuning, as well as reinforcement learning fine-tuning (RLFT), to maximize model quality for customer use cases
- • Design and implement production-ready LLM-based applications using Nebius Token Factory’s serverless inference and fine-tuning platform
- • Build scalable AI solutions leveraging multimodal models and domain-specific open-source LLMs through the platform’s APIs and custom endpoints
- • Deliver technical guidance on prompt engineering, RAG architectures, and model selection to ensure optimal application performance and reliability
- • Collaborate with customers to transition prototypes into scalable production deployments with emphasis on latency, throughput, and cost optimization
- • Work closely with product and engineering teams to translate customer feedback into platform improvements and roadmap priorities
- • Conduct LLM evaluation using task-specific benchmarks and offline/online evaluation pipelines, including LLM-as-a-judge methodologies
- • Apply in-house optimizations such as speculative decoding, quantization, and cache-aware routing to enhance inference efficiency and reduce operational overhead
- • Serve as a technical liaison between customers and internal engineering teams to ensure seamless integration and resolution of platform-related challenges
- • Maintain deep expertise in the LLM ecosystem, including model architectures, fine-tuning techniques, and inference frameworks to provide authoritative guidance
- • Demonstrate fluency in Mandarin Chinese to effectively communicate with clients and stakeholders in Mandarin-speaking markets
🎯 Requirements
- • 5+ years of experience in ML/AI systems, with at least 2 years focused on LLMs and generative AI
- • Deep knowledge of the LLM ecosystem, including model architectures and fine-tuning approaches
- • Hands-on experience running LLMs in production, including deployment and operation of inference workloads
- • Strong Python programming skills
- • Excellent communication skills with the ability to clearly explain technical concepts to diverse audiences
- • Must be fluent in Mandarin Chinese
🏖️ Benefits
- • Competitive compensation
- • Career growth and learning opportunities
- • Flexibility and ownership
- • Collaborative and innovative culture
- • Opportunity to work on impactful AI projects
- • International environment and talented teams
Skills & Technologies
See exactly how your profile matches this role — strengths, skill gaps, and what to do about them.
About Nebius Group N.V.
Nebius Group N.V. is a Netherlands-based technology company that operates a full-stack cloud platform designed for AI and machine learning workloads. It provides scalable GPU and CPU infrastructure, managed Kubernetes, object storage, and specialized AI services to enterprises and research organizations worldwide. The company was formed from the restructuring of Yandex N.V. and continues to serve global markets with data centers across Europe and North America.
Subscribe to the weekly newsletter for similar remote roles and curated hiring updates.
Newsletter
Weekly remote jobs and featured talent.
No spam. Only curated remote roles and product updates. You can unsubscribe anytime.
Similar Opportunities

xAI Corporation
3 months ago

Ramp Business Corporation
2 months ago

Hims & Hers Health, Inc.
2 months ago
2 months ago
