
Job Overview
Location
Paris Offices
Job Type
Full-time
Category
Engineering
Date Posted
August 18, 2026
Full Job Description
📋 Description
- • Meet Arago and the Aragonians, a company re-engineering the foundations of computing from first principles.
- • Arago's mission is to rethink how processors are built, driven by the explosive growth of AI.
- • The company has developed a proprietary technology that fuses optical and CMOS technologies to deliver an order-of-magnitude increase in performance.
- • Arago is the fastest and currently the only company to have built such a processor.
- • The company is backed by leading deep-tech investors and industry leaders.
- • The work is guided by three clear values: do great things, move with high velocity, and operate as one unit.
- • The environment is demanding, with constant learning, ownership, and execution expected.
- • Exceptional people have the opportunity to do their life’s work.
- • As a ML Systems Engineer, you will optimize the execution and serving of modern AI models on Arago's custom accelerator.
- • You will work across kernels, model execution, multi-device distribution, runtime, and inference serving, while helping shape the software stack around the capabilities of Arago's hardware.
- • You will analyze modern AI workloads and identify kernel-, runtime-, memory-, and system-level bottlenecks on Arago's accelerator.
- • You will develop and optimize custom kernels, fused operators, and execution strategies to maximize device utilization.
- • You will design efficient mappings of models and operators across multiple Arago devices, including communication and synchronization strategies.
- • You will develop inference-serving techniques such as continuous batching, paged KV caches, prefix/context caching, chunked prefill, and prefill/decode interleaving or disaggregation.
- • You will build profiling, benchmarking, and performance-analysis infrastructure spanning kernels, full models, and serving workloads.
- • You will work closely with Arago's hardware, compiler, and runtime teams to co-design software abstractions and influence future hardware features based on real model workloads.
- • The role requires strong experience in high-performance ML inference, GPU/accelerator programming, or ML systems engineering.
- • You will have a deep understanding of computer architecture, accelerator/GPU execution models, memory hierarchies, parallelism, and performance bottlenecks.
- • You will have experience developing and optimizing custom kernels using CUDA, Triton, ROCm/HIP, or equivalent low-level programming environments.
- • You will have experience with operator fusion, tiling, scheduling, data movement optimization, graph execution, and profiling of compute- and memory-bound workloads.
- • You will have a strong understanding of distributed model execution, including tensor, pipeline, sequence, and/or expert parallelism and communication/computation overlap.
- • You will have hands-on experience with modern inference-serving systems such as vLLM, SGLang, TensorRT-LLM, or equivalent, including KV-cache management, continuous batching, paged attention, and prefill/decode scheduling.
- • You will have strong C++ and Python skills, and comfort working on a custom accelerator stack where compiler, runtime, kernels, and abstractions are actively being developed.
- • Exposure to or experience with MLIR and MLIR dialects is a strong plus.
- • The role offers competitive cash compensation, with final package based on location, experience, and the pay of team members in similar positions.
- • You will have a meaningful stock option plan offered (included in the majority of full-time offers).
- • You will have healthcare coverage (including family-friendly options), pension contributions, professional development support, and 25 days of PTO, in addition to public holidays.
- • You will have ownership of a key technical domain, with significant vertical and/or horizontal growth opportunities, based on performance and individual drive.
🎯 Requirements
- • Strong experience in high-performance ML inference, GPU/accelerator programming, or ML systems engineering.
- • Deep understanding of computer architecture, accelerator/GPU execution models, memory hierarchies, parallelism, and performance bottlenecks.
- • Experience developing and optimizing custom kernels using CUDA, Triton, ROCm/HIP, or equivalent low-level programming environments.
- • Experience with operator fusion, tiling, scheduling, data movement optimization, graph execution, and profiling of compute- and memory-bound workloads.
🏖️ Benefits
- • Competitive cash compensation.
- • Meaningful stock option plan offered.
- • Healthcare coverage (including family-friendly options).
- • Pension contributions.
- • Professional development support.
- • 25 days of PTO, in addition to public holidays.
Skills & Technologies
See exactly how your profile matches this role — strengths, skill gaps, and what to do about them.
About Arago GmbH
Arago GmbH is a German software company that develops and licenses the AI knowledge-automation platform HIRO. The system applies machine reasoning to enterprise IT and business processes, learning from expert knowledge to resolve incidents, optimize workflows, and manage hybrid cloud infrastructures. Founded in 1995 and headquartered in Frankfurt am Main, Arago serves global banks, manufacturers, and service providers, enabling them to cut operational costs and accelerate digital transformation through intelligent automation.
Subscribe to the weekly newsletter for similar remote roles and curated hiring updates.
Newsletter
Weekly remote jobs and featured talent.
No spam. Only curated remote roles and product updates. You can unsubscribe anytime.
Similar Opportunities

NETGEAR, Inc.
2 months ago

Unilever PLC
2 months ago

Silver.com LLC
2 months ago

Latamcent
3 months ago