d-Matrix Corporation logo

Contract Site Reliability Engineer — AI Accelerator Infrastructure

Job Overview

Location

Santa Clara

Job Type

Full-time

Category

DevOps & SysAdmin

Date Posted

July 16, 2026

Full Job Description

📋 Description

  • At d-Matrix, we are focused on unleashing the potential of generative AI to power the transformation of technology.
  • We are at the forefront of software and hardware innovation, pushing the boundaries of what is possible.
  • Our culture is one of respect and collaboration.
  • We value humility and believe in direct communication.
  • Our team is inclusive, and our differing perspectives allow for better solutions.
  • We are seeking individuals passionate about tackling challenges and are driven by execution.
  • Ready to come find your playground? Together, we can help shape the endless possibilities of AI.
  • d-Matrix designs purpose-built AI inference silicon.
  • Our unified infrastructure team, SRE engineers, DevOps engineers, and data center & lab technicians own the physical and virtual layer that every engineering team depends on.
  • As a DC & lab technician, you are the hands and feet of that team: racking servers, running cables, executing hardware bring-ups, and keeping lab environments in the precise, audit-ready state that high-velocity silicon and software development demands.
  • This is a one-year contract with potential for a full-time conversion.
  • This is an ownership role, not a ticket executor role.
  • You operate independently, document what you build, and escalate with precision when something needs engineering attention.
  • You will be responsible for the physical infrastructure and hardware bring-up.
  • You will rack, stack, cable, and decommission servers, PDUs, network gear, and storage in on-premises labs and colocation facilities.
  • You will execute hardware bring-up for d-Matrix AI-accelerated systems and validation test benches, including BIOS configuration, firmware validation, and OS installation (Linux primary).
  • You will replace and upgrade components: PCIe cards, NICs, storage, memory, and accelerator hardware across multiple server generations.
  • You will maintain lab spaces in organized, ESD-compliant, and audit-ready condition at all times.
  • You will be responsible for asset management and inventory.
  • You will own accurate asset tracking and inventory records using DCIM tools (NetBox, Jira, or equivalent); every piece of hardware is accounted for with the current configuration state.
  • You will coordinate equipment moves between lab, staging, and colocation; manage shipping, receiving, and RMA workflows.
  • You will track hardware lifecycle status: warranty, EOL, refresh schedules, and spare parts inventory.
  • You will be responsible for troubleshooting and incident support.
  • You will diagnose and resolve hardware failures, power and thermal issues, and network connectivity problems — escalating to SRE engineers with clear documentation and reproduction steps.
  • You will serve as the first physical responder for infrastructure incidents requiring hands-on intervention in lab or colo environments.
  • You will document all troubleshooting actions in the team’s ticketing and knowledge base systems, contributing to runbooks that reduce repeat escalations.
  • You will be responsible for documentation and collaboration.
  • You will write and maintain SOPs, rack diagrams, and cabling guides clear enough for a new technician to execute without shadowing.
  • You will communicate issues and status across hardware, software, SRE, and silicon validation teams professionally and with precision.
  • You will support parallel workstreams as silicon development programs evolve; operate independently when SREs are heads-down on engineering work.
  • You will bring a strong background in data center operations, hardware validation, or systems technician roles, hands-on with real servers.
  • You will have proven rack-and-stack experience: physical server installation, PDU and patch panel cabling, rack power planning, and cable management to professional standards.
  • You will have hardware bring-up experience: BIOS/UEFI configuration, firmware updates, component replacement, and OS installation across Linux distributions.
  • You will have Linux command-line proficiency: enough to run diagnostics, inspect logs, and execute runbooks independently.
  • You will have asset management discipline: experience with DCIM or inventory tools and a track record of accurate, up-to-date records.
  • You will have strong written communication: you document what you do, escalate with context, and write SOPs others can follow.
  • You will be comfortable operating independently in a fast-moving startup environment.
  • Preferred qualifications include prior colocation data center experience, experience with AI/ML or GPU hardware, basic Ansible exposure, familiarity with high-speed interconnects, and Python or Bash scripting for operational task automation.
  • You will be a critical part of the infrastructure that powers d-Matrix’s AI hardware development programs.
  • In a small, high-ownership team, your work is immediately visible; when you bring up a system cleanly, the engineers depending on it notice.
  • If you take pride in physical infrastructure done right and want to work at the cutting edge of AI silicon development, this is the role for you.
  • d-Matrix is proud to be an equal opportunity workplace and affirmative action employer.
  • We’re committed to fostering an inclusive environment where everyone feels welcomed and empowered to do their best work.
  • We hire the best talent for our teams, regardless of race, religion, color, age, disability, sex, gender identity, sexual orientation, ancestry, genetic information, marital status, national origin, political affiliation, or veteran status.
  • Our focus is on hiring teammates with humble expertise, kindness, dedication and a willingness to embrace challenges and learn together every day.
  • d-Matrix does not accept resumes or candidate submissions from external agencies.
  • We appreciate the interest and effort of recruitment firms, but we kindly request that individual interested in opportunities with d-Matrix apply directly through our official channels.
  • This approach allows us to streamline our hiring processes and maintain a consistent and fair evaluation of all applicants.
  • Thank you for your understanding and cooperation.

Skills & Technologies

Python
Linux
DevOps
Onsite

Ready to Apply?

You will be redirected to an external site to apply.

AI Job Fit Analysis
Pro

See exactly how your profile matches this role — strengths, skill gaps, and what to do about them.

d-Matrix Corporation logo
d-Matrix Corporation
Visit Website

About d-Matrix Corporation

d-Matrix designs silicon for high-efficiency AI inference at scale. Its Corsair compute platform combines in-memory computing with a digital approach to slash latency and energy use in transformer and generative workloads. Targeting hyperscale data centers and edge deployments, the company offers hardware and software stacks that integrate into existing AI pipelines. Founded in 2019 and headquartered in Santa Clara, California, d-Matrix serves cloud and enterprise customers seeking cost-effective alternatives to GPUs for large language model serving.

Get more remote jobs like this

Subscribe to the weekly newsletter for similar remote roles and curated hiring updates.

Newsletter

Weekly remote jobs and featured talent.

No spam. Only curated remote roles and product updates. You can unsubscribe anytime.

Similar Opportunities

Pragmatike Soluciones Tecnológicas S.L. logo

Pragmatike Soluciones Tecnológicas S.L.

Albania
Full-time
Expires Aug 29, 2026
Python
Node.js
Kubernetes
+4 more

24 days ago

Expired
ARGENTINA
Full-time
Expired Jun 20, 2026
Python
JavaScript
TypeScript
+5 more

3 months ago

Expired
Singapore
Full-time
Expired Jun 16, 2026
AWS
Azure
Kubernetes
+2 more

3 months ago

Expired
ARGENTINA
Full-time
Expired Jun 20, 2026
AWS
Docker
Kubernetes
+4 more

3 months ago