仕事内容
<h1><strong>AI Infrastructure Systems Engineer</strong></h1>
<p><strong>Hybrid at our office in Amsterdam or Remote in the UK.</strong></p>
<p><strong>Build the infrastructure powering the next generation of AI.</strong></p>
<p>At Together AI, you’ll build and operate one of the world’s largest GPU fleets used for frontier model training and inference. This isn’t a traditional infrastructure role—we’re looking for engineers who love building systems, automating everything, and solving problems at massive scale.</p>
<p>If you enjoy writing software more than clicking dashboards, obsess over eliminating manual work, and want to build infrastructure that manages tens of thousands of GPUs autonomously, we’d love to talk.</p>
<h2><strong>What You’ll Build</strong></h2>
<ul>
<li>Design and build <strong>fleet automation systems</strong> that provision, validate, deploy, upgrade, repair, and retire GPU clusters with minimal human intervention.</li>
<li>Build <strong>AI Infrastructure Agents</strong> that automate deployment, root-cause failures, incident triage, and autonomous remediation.</li>
<li>Develop <strong>Fleet Intelligence</strong> platforms that continuously monitor hardware health, firmware, networking, storage, thermals, and workload performance to predict failures before they impact customers.</li>
<li>Build software that maximizes <strong>GPU availability, utilization, performance, and reliability</strong> across thousands of accelerators.</li>
<li>Create automated validation systems for GPUs, InfiniBand/RoCE fabrics, NVLink/NVSwitch, storage, and distributed AI workloads.</li>
<li>Build internal platforms and developer tools that allow infrastructure to be managed through software—not manual operations.</li>
<li>Continuously improve deployment velocity, reliability, and operational efficiency through automation.</li>
<li>Partner closely with hardware, networking, platform, and AI teams to push the limits of AI infrastructure.</li>
</ul>
<h2><strong>What We’
求めるスキル
Python
CUDA
Kubernetes
Rust