仕事内容
<h3>About the Role</h3>
<p>Together AI is building the AI Native Cloud, an end-to-end platform for the full<br>generative AI lifecycle, combining the fastest LLM inference engine with state-of-the-art<br>AI cloud infrastructure. The Together Cloud team builds the Together GPU Clusters<br>product, which provides high-performance, AI-ready GPU clusters through a self-serve<br>cloud console and is the virtualized infrastructure layer powering Together’s inference,<br>RL, and fine-tuning products.</p>
<p><br>As a Senior Software Engineer in Together Cloud Infrastructure, you will play a key role<br>in building the next generation AI cloud platform – a highly available, global, blazing-fast<br>cloud infrastructure that virtualizes cutting-edge ML hardware (GB200s/GB300s,<br>BlueField DPUs). You'll enable state-of-the-art ML practitioners with self-serve AI cloud<br>services, such as on-demand + managed Kubernetes and Slurm clusters, for both our<br>internal SaaS products (inference, fine-tuning, RL) and our external cloud customers,<br>spanning dozens of data centers across the world.</p>
<h3><strong>Responsibilities</strong></h3>
<ul>
<li>Design, build, and maintain performant, secure, and highly-available backend services/operators that run in our data centers and automate hardware management, such as Infiniband partitioning, in. DC parallel storage provisioning, and VM provisioning.</li>
<li>Design and build out the IaaS software layer for a new GB200 data center with thousands of GPUs.</li>
<li>Design and build distributed GPU scheduling and the global management plane that power on-demand and managed clusters across dozens of data centers</li>
<li>Develop infrastructure that powers our internal inference, RL, and fine-tuning products in addition to external cloud customers</li>
<li>Design and build systems that scale per-cluster capacity limits and automate the onboarding of new capacity</li>
<li>Work on a global multi-exabyte high
求めるスキル
CUDA
LLM
Kubernetes
AWS
GCP
Azure