プロンプトエンジニア 求人

Senior Software Engineer - Together Cloud Infrastructure

0万円 〜 0万円
San Francisco
正社員・契約社員
経験年数:
閲覧数:0

仕事内容

<h3>About the Role</h3> <p>Together AI is building the AI Native Cloud, an end-to-end platform for the full<br>generative AI lifecycle, combining the fastest LLM inference engine with state-of-the-art<br>AI cloud infrastructure. The Together Cloud team builds the Together GPU Clusters<br>product, which provides high-performance, AI-ready GPU clusters through a self-serve<br>cloud console and is the virtualized infrastructure layer powering Together’s inference,<br>RL, and fine-tuning products.</p> <p><br>As a Senior Software Engineer in Together Cloud Infrastructure, you will play a key role<br>in building the next generation AI cloud platform – a highly available, global, blazing-fast<br>cloud infrastructure that virtualizes cutting-edge ML hardware (GB200s/GB300s,<br>BlueField DPUs). You'll enable state-of-the-art ML practitioners with self-serve AI cloud<br>services, such as on-demand + managed Kubernetes and Slurm clusters, for both our<br>internal SaaS products (inference, fine-tuning, RL) and our external cloud customers,<br>spanning dozens of data centers across the world.</p> <h3><strong>Responsibilities</strong></h3> <ul> <li>Design, build, and maintain performant, secure, and highly-available backend&nbsp;services/operators that run in our data centers and automate hardware&nbsp;management, such as Infiniband partitioning, in. DC parallel storage provisioning,&nbsp;and VM provisioning.</li> <li>Design and build out the IaaS software layer for a new GB200 data center with&nbsp;thousands of GPUs.</li> <li>Design and build distributed GPU scheduling and the global management plane&nbsp;that power on-demand and managed clusters across dozens of data centers</li> <li>Develop infrastructure that powers our internal inference, RL, and fine-tuning&nbsp;products in addition to external cloud customers</li> <li>Design and build systems that scale per-cluster capacity limits and automate the&nbsp;onboarding of new capacity</li> <li>Work on a global multi-exabyte high

必須要件

求めるスキル

CUDA LLM Kubernetes AWS GCP Azure

勤務条件

勤務時間
雇用形態 正社員・契約社員
勤務地 San Francisco
リモートワーク 不可
Together AI 公式採用ページ掲載求人

この求人に応募する

18日前に掲載

公式ページで応募する

※ 企業の公式採用ページへ移動します

人気求人

他の人気求人をチェック

求人一覧を見る

メールアドレスで無料会員登録

正しいメールアドレスを入力してください
※半角英数記6~40文字
パスワードは6文字以上で入力してください
利用規約プライバシーポリシー をご確認のうえ、「同意して登録する」を押してください。
すでにアカウントをお持ちの方

求職者ログイン

初めての方
掲載企業様の方はこちら

企業様 新規登録

求人掲載をご希望の企業様向けの登録フォームです
正しいメールアドレスを入力してください
※半角英数記6~40文字
パスワードは6文字以上で入力してください
すでにアカウントをお持ちの方

企業ログイン

初めての方
求職者の方はこちら

パスワードリセット

ご登録いただいたメールアドレスを入力してください。
パスワードリセット用のリンクをメールでお送りします。

正しいメールアドレスを入力してください
アカウントをお持ちの方

企業様 パスワードリセット

ご登録いただいたメールアドレスを入力してください。
パスワードリセット用のリンクをメールでお送りします。

正しいメールアドレスを入力してください
アカウントをお持ちの方

新しいパスワードを設定

新しいパスワードを入力してください。

※半角英数記6~40文字
パスワードは6文字以上で入力してください