仕事内容
<p>Scale is powering this generative AI wave by providing the data and infrastructure for companies to build large-scale foundation models. AI is rapidly changing the world, and Scale is growing to meet that rapid demand. Our customers include OpenAI, Microsoft, Adept, Stability AI and many more major players in this space!</p>
<p>The Platform team is responsible for building the core abstractions and infrastructure on which the products can be built and iterated rapidly. We are looking for an Operations-focused AWS Engineer with deep experience in infrastructure management, lifecycle maintenance, and security hardening. We are looking for Ops specialists who thrive on optimizing, securing, and maintaining the health of large-scale AWS environments. You have a growth mindset and are comfortable learning new technologies.</p>
<p>You Will: </p>
<ul>
<li><strong>Manage Infrastructure Lifecycle</strong>: Lead the end-to-end lifecycle of our AWS infrastructure, including routine patching, version upgrades, and system maintenance to ensure high availability.</li>
<li><strong>Vulnerability Management</strong>: Work closely with the Security Team to identify, prioritize, and remediate vulnerabilities across our cloud footprint.</li>
<li><strong>AWS Optimization</strong>: Continuously monitor and optimize AWS resource utilization for performance, reliability, and cost-efficiency.</li>
<li><strong>Automate Operations</strong>: Use scripting and automation tools to streamline repetitive operational tasks such as fleet-wide patching and configuration audits.</li>
<li><strong>Incident Response & Troubleshooting</strong>: Provide technical expertise for troubleshooting infrastructure-level issues and participate in operational health reviews.</li>
<li><strong>Security Compliance</strong>: Build and maintain systems that adhere to strict security standards, ensuring our environment remains compliant through response and proactive mitigations.</li>
</ul>
<p>Qualifications:
求めるスキル
Python
Kubernetes
AWS
GCP
Azure