仕事内容
<p class="p1">Deeply understanding what’s in the enterprise data has been a challenge that Databricks has been addressing by providing analytics and machine learning tools. From data warehousing with Databricks SQL to large-scale distributed processing with Spark and advanced ML tools for experimentation and model serving, we empower our customers to gain insights and drive innovation.</p>
<p class="p1"> </p>
<p class="p1">To enable all of this on Databricks, making data ingestion seamless is crucial. That’s the mission of the Ingestion Core Team: to make the ingestion of all data—structured and unstructured—simple, reliable, and efficient. Simplifying the complex is hard, and that’s where you come in. This role requires building distributed platform systems to incrementally ingest high-volume, petabyte-scale data from diverse sources—including cloud storage (SQS, ADLS, GCS), <strong>databases (Oracle, SQL Server, MySQL, Postgres),</strong> and file sources (Google Drive, SharePoint)—at high throughput and low cost. The data includes structured formats (JSON, Parquet, CSV) as well as unstructured data (text, images, docs, PPTs, and blobs), all of which land in Delta Lake with schema evolution and change data capture (CDC) capabilities.</p>
<p class="p1"> </p>
<p class="p1">Join us in making data ingestion effortless and be part of the team that powers the future of AI + data at Databricks!</p>
<p class="p1"> </p>
<p class="p1">As an engineer on the team, you will work on projects that:</p>
<ul class="ul1">
<li class="li1">Build distributed infrastructure to ingest data from diverse sources and support streaming ingestion, incremental processing, and replication. This isn’t just about building plugin connectors.</li>
<li class="li1">Reduce end-to-end latency, increase throughput, and reduce costs from the time data appears in source systems to when it is available in Delta Lake.</li>
<li class="li1">Design and optimize streaming and distributed wo