Data Infrastructure Engineer for AI Data Pipelines
Summary
Build and maintain petabyte-scale audio data pipelines for AI model training on GCP, sourcing new datasets and optimizing ingestion for cost, throughput, and quality.
Speechify is hiring for the Data side of its AI team to own data collection for model training. You will help build petabyte-scale datasets by tightly integrating infrastructure, engineering, and research, and join a solid software-engineering focus.
You’ll source new audio data, maintain ingestion pipelines on GCP, and collaborate with scientists to optimize cost, throughput, and quality while shaping Speechify’s data roadmap for scalable products.