data engineer for cloud cost management
Summary
Data engineer who designs, builds, and owns scalable data pipelines and ETL infrastructure on GCP (BigQuery, Dataproc, Dataflow) using SQL and Python/NodeJS, plus big data tools like Spark, to power analytics and business insights for Box's cloud content management platform. Hybrid role: at least 3 days per week from the Warsaw office.
Описание
Box is a leader in Intelligent Content Management. Its platform helps organizations collaborate, manage the content lifecycle, secure critical content, and transform business workflows with enterprise AI.
Задачи
- Identify business opportunities and design and build scalable data solutions;
- Build and own data pipelines that clean, transform, and aggregate data from disparate sources;
- Create and maintain optimal data pipeline architecture;
- Assemble large, complex data sets meeting functional and non-functional business requirements;
- Identify, design, and implement internal process improvements by automating manual processes, optimizing data delivery, and redesigning infrastructure for scalability;
- Build infrastructure for extracting, transforming, and loading data from diverse sources using GCP BigQuery and Spark;
- Build analytics tools that use data pipelines to provide actionable insights into operational efficiency and business performance metrics;
- Support Executive, Product, Data, and Design stakeholders with data-related technical issues and infrastructure needs;
- Create data tools that help analytics and data science team members build and optimize the product;
- Influence across teams and functions and build organizational best practices.
Требования
- 3+ Years of relevant industry or academic experience working with large amounts of data;
- Experience building and optimizing scalable data pipelines, architectures, and data sets;
- Awareness of emerging technology trends and opportunities to improve development processes;
- Expert SQL skills;
- Experience with Python or NodeJS;
- Experience with GCP, including BigQuery, Dataproc, and Dataflow/Fusion;
- Experience with big data tools such as Hadoop, Spark, or Kafka;
- Strong analytical skills for working with structured and unstructured datasets;
- Experience supporting and working with cross-functional teams in a dynamic environment;
- Familiarity with virtualization, container abstractions, and orchestration such as Kubernetes or Docker;
- Familiarity with visualization software such as Tableau;
- Familiarity with frontend web frameworks such as React;
- Nice to have: Scala or Java.
Условия
- Work from the assigned office at least 3 days per week;
- Equal opportunity employer with a commitment to diversity.