Backend Engineer, Data
Summary
Designs and maintains data pipelines, models, and tools using Scala, Spark, and Airflow to power Stripe's product, data science, and GTM teams.
Who we are
About Stripe
Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world's largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a staggering amount of work ahead. That means you have an unprecedented opportunity to put the global economy within everyone's reach while doing the most important work of your career.
About the team
The Data Foundations team drives Data Engineering and Data Apps and Tooling work across Stripe, enabling Stripes to leverage data to make informed decisions and build user-centric products. We provide tools and infrastructure to move, store, process, and analyze data, both at rest and in motion. We are looking for talented data-minded software engineers to help us manage business-critical data leveraged across the entire organization. If you are passionate about data, excited about designing data pipelines and data-driven user experiences, and motivated by having an outsized impact on the business, we want to hear from you.
Team Matching: exact team matching for one of the subteams will begin during final stages. Please note we may also consider you for different orgs based on your experience, location, etc. More information on our team matching process can be found here.
What you'll do
Every record in our data warehouse is vitally important for the businesses that use Stripe, so we're looking for people with a strong background in software engineering and data to help us scale while maintaining correct and complete data. You'll work with a variety of internal teams across Product, Data Science, and GTM to help them solve their data needs. Your work will provide visibility into how these stakeholders and the Data Foundations organization are performing and how we can deliver a better experience to Stripe's customers.
Responsibilities
- Design, develop, and own data pipelines, models, and products that power the Product, Data Science, and GTM functions
- Develop strong subject matter expertise and manage the service level agreements for both data pipelines and full stack web applications that support these critical stakeholders
- Build and refine Stripe's data foundations - infrastructure, pipelines, and tools to enable various teams at Stripe - working with Scala, Spark, and Airflow
- Leverage LLMs and Agents at scale to produce high-quality data on ambiguous problems
- Refine our existing data marts that help the GTM organization forecast the future potential performance of the business and reliably measure ongoing attainment toward targets
- Build data services that track key product metrics and measure the impact of different strategies employed by teams in the field
- Our tech stack is Spark, Scala, Java, SQL, and Python - and while we don't expect everyone on the team to be an expert in all of these, you will work across all of these technologies throughout your tenure on the team.
Who you are
We're looking for someone who meets the minimum requirements to be considered for the role. If you meet these requirements, you are encouraged to apply. The preferred qualifications are a bonus, not a requirement.
Minimum requirements
- Must have 6+ years of experience in a Software Engineering role, with a focus on building and maintaining data services or data-intensive applications
- A strong engineering background and an interest in data
- Prior experience with writing and debugging data pipelines using a distributed data framework (Spark, Hadoop, Pig, etc.)
- An inquisitive nature in diving into data inconsistencies to pinpoint issues and resolve deep-rooted data quality issues
- Knowledge of a backend development language (such as Scala, Java, or Go) and strong SQL experience
- The ability to communicate cross-functionally, derive requirements, and architect shared datasets
Preferred qualifications
- Experience creating and maintaining data marts to power business reporting needs • Experience working with Product or GTM (Sales and Marketing) teams
Skills
As published by greenhouse · 16 questions
Basics
First Name, Last Name, Email, Phone, Resume/CV, Cover Letter, Location
Short answers (3)
- Who is your current or previous employer?
- What is your current or previous job title?
- What is the most recent school you attended?
Pick from a list (13)
- Please select the country where you currently reside.
- Please select the country or countries you anticipate working in for the role in which you are applying.
- Do you have more than 6 years of industry experience? Not including internship, co-op or academic.
- Do you live in the bay area, California?
- Do you live in Seattle OR New York?
- Do you have less than 2 years of industry experience (not including of internship, co-op, and academic)?
- Are you authorized to work in the location(s) you selected in your previous response?
- Will you require Stripe to sponsor you for a work permit now or in the future for the location(s) you selected in in your previous response?
- If this role offers the option to work from a remote location, do you plan to work remotely?
- Have you ever been employed by Stripe or a Stripe affiliate?
- Do you opt-in to receive WhatsApp messages from Stripe Recruiting?
- To ensure we can focus fully on our conversation, we use BrightHire to record and transcribe interviews. This helps us get to know you better without the distraction of note-taking. We also use BrightHire to help us analyze interview notes and for internal Stripe interview quality training and assurance purposes. Please confirm if you consent to BrightHire being enabled for this interview and for your personal data to be collected and processed for the purposes mentioned above, by responding below.
- What is the most recent degree you obtained?
Y Combinator