Manager, Machine Learning Engineering
Summary
Player-coach manager leading Tala's ML Platform team (4–6 MLEs), building the platforms and infrastructure that let Data Science securely train, deploy, monitor, and operate ML models at scale — focused on real-time inference and streaming systems, using Python, AWS/GCP, Kubernetes, Kafka, and Airflow.
The Role
We’re looking for a Manager, Machine Learning Engineering to lead Tala’s ML Platform team. This person will manage a team of Machine Learning Engineers responsible for building the platforms, frameworks, and infrastructure that enable our Data Science teams to securely train, deploy, monitor, and operate machine learning models at scale.
This is a player-coach management role. You’ll be responsible for developing and growing the team while also providing enough technical leadership to guide architecture, engineering practices, reliability, and production systems. The role has a particular focus on real-time machine learning inference and streaming data systems, as well as the platforms that support batch model development and deployment.
What You'll Do
Lead & Grow the Team
- Manage and develop a team of 4–6 Machine Learning Engineers across mid-to-senior levels.
- Hire, source, interview, and close strong MLE talent.
- Establish clear expectations, provide regular feedback, and create development plans for direct reports.
- Coach engineers toward growth and promotion while addressing performance gaps directly and thoughtfully.
- Create opportunities for engineers to take on challenging projects and grow their technical leadership.
Own Engineering Delivery
- Set quarterly goals and ensure the team consistently delivers against them.
- Own prioritization across product roadmap work, run-the-business activities, and operational excellence.
- Balance team capacity across new development, maintenance, technical debt, and production support.
- Improve team productivity by reducing context switching and delegating effectively.
- Partner with engineers and technical leads to estimate and scope complex work.
Provide Technical Leadership
- Guide the development of platforms and frameworks that allow Data Scientists and Analysts to explore data, develop features, and train, test, deploy, and monitor ML models.
- Provide technical leadership across model infrastructure, real-time inference, streaming feature extraction, batch processing, and production ML systems.
- Drive strong engineering practices around testing, automation, observability, fault tolerance, infrastructure-as-code, and deployment.
- Own and improve SLOs, on-call health, capacity planning, reliability, and incident response.
- Review technical designs and help drive architectural standards and technical debt reduction.
Partner Across the Organization
- Work closely with Data Science, Data Engineering, Data Platform, Product, Credit, and Business Development teams.
- Translate business and technical needs into scalable ML platform solutions.
- Coordinate dependencies and delivery across multiple engineering and data teams.
- Help create structure and clarity in an environment where priorities and requirements can evolve.
What You'll Need
Management Experience
- 2+ years of directly managing engineers, including hiring, performance management, coaching, and career development.
- Experience managing a team through at least one full performance cycle.
- Demonstrated ability to coach engineers toward promotion and address underperformance effectively.
- Experience owning team goals, prioritization, estimation, and delivery.
- Experience with production on-call, incident response, and capacity planning.
- Willingness to be actively involved in sourcing, interviewing, and closing engineering talent.
Technical Experience
- 6+ years of backend software engineering experience in consumer-scale applications.
- At least 3 years of hands-on Python experience.
- Experience building and operating machine learning or causal inference systems in production.
- Earlier-career experience personally building and deploying ML models or ML infrastructure.
- Ability to participate in technical architecture and system-design discussions and provide technical direction without needing to be the primary coder.
- Strong understanding of software quality, security, reliability, testing, and production operations.
Technical Skills
We’re particularly interested in candidates with experience across:
- Languages: Python, SQL
- Machine Learning: Jupyter, Pandas, Scikit-Learn, XGBoost, TensorFlow, PyTorch, Hugging Face
- Cloud & Infrastructure: AWS, GCP, Azure, Kubernetes, Docker
- Streaming: Kafka, Kinesis, Beam, Flink, Spark Streaming
- Batch Processing: Airflow, Metaflow
- Databases: MySQL, PostgreSQL, Cassandra, Snowflake, Druid, and/or similar technologies
- APIs: REST, GraphQL, gRPC, Protocol Buffers
- Production Engineering: DevOps, SLOs, monitoring/observability, on-call, capacity planning, root-cause analysis
- ML/Analytics: Machine learning, causal inference, scalable algorithms
Skills
- AI
- Airflow
- Analytics
- API
- Automation
- AWS
- Azure
- Business Development
- Cassandra
- Causal Inference
- Cloud
- Data Engineering
- Data Science
- DevOps
- Docker
- Fintech
- Flink
- GCP
- GraphQL
- gRPC
- Hugging Face
- Infrastructure as Code
- Jupyter
- Kafka
- Kinesis
- Kubernetes
- Machine Learning
- MySQL
- Observability
- pandas
- Performance Management
- PostgreSQL
- Python
- PyTorch
- scikit-learn
- Snowflake
- Spark
- SQL
- TensorFlow
- XGBoost
As published by lever · 8 questions · 7 written answers
Basics
Resume/CV, Full name, Email, Phone, Current location, Current company, LinkedIn URL, Twitter URL, GitHub URL, Portfolio URL, Other website, What is your age range?, I identify my ethnicity asSelect all that apply, What gender do you identify as?
Short answers (1)
- Will you now or in the future require sponsorship or transfer for employment visa status ?
Written answers (7)
- Why are you interested in joining Tala?
- How soon can you join us?
- This role is located in the US. Do you currently live in the US? optional
- How many years of direct-people management experience do you have? Have you directly managed engineers, including hiring, performance management, and growth/compensation conversations? optional
- Do you have at least 6 years of backend software engineering experience in consumer-scale applications, including at least 3 years working in Python? optional
- Are you comfortable working in a global/remote environment with meaningful overlap with US Pacific and East Africa working hours? optional
- Briefly describe a machine learning model or ML infrastructure you personally helped build and deploy to production. What was your role, and what did you own? optional