Lead Data Engineer
Summary
Lead Data Engineer (first senior data hire) for an AI platform in the property sector, architecting scalable real-time and batch data pipelines, vector search, and ML data workflows using Python, PostgreSQL, and distributed systems.
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Lead Data Engineer based in India.
As the first senior data hire, you will shape the data foundations behind an intelligent AI platform transforming operational workflows in the property sector. You will architect scalable infrastructure spanning real-time pipelines, analytics, vector search, and machine-learning data workflows. The role combines hands-on engineering with high-level architectural ownership and strategic decision-making. You’ll collaborate closely with AI and backend engineers, product teams, founders, and senior leadership. Your work will directly influence product capabilities, data reliability, and the long-term technical direction of the platform. You will also help establish engineering standards, data culture, and the future hiring bar as the function grows. This is an opportunity to have significant ownership in a fast-moving, AI-focused environment.
Accountabilities
- Architect, build, and scale robust data pipelines and infrastructure supporting both AI products and operational applications.
- Design and maintain data ingestion, transformation, processing, and storage architectures for batch and real-time workloads.
- Develop scalable systems for vector search, retrieval, and machine-learning data workflows.
- Build and optimise data models, distributed processing systems, and large-scale query infrastructure.
- Establish frameworks for data reliability, quality, security, governance, monitoring, and observability.
- Collaborate with AI and backend engineering teams to support model training, inference, and data-driven product capabilities.
- Contribute to technical architecture decisions, engineering standards, and the long-term data strategy.
- Partner with founders and product leadership to translate data capabilities into meaningful product and business decisions.
- Take ownership as the founding data specialist, helping define the team’s technical culture, standards, processes, and future hiring requirements.
- 7+ years of professional experience, with significant experience in dedicated data engineering roles and ownership of complex data infrastructure.
- Strong experience designing and building scalable data pipelines and distributed data systems.
- Solid experience with relational databases, preferably PostgreSQL; experience with MySQL or comparable technologies is also relevant.
- Experience working with NoSQL databases and modern vector databases used in AI applications.
- Strong Python programming skills, including experience with data-processing libraries such as Pandas or Polars.
- Demonstrated ability to make, communicate, and justify architectural decisions rather than simply implementing predefined solutions.
- Experience building scalable backend systems and designing robust data models and storage architectures.
- Strong understanding of data processing performance, scalability, and optimisation.
- Experience with technologies such as Apache Spark, Apache Airflow, Kafka, Elasticsearch, or OpenSearch is highly desirable.
- Experience with PostgreSQL, MongoDB, and vector technologies such as Qdrant, Milvus, or pgvector is highly desirable.
- Strong collaboration and communication skills, with the ability to work effectively with technical and non-technical stakeholders.
- Experience in AI/ML platforms, event-driven architectures, cloud infrastructure, or high-growth startup environments is a plus.
- Remote-first working environment from locations across India.
- Opportunity to take significant ownership as the first senior data engineering hire.
- High-impact role working at the intersection of data engineering, AI, and intelligent automation.
- Direct collaboration with AI engineers, backend engineers, product teams, founders, and leadership.
- Opportunity to shape the data architecture, engineering standards, culture, and future team.
- Exposure to modern technologies including real-time data pipelines, vector databases, machine-learning infrastructure, and distributed systems.
- Strong opportunities for professional growth and technical leadership in a rapidly evolving environment.