Senior Software Engineer - Data License Analytics, Pune
- Build GenAI tools and AI agents — RAG pipelines, contextual embedding systems, and MCP tool servers that give LLM agents structured access to live pipeline data
- Build and maintain data ingestion pipelines into Apache Iceberg, orchestrated via Apache Airflow
- Develop semantic search capabilities using vector databases and hosted embedding models
- Build self-service ingestion APIs enabling partner teams to onboard to the platform without managing Iceberg or storage infrastructure
- Design microservices and backend APIs — with opportunities to contribute to React/TypeScript interfaces embedded in production tools
- Instrument systems with OpenTelemetry and contribute to anomaly detection capabilities on the near-term roadmap
- Languages & Frameworks: Python (3.11 / 3.12 / 3.13), TypeScript, FastAPI, React
- Data & Orchestration: Apache Iceberg, Apache Airflow, PyArrow, PyIceberg, Spark, Kafka, RabbitMQ, Parquet
- Databases: PostgreSQL, vector databases, Redis, Solr, distributed SQL
- Infrastructure: Docker, Kubernetes, NGINX, OpenTelemetry
- Cloud: GCP, AWS (S3, Redshift), Snowflake, Databricks
- AI & ML: RAG frameworks, embedding models, contextual chunking, semantic search, LLM integration
- Agent Frameworks: MCP (Model Context Protocol), Agent-to-Agent (A2A)
- 6+ Software engineering experience in production environments
- Proficiency in Python or a similar language
- Familiarity with relational databases and SQL
- Interest or experience in data engineering — pipelines, batch processing, or data lake technologies
- Understanding of distributed systems and service architecture
- Interest in AI/ML — particularly LLMs, RAG, or embedding-based search
- Experience with Apache Iceberg, PyArrow, or data lake engineering
- Familiarity with Apache Airflow or workflow orchestration
- Experience with vector databases, embedding models, or semantic search
- Background in anomaly detection or data quality frameworks
- Exposure to MCP, A2A, or agent communication protocols
- Experience with observability tooling — tracing, metrics, structured logging
- Curiosity about financial data and reliability engineering
- Ship GenAI tooling used daily in production — MCP servers, RAG pipelines, and AI agents for real incident triage
- Be a founding member of a greenfield engineering hub — shape the culture and standards from day one
- Collaborate with engineers across Dublin, London, and New York
- Work at scale — millions of securities, billions of data points, clients depending on it around the clock
- Join a team that values curiosity, learning, and measurable impact
If this sounds like you:
Skills
- Agentic AI
- AI
- Airflow
- Analytics
- Anomaly Detection
- API
- AWS
- Cloud
- Data Analytics
- Data Engineering
- Data Ingestion
- Data Lake
- Data Quality
- Databricks
- Distributed Systems
- Docker
- Embeddings
- FastAPI
- GCP
- Generative AI
- Iceberg
- Kafka
- Kubernetes
- LLM
- Machine Learning
- MCP
- Microservices
- Nginx
- Observability
- OpenTelemetry
- Parquet
- PostgreSQL
- Python
- RabbitMQ
- React
- Redis
- Redshift
- Semantic Search
- Snowflake
- Solr
- Spark
- SQL
- TypeScript
- Vector Databases
- Workflow Orchestration