Principal AI Data Engineer
Summary
Principal AI Data Engineer builds and deploys GenAI and agentic AI systems in Azure, using Databricks, LangChain, and MLflow to prototype and scale AI solutions for enterprise clients in Trading & Supply.
Contract role: Principal AI Data Engineer
Key Responsibilities
- Develop and evaluate AI/GenAI/AgenticAI prototypes
using tools like Copilot Studio, AI Foundry and Copilot Analyst Agent, Mosiac
AI, Genie, AgentBricks, MLflow with a focus on quick wins and enterprise
integration.
- Build and tune Retrieval-Augmented Generation (RAG)
systems, including embedding model selection, prompt engineering, and traceable
evaluation.
- Design and deploy basic AI agents using frameworks
such as LangChain, AutoGen, and smolagents
- Communicate complex AI concepts clearly to business
stakeholders and cross-functional teams.
- Collaborate on E platform enhancements and work within
its current limitations.
- Deploy models and applications using Azure OpenAI, Azure
AI Foundry, Databricks Mosaic Gateway, and Docker.
- Follow DevOps best practices including CI/CD
pipelines, testing, linting, and GitHub workflows.
- Write modular, reusable code using OOP design patterns
in Python (Pydantic, PyTorch, etc.).
- Operate in agile teams and contribute to sprint
planning, reviews, and retrospectives.
- Deliver hands on GenAI/AgenticAI systems used directly
by commercial teams within Trading & Supply, taking solutions from
prototype to production
- Apply engineering skills (emphasis on Databricks) and
research skills across experimentation, rapid prototyping, and iterative
delivery. Someone who puts emphasis on reproducibility and open source, manages
large-scale text and structured datasets on Databricks.
- Build AI capability, manage stakeholders and
communicate effectively to ensure alignment between business needs and AI
solutions, and a quick understanding of commercial operations that happen in
T&S
- Design and run evaluation and testing frameworks for
GenAI systems, including benchmarking, reproducibility checks, and structured
model assessments
- Build solutions using Databricks infrastructure,
Genie, MLflow (deployment and tracing and evaluations), LangChain, and
LangGraph, and integrate them into scalable AI workflows and architectures
- Contribute to system planning, architectural design,
and structured testing to ensure long term reliability, performance, and
maintainability
- Preferably also someone who can set the building
blocks and lead building out the backlog
Required Skills
- Bachelor or Master or equivalent in Statistics,
Mathematics, Econometrics or similar discipline with at least 8-12 years’
experience on data science/AI projects.
- Deep understanding of LLM families (GPT, Llama,
Claude, Mistral) and their reasoning capabilities.
- Strong experience with Databricks- DLT,
Delta Lake concepts, UC governance.
- Solid understanding of streaming
technologies (e.g., Spark Structured Streaming, Autoloader)
- Programming skills in Python, SQL, or Scala.
- Proficiency in data modelling,
ETL/ELT processes, and data architecture.
- Strong analytical background with problem-solving
skills.
- Performance tuning concepts like watermarking, late
data handling, parallelism & checkpointing.
- Hands-on expertise in ADF, and Qlik
Replicate for data ingestion and replication.
- Experience working in Azure cloud
environments.
- Experience with GenAI evaluation frameworks and benchmarking
methodologies.
- Experience in MS Copilot, AI Foundry , Databricks
(MosiacAI, MLflow, Agentbricks, Genie)
- Strong Git practices and collaborative coding
standards.
- A passion for and expertise in practicing data science
to solve real-world problems.
- Excellent oral and written communication skills.
- Strong interpersonal skills and enthusiasm for
teamwork, as well as the ability to work independently.
- Familiarity with the enterprise AI platforms and
governance models is a plus.
- Strong decision-making abilities, using data-driven
insights to make informed choices that align with organizational goals.
- Skills in managing conflicts and facilitating
effective resolutions to maintain a positive and productive team dynamic.
- Ability to engage with and manage expectations of
various stakeholders, including executives, project managers, and other teams.
- Proficiency in identifying potential risks in data
projects and implementing strategies to mitigate them.
- Strong commitment and ownership of project delivery.