Senior AI/ML Engineer
Summary
Staples is hiring a Senior AI/ML Engineer in Chennai to architect and scale production agentic AI systems — designing multi-agent orchestration, RAG pipelines, and LLM fine-tuning while mentoring engineers and setting best practices. Core stack: Python, LLM APIs (OpenAI, Anthropic, Azure OpenAI), and agentic frameworks like LangChain and CrewAI.
- Architect end-to-end agentic AI systems including agent design, orchestration, and integration with enterprise systems
- Design advanced agent patterns including hierarchical agents, multi-agent collaboration, and dynamic tool management
- Lead the development of sophisticated reasoning architectures with planning, reflection, and self-correction mechanisms
- Establish and implement evaluation frameworks, bench-marking methodologies, and cost optimization strategies for agentic systems
- Build scalable prompt engineering pipelines and fine-tune LLMs for specific use cases and domains
- Design and implement RAG systems with advanced retrieval strategies, knowledge management, and semantic search
- Develop production-grade monitoring, observability, and error handling for AI systems at scale
- Mentor junior engineers on agentic AI patterns, best practices, and system design principles
- Lead technical design reviews and architectural decisions for AI/ML initiatives
- Collaborate with product, infrastructure, and security teams to ensure responsible AI deployment
- Drive innovation in agentic AI approaches and evaluate emerging frameworks and models
- Optimize system performance, latency, and costs across production agentic workflows
- Establish CI/CD and testing practices specific to AI/ML systems
Requirements
- 5-7 years of professional software development or machine learning experience, with 2+ years focused on agentic AI
- Proven track record building production agentic AI systems with measurable business impact
- Expert-level proficiency in Python and software engineering best practices (design patterns, testing, documentation)
- Deep experience with multiple LLM providers and APIs (OpenAI, Anthropic Claude, Azure OpenAI, open-source models)
- Advanced knowledge of prompt engineering, prompt chaining, chain-of-thought reasoning, and few-shot learning
- Hands-on expertise with agentic frameworks (LangChain, CrewAI, Autogen, Semantic Kernel, or equivalent)
- Strong understanding of tool use, function calling, and dynamic agent capability management
- Experience architecting retrieval-augmented generation (RAG) systems with vector databases and semantic search
- Knowledge of model fine-tuning, domain adaptation, and custom model training
- Proficiency with asynchronous programming and handling high-concurrency systems
- Experience with API design and integration patterns for complex distributed systems
- Strong background in software testing, including evaluation frameworks for AI systems
- Experience deploying and managing agentic systems in production at scale
- Cloud platform expertise (Azure, AWS, GCP) for AI/ML services and infrastructure
- Background building observability, monitoring, and debugging tools for ML systems
- Experience with multi-agent systems, agent communication protocols, and collaborative agent patterns
- Knowledge of reinforcement learning from human feedback (RLHF) or similar alignment techniques
- Familiarity with knowledge graphs, semantic databases, and advanced retrieval strategies
- Experience optimizing token usage and managing costs for LLM-based applications
- Background in enterprise software or complex system integration
- Experience with MLOps, model versioning, and continuous integration for AI systems
- Leadership or mentoring experience with junior engineers
- Published research, papers, or significant open-source contributions in AI/ML
- Experience with agent-as-a-service architectures or API platforms
- Background in building customer-facing AI products
- Knowledge of adversarial testing and robustness evaluation for AI systems
- Experience with specialized domains (financial services, healthcare, legal, etc.)
- Understanding of AI safety, alignment, and ethical considerations in agent design
- Experience building domain-specific large language models or specialized agent systems