AI Engineer/ Data Scientist
This is a remote position.
- Design, build, and deploy AI models, tools, and agents to automate intelligence-heavy workflows in investment research, portfolio management, operations, and client reporting.
- Partner with the Head of Data & AI to define the technical architecture for AI use cases, integrating with Snowflake, Databricks, internal APIs, and event streams in a secure and governed way.
- Implement retrieval augmented generation (RAG) pipelines and prompt conditioning patterns that ground LLMs on Marathon’s proprietary data, documents, and knowledge assets.
- Design and implement Model Context Protocol (MCP)–based integrations so AI assistants and agents can securely discover and connect to internal systems, databases, and services through standardized MCP servers and clients.
- Build and maintain MCP servers that wrap key enterprise services (data warehouses, document stores, workflow systems) and expose tools, resources, and prompts to AI clients in a standardized way.
- Establish patterns for MCP host/client configuration, access control, and observability to ensure reliable, auditable AI interactions with enterprise systems.
- Implement MLOps and LLMOps practices for both model and MCP-based integration lifecycles, including deployment automation, monitoring, logging, and rollback strategies.
- Collaborate with data engineers and platform teams to ensure clean, secure, and well-structured data access for AI consumption, including governance of which systems are exposed via MCP.
Requirements
- 6+ years of hands-on Python — you write production code, build packages, and care about performance and readability.
- Proven experience building and shipping LLM-powered applications (not just wrappers — you understand what's happening under the hood).
- Deep understanding of RAG: dense retrieval, sparse retrieval, hybrid, re-ranking, context window management, and evaluation.
- Strong prompt engineering skills: you know when to use CoT, how to design structured outputs, and how to debug hallucinations systematically.
- Hands-on AWS experience — SageMaker, S3, Lambda, CloudWatch, and ideally Bedrock or SageMaker JumpStart.
- Solid ML fundamentals: model training, evaluation, bias/variance, feature engineering, and experimental design.
- Comfortable with embedding models, tokenization internals, KV-cache, quantization trade-offs, and latency optimization.
- Experience with agentic frameworks (LangChain, LlamaIndex, CrewAI, or custom tool-use implementations).
- Fine-tuning experience — LoRA, QLoRA, instruction tuning, RLHF familiarity.
- Familiarity with evaluation frameworks: RAGAS, DeepEval, or building custom eval harnesses.
- MLOps tooling: MLflow, DVC, Weights & Biases, or SageMaker Pipelines.
- Knowledge of streaming inference, async serving, and cost optimization for token-heavy workloads.
- Prior work in analytics, BI, or domain-specific NLP (finance, healthcare, e-commerce).