AI/ML Engineer
About the Role:
Reporting to the Director of Data & AI Engineering, the AI/ML Engineer is a mid-level, hands-on technical role responsible for building and maintaining the data pipelines, AI models, and intelligent features that power the GLOBO platform. This role spans the full lifecycle of AI development—from cleaning and preparing data, to building and evaluating models, to shipping production features that directly improve operational efficiency and customer experience.
The AI/ML Engineer works across GLOBO’s modern data stack (Fivetran, dbt, Snowflake) and AI infrastructure (AWS Bedrock, LLMs, agentic frameworks) to deliver reliable, well-tested solutions. This person is equally comfortable wrangling messy data and prompt-engineering an LLM, and takes pride in writing clean, tested code that other engineers can build on.
Data Engineering & Pipeline Development:
- Build and maintain reliable ingestion pipelines usingSnowflake Openflow, Python, and Snowflake, including API, PostgreSQL, and CDC-based integrations.
- Develop incremental synchronization, cursor/state management, retry logic, schema-drift handling, soft-delete propagation, and source-to-target reconciliation.
- Transform raw source data through staging, intermediate, and core models into trusted datasets for analytics, reporting, and machine-learning workloads.
- Apply data-quality checks for freshness, completeness, uniqueness, referential integrity, valid relationships, and business-rule compliance.
- Maintain source definitions, model documentation, lineage, metadata, and data contracts.
- Collaborate with data owners to ensure PII/PHI classification, masking, retention, and deletion requirements are implemented throughout the pipeline.
Model Development, Testing & Evaluation:
- Implement monitoring and alerting for ingestion failures, pipeline freshness, schema changes, data-quality failures, transformation errors, model drift, and inference degradation.
- Establish automated regression testing for dbt models, features, evaluation datasets, prompts, and model outputs.
- Validate that sensitive data is appropriately masked, redacted, access-controlled, and excluded from unauthorized model training or data-sharing workflows.
- Build safeguards for PII/PHI in recorded-call, transcript, and AI/ML processing pipelines, including verification that redaction and deletion workflows complete successfully.
- Ensure AI/ML outputs are traceable to their source data, model or prompt version, feature set, and evaluation results.
- Define recovery procedures, data-quality escalation paths, and operational runbooks for critical pipelines and models.
- Support human review and approval for model outputs that may affect customers, interpreters, employees, financial activity, or service quality.
Feature Development & Integration:
- Collaborate with Product and Engineering to ship AI-powered features into the GLOBO platform.
- Build and deploy LLM integrations (AWS Bedrock, Anthropic Claude) and agentic workflows (CrewAI, LangChain).
- Write production-quality code with proper tests, documentation, and error handling.
Reliability & Safety:
- Implement guardrails, monitoring, and alerting for AI services in production.
- Ensure AI outputs are consistent and trustworthy.
- Contribute to evaluation datasets, prompt versioning, and regression testing for deployed models.
Performance & Cost Optimization:
- Monitor and optimizeSnowflake compute, storage, query performance, dbt execution, Openflow runtime usage, and model-inference costs.
- Design efficient incremental models, CDC pipelines, materializations, clustering strategies, and warehouse/task schedules.
- Compare and optimize ingestion costs as Globo transitions from Fivetran to Snowflake Openflow.
- Reduce unnecessary full refreshes, duplicate processing, excessive data movement, and inefficient feature recomputation.
- Optimize model selection, prompt size, token usage, batching, caching, inference frequency, and routing between model providers.
- Measure model performance against operational cost, latency, throughput, and data-freshness requirements.
- Establish practical service-level targets for critical datasets, transformations, batch jobs, and model-serving workflows.